deslopify.me

← Research

Research

How text watermarking works

A watermark in text can live in two places: in the statistics of word choice, put there by the model provider at generation time, or in characters you can't see, injected into the text itself. The two have very different robustness stories. Here is what the published research says about each — and what our diagnostic and rewriter actually do about them.

A signature hidden in word choices

Kirchenbauer et al., ICML 2023 (arXiv:2301.10226) introduced the scheme most later work builds on. At every generation step, a pseudorandom function seeded by the preceding tokens splits the vocabulary into a "green" and a "red" list, and the model is nudged to prefer green words. Any single word choice is unremarkable; over a few hundred words, the surplus of green tokens becomes a statistical signature. The text reads normally. Nothing is visible. Detection requires the secret key that generated the green lists — which means the model provider, not the reader, decides who can check.

SynthID-Text: the production version

Dathathri et al., Nature 2024 (Nature 634, 818–823) describes SynthID-Text, Google DeepMind's deployment of this idea in Gemini — the first statistical text watermark shipped at consumer scale, and open-sourced for other providers. The paper is candid about the limits: detection stays on the provider's side (a reader without the key can verify nothing), reliability drops on short or low-entropy text — factual answers, translations, code — and robustness is evaluated under mild edits, not full rewrites.

Who watermarks text today (August 2026)

Google runs SynthID-Text in Gemini production and operates a gated SynthID Detector portal — text checking remains provider-side, behind a waitlist. Anthropic announced on August 11, 2026that Claude models released after August 2, 2026 will carry an invisible text watermark at the model level, API included, with no opt-out; the mechanism is undisclosed and no public detection tool exists yet. This supersedes the earlier state of the Claude API and matters for this site's own compliance record. OpenAI still ships text unmarked — its detector-grade text watermark has been built and withheld since 2024; its provenance work covers images and audio via C2PA, a standard that by its own spec cannot travel with pasted plain text.

One fact frames everything on this page: without the provider's key, no one can check a single pasted text for a modern statistical watermark. For well-designed schemes this is not an engineering gap but a proof — Christ, Gunn & Zamir (arXiv:2306.09194) construct watermarks that are computationally undetectable without the key. The published black-box tests that do work (arXiv:2405.20777) interrogate the generating model with hundreds of chosen queries — not a paste box. Any tool that claims to see a SynthID or Claude watermark in your paste is selling something it cannot do. That includes us: our limitations page says so in as many words.

The crude cousin: characters you can't see

The other family doesn't touch the model at all. It hides marks in the text itself: zero-width spaces between letters, direction-control characters, filler codepoints, or Latin letters swapped for identical-looking Cyrillic ones (homoglyphs). Some tools use these as crude provenance marks or copy-paste trackers; the same characters are also used adversarially — a zero-width space inside a word splits it for any tokenizer-based detector, and the bidirectional override characters are the primitive behind the "Trojan Source" attacks described by Boucher & Anderson (arXiv:2111.00169). No regulation treats these as a sanctioned marking scheme, and they are trivially fragile: retyping the text, or any normalization pass, removes them.

This is the part this site acts on. The diagnostic's Provenance panel counts invisible characters, direction controls, steganographic spacing characters, and homoglyph-spoofed words (via a curated confusables table — whole words in a foreign script never flag) in every paste — they never change your score, because they say nothing about how the text reads. The rewriter strips them from its input and maps spoofed letters back to Latin.

How much survives paraphrase?

Krishna et al., NeurIPS 2023 (arXiv:2303.13408) put detectors against DIPPER, a paraphrase model. Classifier-style detection collapsed (DetectGPT: 70.3% to 4.6% true positives at a 1% false-positive rate). Watermarking was the most robust method tested — and still lost roughly a fifth of its detections to a single paraphrase pass. Sadasivan et al. 2023 (arXiv:2303.11156) pushed further: recursive paraphrasing degrades watermark detection toward chance, and a spoofing attack can even forge a watermark onto text a model never wrote.

The counterpoint is real and deserves the same visibility. Kirchenbauer et al.'s reliability follow-up (arXiv:2306.04634) found that watermarks survive both machine and human paraphrase better than early results suggested — given enough text. Detection is a function of length: a few hundred words of paraphrased output often still carry a readable signature; a tweet-sized snippet doesn't. The SynthID-Text paper reports the same shape. Short, heavily-edited text is where every scheme is weakest; long, lightly-edited text is where watermarks genuinely work.

What this means for deslopify

Our rewriterdoesn't paraphrase sentence by sentence; it regenerates the text with a different model, preserving meaning, facts, and direct quotes. A statistical watermark lives in the original model's word choices, so regenerated text no longer carries it — except in whatever is quoted verbatim. That is a property of paraphrase established in the literature above, not a feature we built, and we don't measure, promise, or optimize any of it.

The reason is the same one behind our position on AI detection: provenance should come from disclosure, not from an arms race. If you rewrite AI-generated text and publish it to inform the public, EU transparency rules may still require you to disclose that — a watermark disappearing does not make the duty disappear. The rewriter says this next to its output, and our Article 50 guide and marking guidecover what disclosure actually requires. Watermarks are one honest answer to "where did this text come from". Saying so yourself is a better one.