Interactive demo · based on OpenAI's textGrain report (Oct 2026)

textGrain watermark, simulated

A watermark hides a statistical fingerprint in which tokens the model picks. textGrain does it with a secret key, a small optimal-transport problem per token, and one knob: how much sampling randomness (entropy) you are willing to give up.

1 · GroupThe key and the last few tokens split the vocabulary into secret blocks.
2 · TiltA keyed table of random scores says which blocks are "favored" right now. Optimal transport tilts block probabilities toward them, within the entropy budget β.
3 · TestThe detector recomputes the scores from text + key alone. Watermarked text keeps landing on high scores; human text doesn't.
Part 1 · one generation step

Watermark the next word

Prompt: “The morning was ___”. The model's own probabilities are fixed; change the key or β and watch the sampling distribution shift.

a · model probabilities P, coloured by secret block

Block membership comes from hash(key, context, token). Without the key it looks random.

b · block mass ρ = sum of P inside each block

H(P) = nats. Budget to remove: β·H(P) = .

c · coupling π(block, column) from entropy-constrained OT

Cell shade = keyed Gumbel score (darker = favored). Circle area = π. The key picks one column (outlined). Each row still sums to ρ and each column to 1/m, so averaged over keys nothing is biased.

d · what actually gets sampled, Q(w | column)

Bar = Q. Thin dark line = original P. Entropy removed: nats ( of H(P)).

Same key, same context. Still varied unless β is high.
Part 2 · detection

Spot it in a whole passage

A toy model writes two passages from the same prompt: one with textGrain on, one with plain sampling (stand-in for human or other-model text). The detector sees only the words and a key.

running evidence: Sn − n  (each score Y averages 1 on unwatermarked text)
watermarked passage unwatermarked passage threshold, false-positive rate 0.1%

Watermarked

Unwatermarked

Word shading = score Y for that position (darker = more watermark evidence). Struck-out words repeat an earlier context window, so they are skipped to keep the test honest. Try β = 0, a short passage, or a wrong detector key.

Why it is built this way
β is a literal priceThe KL penalty in the OT problem equals the mutual information between token and key, which is exactly the average entropy lost. β = 0.2 means 80% of the sampling randomness survives, on average across keys.
Unbiased on averageCoupling marginals are pinned to the model's distribution, so averaged over keys the text distribution is unchanged.
Still diverse under one keyGumbel-max watermarks always pick the same token for a fixed key and context. Here, only part of the randomness is spent, so regenerating gives different answers.
Blocks keep it cheapOT runs on B blocks × m columns, not the full vocabulary. Within a block, tokens keep their original relative odds.
Detector is blind to the modelScore Y = −log(1 − F(Z)) is Exponential(1) on unwatermarked text, so the sum is Gamma(n,1). No model, no β needed.

Part 1 uses B = 3 blocks and m = 4 columns like the report's figure; Part 2 uses B = 4, m = 8. Both use a 2-token context window, a 40-word toy vocabulary and a bisection search for λ instead of the report's multiplicative update. The real system runs on full LLM vocabularies.

Full explanation, maths and sources: textGrain explained: how OpenAI watermarks ChatGPT text. Built by Jaskamal Kainth.