A watermark hides a statistical fingerprint in which tokens the model picks. textGrain does it with a secret key, a small optimal-transport problem per token, and one knob: how much sampling randomness (entropy) you are willing to give up.
Prompt: “The morning was ___”. The model's own probabilities are fixed; change the key or β and watch the sampling distribution shift.
Block membership comes from hash(key, context, token). Without the key it looks random.
H(P) = nats. Budget to remove: β·H(P) = .
Cell shade = keyed Gumbel score (darker = favored). Circle area = π. The key picks one column (outlined). Each row still sums to ρ and each column to 1/m, so averaged over keys nothing is biased.
Bar = Q. Thin dark line = original P. Entropy removed: nats ( of H(P)).
A toy model writes two passages from the same prompt: one with textGrain on, one with plain sampling (stand-in for human or other-model text). The detector sees only the words and a key.
Word shading = score Y for that position (darker = more watermark evidence). Struck-out words repeat an earlier context window, so they are skipped to keep the test honest. Try β = 0, a short passage, or a wrong detector key.
Part 1 uses B = 3 blocks and m = 4 columns like the report's figure; Part 2 uses B = 4, m = 8. Both use a 2-token context window, a 40-word toy vocabulary and a bisection search for λ instead of the report's multiplicative update. The real system runs on full LLM vocabularies.
Full explanation, maths and sources: textGrain explained: how OpenAI watermarks ChatGPT text. Built by Jaskamal Kainth.