Dominic Chiappe · People, capability & transformation

Thinking about how organisations perform in an AI-enabled world

AyEye Today ·

The watermark is not the proof. The detector is.

OpenAI is moving text provenance inside the model's sampling process, but its own tests show the signal can collapse under ordinary editing and its detector will initially be restricted. Trustworthy provenance therefore needs due process: disclosed test conditions, corroborating evidence, independent interpretation and a route to challenge the result.

Bold painterly scene of an editor comparing two sheets as a long paper ribbon passes through an inspection arch and a distant operator reveals a hidden dotted pattern with a narrow beam of light
Original illustration · Chiappe × OpenAI.

LEAD SUMMARY — ANALYSIS. A watermark sounds reassuringly physical. You picture a mark pressed into paper: fixed, visible against the light and difficult to argue with. Text generated by an AI is less accommodating. Its mark is a statistical tendency hidden among word choices, readable only with the right detector and vulnerable to the ordinary human act of editing.

On 5 October, OpenAI began offering text watermarking to API customers and said it would add the technology to eligible ChatGPT and Codex output in the European Union. The launch matters because a regulatory demand for machine-readable provenance is now reaching inside model sampling. But OpenAI's own results also make the limit unusually clear: a detected watermark is evidence of a signal, not a verdict about authorship, ownership, accuracy or responsibility.

The useful mental shift is from watermarking to provenance due process. If a statistical mark can affect a student's grade, an employee investigation, a publisher's decision or a platform sanction, the institution needs more than a detector. It needs a disclosed threshold, corroborating evidence, accountable interpretation and a route to challenge the result.

1. Regulation has entered the sampling loop

CONFIRMED — PRIMARY SOURCE. OpenAI says its new system, textGrain, adds an invisible statistical signal to word choices. API customers worldwide can opt in for selected models from 5 October; eligible ChatGPT and Codex text in the EU will be watermarked over the following weeks. The company links the regional rollout to the EU AI Act's requirement that generated text be identifiable in machine-readable form. OpenAI: approach to EU text provenance rules

CONFIRMED — TECHNICAL REPORT. The accompanying paper describes a keyed process: pseudorandom values derived from a secret key and the preceding context influence token sampling, while a detector uses the text and the key to test for statistical dependence. The design assigns an explicit budget to the sampling entropy surrendered in exchange for a stronger watermark signal. textGrain technical report

ANALYSIS. This is a quiet architectural change with a large implication. Compliance is no longer something applied only after generation as a label, policy or audit trail. It can alter the probability path through which the output is produced. The law has entered the inference loop.

That does not mean the watermark makes each answer worse. OpenAI reports no meaningful performance difference across its listed benchmarks, although those are the company's evaluations rather than an independent audit of every use. It means provenance now has an engineering budget: enough dependence on a key to be detectable, but not so much constraint that useful variation or quality is damaged.

2. The signal weakens precisely where human work begins

CONFIRMED — PRIMARY SOURCE. At a target false-positive rate of 1%, OpenAI says its detector found watermarks in about 80% of 200-token psychology passages and about 95% of 400-token passages. Results were substantially lower for constrained material such as mathematics. In another evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%; replacing 25% reduced it to 17%.

IMPORTANT LIMIT. These figures describe particular English-language evaluations, models, passage lengths, content types and attack methods. They are not universal detection rates. OpenAI explicitly warns that strong performance under ideal conditions does not guarantee reliable detection in everyday use.

ANALYSIS. The awkwardness is almost comic: the more a person rewrites, translates, shortens or combines a model's draft, the less confidently the technical system may recognise the model's contribution. The detector is strongest when the text remains close to machine output and weaker when authorship becomes genuinely mixed.

That reverses a common intuition. Human editing does not merely make provenance harder to see; it is also the activity institutions are often trying to distinguish. A teacher, employer or publisher may care deeply about the difference between unreviewed model output, heavily edited assistance and a human-authored text polished by a tool. A binary watermark result cannot measure that contribution.

Concept to learn today: Provenance due process

PROVENANCE DUE PROCESS is the set of evidentiary and procedural safeguards needed before an AI-origin signal is used to make a consequential judgement about a person or piece of work.

Preserve the artefact

Keep the original text, relevant versions and collection context. Pasting a fragment into a detector should not erase how the material was obtained, edited or combined.

Name the test

Record the detector, key version, supported models, language, passage length, threshold and date. A result without its test conditions is difficult to reproduce and easy to overstate.

Corroborate the signal

Use platform records, document history, declared tool use and other provenance evidence where lawful and proportionate. A watermark should add weight to a case, not become the whole case.

Separate detection from judgement

The operator should report what the detector found and its known error profile. A different accountable person or process should decide what the result means under the relevant policy.

Provide a route to challenge

Allow retesting, disclosure of the applicable method and human review before a high-impact sanction. False positives and false negatives are not edge cases that disappear because the interface displays a confident badge.

ORIGINAL SYNTHESIS. The detector is becoming a new institutional witness. It observes a hidden statistical relationship and offers an opinion about origin. Like any witness, its evidence has conditions, blind spots and a chain of custody. Provenance becomes trustworthy only when the institution can explain how that witness was heard.

3. The secret key creates an evidence asymmetry

CONFIRMED — PRIMARY SOURCE. OpenAI is accepting applications for detector access, initially limited to approved researchers and expert organisations. It says the restriction reflects the risk of missed watermarks and false positives. The detector will report whether it finds an OpenAI watermark without identifying the user or revealing prompts or conversations.

ANALYSIS. Restriction is defensible during an early deployment: uncontrolled access can enable gaming, encourage misuse or turn unreliable results into a public accusation machine. Yet restriction also concentrates interpretive power. The provider controls the key, the detector, the calibration and the conditions under which outsiders may test claims.

An independent September paper made this governance problem its central point. Studying the open implementation of SynthID-Text rather than textGrain, the authors argued that neither public objections nor vendor assurances could be fully checked without access to deployed systems. They proposed matched outputs, configuration disclosure, accredited audits, a shared evaluation protocol and interoperable detection. Nemecek, Chaudhary and Ayday: Watermarks Without Verification

IMPORTANT DISTINCTION. That paper does not evaluate OpenAI's new system and should not be read as evidence that textGrain fails. It supplies a governance test: when detection depends on a provider-held key, what evidence can another institution independently verify?

INFERENCE. Detector access will become a form of delegated authority. Universities, platforms, courts, publishers and auditors may need certified routes to use or challenge vendor-specific detectors. The hard interoperability problem is therefore not only whether several watermarks can coexist. It is whether an adverse decision can be reproduced across institutional boundaries without distributing the secret needed to defeat the mark.

4. A provenance result is narrower than the question people ask

CONFIRMED — PRIMARY SOURCE. OpenAI states that a watermark does not measure human contribution, establish ownership or responsibility, identify the user, or verify accuracy. Absence of a detected watermark does not prove human authorship: the text may be too short, edited, translated, produced by an unsupported model, generated before the system was introduced or made by another provider.

ANALYSIS. Most real disputes ask broader questions. Did this student do the required work? Did an employee disclose assistance? Is a passage accurate? Was a contract term authorised? Who is responsible for a harmful claim? The detector answers none of them directly. It answers a smaller technical question about whether a particular keyed statistical pattern is present strongly enough to cross a threshold.

That smaller answer can still be useful. It may support content labelling, platform triage, research into synthetic media or investigation of coordinated influence operations. Trouble begins when a narrow origin signal is promoted into an all-purpose theory of authorship. A smoke alarm is valuable because it detects smoke; it becomes dangerous when somebody asks it to identify the arsonist.

5. Provenance needs layers, not a universal badge

CONFIRMED — PRIMARY SOURCE. OpenAI says no single provenance technique is sufficient. Its broader approach combines metadata standards, invisible watermarks and verification tools. Content Credentials can record origin and history, while a watermark may survive after metadata is stripped. Policies, abuse detection, reporting, investigations and enforcement supply further context.

ANALYSIS. These layers answer different questions. Metadata can describe a file's declared history but may be removed. A watermark may persist inside content but has probabilistic detection and can weaken through transformation. Platform logs can link activity to accounts but raise privacy and access concerns. Human review can interpret purpose and responsibility but is slower and inconsistent.

INFERENCE. The sensible architecture is cumulative rather than binary. Low-stakes labelling might rely on one signal. A disciplinary, contractual or legal decision should require independent corroboration and a documented path from artefact to conclusion. The higher the consequence, the less acceptable it is for a single opaque detector to carry the burden.

Noise: invisible does not mean personal tracking

NOISE CHECK. A secret-key watermark can sound like an individual identifier hidden in every sentence. OpenAI says textGrain does not associate a passage with a person, organisation, account, prompt or conversation. The technical report describes a signal for detecting keyed statistical dependence, not a payload containing identity.

That claim still deserves independent scrutiny as deployment expands, especially around key management and detector governance. But concerns should match the mechanism. The immediate risk is less a covert name tag in the prose than institutional overconfidence in a probabilistic signal.

Mental-model update

Yesterday: frontier AI needs high-reliability institutions with independent stop authority, diverse safeguards and the ability to interpret weak signals before discovery outruns control.

Today add: those institutions also need rules for evidence produced by the model ecosystem itself. A weak signal can guide inquiry without becoming a verdict. The quality of the institution will be visible in whether it distinguishes detection from judgement, makes thresholds contestable and gives affected people a practical route to challenge the machine's testimony.

In the illustration, the mark becomes visible only inside a beam controlled behind a partition. The editor can hold up two pages; the waiting figure can carry an appeal; neither controls the light. That imbalance is not an argument against watermarking. It is the reason provenance needs procedure.

Questions to carry forward

  • Which decisions may use an AI watermark as a screening signal, and which require independent corroboration?
  • What detector details must be recorded so a result can be reproduced or challenged later?
  • Who qualifies for detector access, and who audits the organisations that receive it?
  • How should policies distinguish unedited model output from substantial human revision or translation?
  • What appeal process exists when a probabilistic provenance result causes a real-world sanction?