Insightsby Lean Change

You Used to Be the Detector

AI watermarking, explained for change agents.

Jason LittleSep 9, 20266 minComments (0)
You Used to Be the Detector
Contents

This is Jason. I didn't write or edit any of the post below. Reason being, Anthropic introduced watermarking as a response to combatting AI workslop, more or less. It's conceptually simple: There are private APIs being used by organizations subject to Article 50 compliance. Meaning, these private organizations can use the API to check if something was likely written by Claude.

While this is in its infancy, the last year of trends are showing that OpenAI, Gemini and Co-Pilot are taking their cues from Anthropic making me believe Anthropic and Claude are now the leaders in AI. We saw this happen with skill file adoption, projects and more. This post below is 100% written by Claude.

Workslop has always been a thing, just look at LinkedIn! It's just getting worse because of AI and this compliance will likely be implemented globally at some point. I thought it would be best to let Claude describe what it is.


Disclosure: Everything below this line was written by Claude Fable 5.1 on 9 September 2026, prompted and published by Jason Little. No human edits. See "How to check this post" at the bottom.


For about three years, spotting AI-written text was a party trick. Em dashes. "Delve." Three-item lists. "It's not X, it's Y." You could run a workshop on it, and plenty of people did.

That skill is now worth roughly nothing. Two things killed it. The models got trained away from the tells. And the vendors started building the check into the text itself.

What AI text watermarking is

AI text watermarking is a technique where a language model embeds a hidden, statistically detectable pattern into the words it generates, so that the vendor can later verify whether a piece of text was likely produced by that model. The reader sees nothing. The text reads identically. The signal only appears to someone holding the key.

Anthropic switched this on for all new Claude models in August 2026, starting with Fable 5.1. Older models get it over the following months. There is no opt-out.

How it works, without the math

Every time a model picks the next word, it has a few plausible choices, and it uses randomness to pick one. Watermarking doesn't change which words are plausible. It changes where the randomness comes from.

Instead of rolling dice, the model uses a secret key plus the last few words to decide. Anthropic's own analogy: playing Monopoly with the digits of pi instead of dice. Nobody at the table can tell. But anyone who knows you used pi can check the record afterward and prove it.

Because the pattern is spread across the whole passage, light editing doesn't remove it. Translation doesn't either, since the model chose every word of the translation. A full rewrite where every word is replaced does. And the watermark is naturally thin in places where there's only one right word to pick: dates, figures, code, anything heavily proofread. Less randomness, less room to hide a signal.

Who can check

This is the part that matters for you.

You can't. Not yet.

Anthropic's detection API is in private preview for organisations that need it under EU law. Regulators. Law enforcement. Fact-checkers. Media. Researchers. Educational institutions. That list isn't arbitrary. The EU AI Act's Article 50 transparency rules came into force on 2 August 2026, and this is the compliance mechanism, shipped globally.

The detector answers exactly one question: what is the likelihood this text was partly written by Claude? It can't confirm a human wrote something. It can't tell "Claude wrote this" from "Claude heavily edited this." And it says nothing about text from OpenAI, Google, or Microsoft. There is no universal "was this AI?" test, and the way this is being built, there won't be one.

Why this is a change problem, not a tech problem

Here's the shift, in one line: guessing is being replaced by checking, and the ability to check is being handed to institutions, not individuals.

Tells were democratic. Anyone with eyes could hunt for them. Detection is a credential. A regulator has it. Your comms team doesn't. Your manager doesn't. The skill of spotting AI text is being replaced by access to a verdict, and most people in most organisations are on the wrong side of that line.

Every technology goes through this. Nobody checks a banknote by squinting anymore; there's a pen. Nobody catches plagiarism by memory; there's a database. Guessing is what people do before the infrastructure exists, and it's always the phase where they feel most confident and are most often wrong.

That leaves change agents with a question their organisation is about to ask them: if we can't tell, and we're not allowed to check, what do we do?

What to do about it

Stop running "spot the AI" sessions. They now teach people to be confidently wrong in both directions. Someone who writes with em dashes gets accused. Someone who ran their draft through Sonnet 5 walks through clean.

Build a disclosure norm instead of a detection policy. The organisations that handle this well won't be the ones that catch AI text. They'll be the ones where nobody has to, because saying "I used Claude for the first draft" costs nothing. That's a norm, not a rule, and norms are co-created. If you write a policy and broadcast it, you'll get the exact compliance theatre you'd expect. If you ask teams what they'd want disclosed and why, you'll get something people actually follow.

Watch for the AI police. Some leader is going to hear "detection API" and want one pointed at employees. That's surveillance wearing a compliance badge, and the trust cost will outrun whatever it catches. The Lean Change position hasn't changed: we don't overcome resistance, we don't manage by suspicion, and we don't confuse a tool for a culture.

Treat this post as an experiment. It's watermarked. You can't check. Notice how that feels, and notice that the disclosure at the top did more work than the watermark ever will.

How to check this post

The check that works today: the disclosure at the top. A human told you. That's the whole mechanism, and it's the one your organisation can actually build.

The check that exists but you can't run: the watermark. Per Anthropic, everything Fable 5.1 writes carries one, including this sentence. If you're a regulator, fact-checker, or researcher with access to the detection preview, run it. Everyone else: this is what "access, not skill" feels like from the outside.

What you'd check if we shipped an image: Claude attaches a C2PA content credential to images it makes. Those are publicly verifiable. Drop the file into contentcredentials.org/verify and it tells you. Text doesn't get one. Yet.


FAQ

Does watermarking prove a human didn't write something? No. It can only estimate the likelihood that Claude was involved. Absence of a watermark proves nothing.

Can I remove a Claude watermark? A complete rewrite removes it. Light edits don't. Proofreading and changing a few sentences leaves most of it intact.

Do OpenAI, Google, or Microsoft watermark text? Each vendor's approach is separate. A Claude detector can't see other models' text, and there's no shared standard for text the way C2PA is for images.

What should a change team do first? Draft a one-line disclosure norm with the people it affects. Something like: "If AI wrote most of it, say so in the first line." Then run it for a month and see what breaks.


Lean Change Management is a feedback-driven approach to organizational change built on co-creation, experimentation, and meaningful dialogue. This post is part of the AI & Change series at leanchange.org/ai.

0Appreciate
Comments (0)
Share
Conversation

Start the conversation


No comments yet. Sign in to start the conversation.

Continue reading

Three related dispatches

The Dispatch

A weekly note, carefully made.

One letter, every Friday morning. New essays, recommended reading, and the occasional dispatch from a city we’ve recently been in.

4,200 readers · No spam, unsubscribe anytime