The Mark Is Not the Verdict
A lot of people have been angry about the wrong thing this week.
On August 11, Anthropic said it would begin watermarking text produced by Claude. On August 14, it published the details. Somewhere between those two dates, a story took hold: that the company had begun tagging people’s writing, that machine-generated tags could be traced back to them by someone, anyone, that work done in private was now stamped and filed. Dozens of people said publicly that they were canceling their subscriptions. In the threads I read, the word that kept appearing was surveillance.
That story isn’t accurate, and it’s worth saying so clearly: the real conversation is more interesting than the one that’s been happening.
The watermark carries no identifying information. Anthropic’s own account describes the mechanism plainly: nothing in the mark, or in the key used to read it, is built to let anyone recover information about a user, their organization, or their conversations. It is not a serial number. It doesn’t know your name. It encodes only one thing: that a Claude model chose some of these words.
It’s also not what most people picture when they hear “watermark.” There are no hidden characters, no invisible ink, no metadata to strip from the text. Generated files are a different case: Claude attaches standard C2PA provenance metadata to those, the kind an image tool can already read. The mark lives in the word choices themselves: the model’s selection among plausible next words is nudged by a key, and the pattern of those nudges is what a detector reads. That is why it survives a copy-and-paste, why a light edit probably won’t clear it, and why a complete rewrite will.
It’s going to be okay, really.
Although this next part deserves more of your attention.
A watermark reports one fact: a machine was involved in producing this text.
That is a narrow claim, and Anthropic makes it narrowly. The company says the mark “cannot distinguish ‘Claude wrote this’ from ‘Claude heavily edited this.’” A passage a human wrote and then handed to Claude for a heavy line edit, the kind that touches structure and phrasing throughout, can come back carrying a mark indistinguishable from something Claude drafted outright.
Clients, editors, admissions committees, and hiring managers will read a green light or a red one. The detector will say mark present, and the institution will hear, “A machine wrote this.” No amount of careful documentation on a vendor’s website closes that distance. Caveats don’t survive contact with a workflow. Verdicts do.
We’ve watched this play out before. A plagiarism score became an accusation. A credit score became a character reference. A pretrial risk-assessment score, its error rates published openly by the researchers who built it, became a factor in whether someone sat in jail before trial. In every case, a probabilistic signal, honestly described by the people who built it, hardened on contact with an institution that needed a decision faster than it needed the truth. The signal didn’t change. The reading did.
Anthropic hasn’t published a false-positive rate for this deployment. The underlying method, SynthID-Text, was developed at Google DeepMind and studied elsewhere; Anthropic’s own implementation hasn’t been, and there’s no public detector yet for anyone outside the company to test. Anthropic says an API is coming. When it arrives, there will be people entitled to ask whether your work is machine-made, and there is at present no published rule about who they are, what they may do with the answer, or whether you’ll ever know they asked. That’s not a scandal. It’s an unfinished building, and people are already moving in.
This is what bothers me, and it has nothing to do with privacy.
We’ve just built infrastructure to detect machine involvement in writing. We’ve built nothing to detect human judgment.
Machine participation is now the fact that requires special accounting: measured, marked, flagged, reportable. The framing of the question, the rejection of the first four answers, the decision about what to leave out, the willingness to put a name on it and take the consequences — none of that leaves a trace any detector can find. It’s invisible to the instrument. And what is invisible to the instrument tends, over time, to become invisible to the institution.
So we are constructing a world with a legible measurement of the least important variable and no measurement at all of the most important one. Legible is not the same as accurate. Nobody has published what this measurement gets right or wrong, and an unvalidated signal that institutions can act on is worse than an accurate one, because it carries authority without warrant. Then we’ll make decisions with it, because it’s the number we have.
That is the failure mode. Not the mark. The reading laid over it.
Authorship is not located in the tokens.
Suppose a machine helped me draft a sentence; that tells you something about my process and nothing about my argument. What makes a piece mine is not the absence of assistance. It’s the presence of someone who decided what question was worth asking, who refused the answer that arrived too easily, who chose what to stand behind, and who remains answerable for it afterward. The mistake runs in both directions: a writer can just as easily read the mark’s absence as proof of their own effort, or its presence as evidence they didn’t really do the work. Neither reading holds up. The mark doesn’t know what you contributed. Only you do.
Answerability is the load-bearing word. It’s the thing a mark cannot carry, because it’s not a property of text at all. It’s a relationship between a person and a claim, sustained over time, in public, at cost. You can hash a document. You cannot hash the willingness to be wrong in front of people.
That’s why I don’t object to provenance. In a content channel filling with machine-made language at a rate no one can moderate, an honest signal that a machine was involved is genuinely useful. I want it. What I object to is the promotion of that signal from a fact about production into a verdict about authorship, a promotion nobody has authorized, that the vendor explicitly disclaims, and that will happen anyway because institutions need something to point at.
(I make a version of this same argument, at more length, in an essay called The Last Frontier.)
The instinct in the replies has been to look for the exit: cancel, switch, find the tool that doesn’t mark. The EU obligation that produced this isn’t going away. Anthropic is one of roughly 190 companies that signed the July 2026 Code of Practice. The other labs are not exempt, and the honest ones will get there first. The exit closes.
A better move is available today.
Publish your account of your own process before someone else’s instrument supplies one.
Write the page. What tools you use. What they do. What they do not decide. Where your judgment enters and where it stops. Put it under your name, at a stable address, and keep it current. Then, when a detector reports a mark on something you wrote, you aren’t scrambling to explain yourself against a number. You’re pointing at an account you made in the open, before you needed it.
That remedy works better if you already have a name and an audience to point people toward. A student in front of an academic integrity board, a freelancer whose submission got flagged by a client’s detector, an applicant whose cover letter tripped a screen — none of them have a stable address anyone already trusts. The account only helps if the person judging you is willing to look for it, and the people most exposed to a mark being read as a verdict are the ones least likely to have that leverage. That gap doesn’t close today. It’s worth saying so rather than pretending the remedy scales evenly.
That’s not a defense. It’s a better offer than the detector’s.
The detector can tell someone that a machine touched your prose. Only you can tell them who framed the problem, who threw out the easy version, and who will answer for the finished piece.
One of those is a measurement. The other is a person standing behind the work.
Here is a brief version of that account. This essay went through several full passes with an AI editor that pushed back on claims, flagged a misattributed quotation, and named an argument I hadn’t answered well enough. I made every one of those calls myself, including the ones I’m still weighing. That’s the whole practice. It doesn’t need to be more complicated than that to be true.
The mark is not the verdict.
On process and disclosure: How I make this →