AI Watermarks Are Solving the Wrong Problem
When AI tells us that a machine was present, It tells us far less about the quality of the thought, the human contribution, or whether the tool was used well.
I’ve written before about why AI detection is unreliable. So the next obvious question is: what if we stop trying to infer AI use after the fact and have the generator leave a signal itself?
Anthropic has now offered exactly that. Instead of asking a detector to guess whether a piece of text came from AI, the generator leaves a signal in the text itself. Claude’s new models are being equipped with imperceptible, machine-readable watermarks designed to survive copying, pasting, and light editing, with third-party verification tools promised as well.
Fair enough. Let us give the idea every possible advantage.
Assume the signal is perfect. No false positives, no false negatives. Editing does not disturb it. Translation does not destroy it. Human and machine prose remain perfectly attributable. I could write another 8,000 words on why text watermarking is unlikely ever to deserve such confidence. If you do not believe me, ask Anthropic.
But that would spoil the experiment. Let us assume the thing works perfectly. Claude was used, and we know this for certain. Now what?
If you are a school, university, publisher, or employer, what exactly are you supposed to do with this information?
This is where the technical problem politely leaves the room and the methodological one is found sitting in its chair.
A watermark can have perfectly legitimate purposes: provenance. Publishers may want to know where material originated. Platforms may have disclosure obligations. The watermark may perform that narrow function admirably.
The trouble begins when an institution asks it to perform a judgment. The signal tells us that a tool was present. It tells us very little about intellectual contribution, competence, or misconduct. Those conclusions require the tiresome business of deciding what the person actually did.
We built a tool everyone should learn to use, then built another tool to tell us that someone used it. This is apparently progress.
Human beings are tool users. We extend ourselves with machines and eventually include mastery of those machines inside the definition of competence. Nobody admires an architect because she personally poured the concrete. A scientist does not become more rigorous by refusing statistical software. A filmmaker earns no artistic distinction by editing with scissors. Yet AI has somehow acquired a faint moral odor, as though intellectual virtue were measurable by the quantity of unnecessary labor one performs personally.
There is a deeper problem. AI is changing a relationship we have long been rather careless about: the relationship between the quality of an intellectual artifact and the quality of the thought behind it.
For a long time, artifact quality told us quite a lot. Vocabulary, grammar, structure, clarity, polish, and control of style were expensive. Producing them required skill. A coherent 2,000-word essay therefore carried useful evidence about the person who produced it.
Useful evidence is still only evidence, it suggests nothing but possibilities, even if the person did not get any help from friends, ghost writers, editors, etc.
A beautifully written essay has never guaranteed a beautiful thought. Shakespeare did not become Shakespeare because he had perfect grammar, and humanity has produced a heroic quantity of impeccably grammatical nonsense. We knew this in principle. In practice, polished language was difficult enough to manufacture that we began treating it as a convenient proxy for intellectual quality.
AI is in the process of breaking that proxy.
Language remains extraordinarily important. Thought has to travel somehow, and language is one of the principal channels through which one mind becomes legible to another. Vocabulary, syntax, cultural convention, education, confidence, and command of the receiving language all affect what survives the journey.
The problem is often not that a person cannot think something, or even that they cannot express it. They usually can. The problem is what arrives at the other end.
A person may already possess the argument, the distinction, the joke, the objection, and the precise degree of hesitation intended. Put all of that through a second language, unfamiliar academic conventions, or a style of prose never properly learned, and the receiving mind gets a lower-resolution version.
We have traditionally judged the source partly by the quality of the received signal. It was convenient. It was never entirely fair.
I have written elsewhere about the old Buddhist image of the finger pointing at the moon. Language is one such finger. An extraordinarily important one, certainly, but still a means by which one mind attempts to direct another toward something beyond the language itself.
Education has spent a great deal of time examining the finger because the moon is inconveniently difficult to grade.
AI can improve the channel.
Its promise is not that everyone can now sound intelligent. We industrialized sounding intelligent some time ago. The more interesting possibility is that more of the intelligence already present at the source can survive the journey.
Someone whose intellectual life happens largely in Chinese, Arabic, Spanish, Hindi, or another language no longer has to acquire decades of English-language polish before an English reader can encounter something closer to the full resolution of the thought. Someone who thinks clearly but writes awkwardly may be able to reduce the distortion between intention and reception.
That does not abolish language. It increases bandwidth.
From this perspective, AI looks less like an engine of intellectual fraud and rather more like an equalizer. Equalizers, naturally, are not welcomed equally.
A person who has spent decades mastering the old bottleneck may have complicated feelings about its sudden removal. The skill remains real. What changes is some of its scarcity value. Someone formed in one linguistic tradition may now find herself competing, on more equal expressive terms, with someone whose intellectual formation happened almost entirely somewhere else.
Neither has suddenly acquired the other’s education. What has changed is the cost of making one legible to the other.
AI does not make linguistic mastery worthless. It merely begins separating its intrinsic value from its gatekeeping value. Those are not quite the same thing, although beneficiaries of the latter have had little reason to insist on the distinction.
A bottleneck is merely an inconvenience until someone has built a career on being unusually good at passing through it.
AI may give millions of people a better finger. Naturally, one of our first institutional responses is to devise a way of identifying who used one.
There is another complication. AI, left largely to itself, is remarkably good at the competent average. Give it a familiar assignment and it will produce something fluent, organized, plausible, and faintly forgettable. It has excellent manners and very little reason to offend.
The interesting skill begins when someone moves the result away from that center.
Give two people the same model. One asks for an essay, accepts the first plausible answer, and submits twelve polished paragraphs of competent sludge. The other brings original notes, source material, preferences, peculiar objections, and enough accumulated irritation to be useful. She rejects mediocre output, changes direction, checks claims, rewrites sections, and keeps going until the result reflects choices the default model would never have made.
Both used AI. The watermark, having discharged its duties impeccably, reports the same fact about both.
Education should perhaps aspire to greater discrimination.
This is where taste enters. The more personal the material, the more tailored the instructions, and the less willing the user is to accept generic output, the more the final result can preserve the peculiarities of the source mind. Taste is the accumulation of decisions about what belongs, what is banal, what is alive, and what deserves to be removed despite being perfectly correct.
Anyone can ask for twenty introductions. The useful ability is knowing why nineteen are wrong, and occasionally why all twenty should be quietly disposed of.
There are perfectly good reasons to prohibit AI in particular exercises. If the purpose is unaided writing, ban AI. If the purpose is arithmetic, ban calculators. Learning to make fire without a lighter also has value when the lesson concerns making fire.
The confusion begins when a training constraint becomes a theory of virtue. There is considerably less educational value in standing beside a box of matches and congratulating oneself on purity. “I did it without AI” may sometimes be admirable discipline. At other times it is rubbing two sticks together while dinner gets cold.
Education will have to decide what it actually wants to measure. This is unfortunate, because judgment is considerably harder to administer than detection. A detector gives you a percentage with two decimal places, which is wonderfully soothing. Evaluating judgment requires another person with judgment. One can see the administrative difficulty.
That is the methodological error underneath the watermark debate. We measure what is easy to measure, then quietly promote it into what matters. Artifact quality once worked as a rough proxy because good artifacts were expensive to produce. AI is making them cheap while also allowing people who once paid a heavy linguistic toll to communicate with greater fidelity.
If communication is ultimately an exercise in alignment, a tool that reduces distortion between intention and reception has a fairly obvious value. The fact that our instinct is sometimes to mark its users suggests that the old bottleneck may have been doing rather more than helping people communicate.
The tollbooth disappeared, so we invented a watermark to identify who failed to pay the old toll. Anthropic’s watermark may aspire to tell us that Claude was present. Fair enough. Let it.
Education has the less convenient task of deciding whether anything worth saying was present as well.
That is a skill worth teaching, if anybody knows how.
**
P.S. This article was originally considerably longer. I was explaining too much. So naturally I asked ChatGPT to tighten it. My instruction was simple: stop proving everything. Point in the right direction so that my smart readers won't get bored.
You may now decide whether ChatGPT did a decent job.
Claude, incidentally, was nowhere near this article.
If Anthropic’s watermark declares AI involvement, it is technically lying while accidentally revealing a partial truth. If it declares no AI involvement... you are fooled.
That says a lot about the watermark.
Now, what if I asked my kid to do the same job? Will more 'human effort' make the article better? That says even more about the watermark.