Anthropic recently started stamping an invisible watermark into everything Claude writes. The watermark is a modest piece of engineering. It was released to satisfy the European Union’s AI Act, which now asks AI providers to mark machine-generated text. Within a day, people shipped free tools to scrub it off.
A few weeks before, a multimillion-dollar book deal fell apart because the author’s own agents could not prove he had written the novel himself.
Two different worlds, same reflex. Rather than judge the work, we reach for a way to trace the tool.
The trace is easier than the judgment. It is also the wrong thing to measure, and that is the mistake worth examining. Brace for a long post as there is quite a bit of nuance to wade through.
The puzzle
We hand large parts of coding to AI and call it good practice. Experienced engineers let models draft and test significant chunks of code, and spend their own time on design, architecture, and review. Nobody calls that fraud.
Hand a sentence to the same AI, and the verdict flips. Using AI to write is looked down on, quietly or loudly.
Same tool. Opposite judgment. Why does the tool that makes you a competent engineer make you a suspect writer?
The answer is verifiability
In the book I lean on one property to predict where AI improves fastest and performs most reliably. Verifiability. A task is verifiable when its success can be specified in advance and checked. Code sits at the high end. You can state what working means before you write a line, and the check runs on its own. It compiles or it does not. The tests pass or they fail. Clean, cheap feedback is exactly what these systems learn from, which is why coding improved faster than anything else.
It is also why we forgive the tool here. When you can check the result easily and at scale, you stop needing to know how it was produced. The proof is in the running system. Where you cannot check the result, you reach for the next best thing, a guess about who or what was involved.
For writing, everything depends on one clarification: verify what, exactly. Verify the goal of the writing, whether it did the job it set out to do. That is a separate question from whether the content is true in the world, and it returns when we get to accountability. And like any task, verifiability lives at the level of the specific goal. “Writing” as a whole has no single answer. That is why “writing” is a trap word. It covers at least three goals that sit in very different places.
Some writing is functional. Its goal is to carry an idea from one head to another. Manuals, release notes, briefings, most business prose. The form is disposable. The goal is specifiable. You can say in advance what it would mean for the idea to land, and you can check whether it did with a rubric and a test reader. You judge it much the way you judge code, and almost nobody gets upset about AI here.
Some writing is expressive. The writing is not just the form or mechanism but also the end goal in itself. Poems, stories, essays, a voice you came for. Its goal is the experience it evokes in the reader. A person can judge that, and good readers agree more than you would guess. But the standard will not reduce to a specification that runs without them. Every verdict needs a human in the loop. Readers come to expressive work for a human voice, and they feel cheated when the voice turns out to be a machine. Bad AI creative work offends the most, because it asks for an emotional response it never earned.
And a lot of writing lives in between. Thought leadership, newsletters, a company’s voice. It carries an idea and represents its author at the same time. Most writing that people actually argue about lives here.
The flood, and the trap it sets
The complaint I hear most is a fair one. People say they can see through AI writing now. There is an ocean of it. They are tired, they are busy, and they want a faster way to sort it than reading every word.
That sounds like verification working. It is pattern-matching on a fingerprint.
When AI prose was rare, the fingerprint and the badness came together. The tells in the style and the emptiness of the content were the same texture, so spotting the style was a decent proxy for judging the quality of content. Volume broke that link. Now the tells sit on top of real thinking and on top of filler alike. The surface no longer tells you which is which. So readers lean harder on the fingerprint in the style, and they throw out the good with the bad.
You can watch this happen in book publishing right now. A recent Wall Street Journal piece described literary agents so overwhelmed that clumsy, human writing has become a relief. One agent said the polished submissions flooding her inbox make “Fifty Shades of Grey” look like Tolstoy. Polish has become a signal of guilt. Competence in style reads as a machine. That is what happens when a fingerprint is the only tool you have.
The same piece put numbers on the deluge. One executive estimated that the overwhelming majority of AI books online exist to trick a buyer. One study, not yet peer-reviewed, found that around a fifth of the Amazon ebooks it sampled showed substantial AI help. The flood is real. The tools for sorting it are the problem.
What the watermark actually reads
The watermark answers one narrow question. Did a Claude model probably touch this text, given enough of it to measure. That is all. It does not know who had the idea. It does not know whether the writing is any good. It does not know whether another AI wrote the whole thing. It cannot even tell generation from light help. Run your own paragraph through Claude for a grammar pass, and it comes back marked.
And it is not alone. The industry is building a whole shelf of these instruments. Producer-side watermarks like Anthropic’s. Reader-side detectors like Pangram, which publishers are already using to vet manuscripts. Honor-system badges like the Authors Guild’s “human authored” certification, which rests on a signed attestation and, as the Journal notes, not much else.
Every one of them reads the same thing. Provenance. Which tool was in the room. None of them reads quality.
The clearest case in publishing is a dystopian romance that climbed the bestseller lists, got picked up by a major publisher, and then drew fire when a detector flagged it. Readers liked it. The market judged it good. And the provenance suspicion overrode that verdict anyway. The question stopped being “is this any good” and became “was a machine involved,” as if the second answered the first.
I will grant one place where provenance genuinely matters. Ownership. AI-generated text cannot be copyrighted, so knowing what a machine produced has real legal weight. That is a fact about property. It says nothing about quality. Keep the two apart and most of the confusion clears.
The tools can miss the wrong people
Detection does not just answer the wrong question. If it answers it badly, it lands hardest on the wrong people.
The detectors produce false positives. Authors deny the charge and have no way to prove a negative. Agents and editors, who signed up to find good books, now find themselves acting as police. The person using AI to clean up grammar in a language they learned as an adult gets flagged the same way as a spam farm. The motivated bad actor, meanwhile, runs the text through another model and walks away clean.
A signal that catches the honest and misses the deceptive protects no one. It is theater.
Slop has two axes, and provenance is neither
So if AI use is not what makes something slop, what does?
In the book I define slop as cheap creation meeting vague intent. It has two moving parts.
The first axis is intent. Is it clear what this is for, and for whom. That “for whom” matters. Frictionless prose that dumps a wall of text on a busy reader is an intent failure too. The writer never decided to respect the reader’s time.
The second axis is execution. Is it done well. Clear, economical, well made, serving the job it set out to do.
Slop is a failure on either axis. Four corners.
Clear intent, good execution, is craft. That is the only corner that is not slop.
Vague intent, good execution, is polished slop. It reads beautifully and serves no purpose, or ignores the reader it was aimed at. This is the dangerous corner, because the quality in style hides the emptiness.
Clear intent, poor execution, is a real point, botched. Still slop.
Vague intent, poor execution, is the pure kind nobody argues about.
The book’s definition looks narrower than this, because it is the same picture under one assumption. The book is about the AI era, and especially about agents, where execution is assumed to have cleared the reliability bar. Assume execution is handled, and that axis drops out. The four corners flatten onto the intent line: clear intent gives craft, vague intent gives slop. The only way left to make slop is to fail on intent. That is the case the book describes.
There is a reason intent is suddenly the axis that matters. Doing the work used to be expensive, and the expense screened out a lot of weak output before anyone saw it. It took effort to write and that effort alone screened out a lot of potentially bad writing. AI removed that screen. The cost of execution no longer filters anything. So intent is the only axis left doing real work, and it is the one no tool touches.
AI mostly lifts execution and leaves intent alone. So it multiplies polished slop, while the hard axis, intent, stays exactly as hard as it always was.
And the watermark? It reads a third axis entirely. Provenance. It runs at a right angle to both of the axes that actually define slop. A human can produce pure slop with no machine anywhere near it. An author with a clear point, using AI to execute well, produces craft. The detector cannot tell them apart, because it is not looking at either thing that matters.
Why writing takes the moral heat
The practical wariness about AI writing has a simple source. We cannot cheaply check the result, so we are left guessing. The moral charge, the sense of betrayal, is a separate thing, and it comes from what effort is a proxy for.
Expressive and in-between writing carry a costly signal. The time you spend is a proxy for how much you care, the same way a thoughtful introduction carries weight because the person made the effort to vouch for you. Spend that time and the reader feels respected. Let a machine spend it in a second, and hide that you did, and it can feel like deception. That is why the reaction to AI writing runs hotter than the reaction to AI code. Code was never carrying that signal.
Where I stand
I use AI in my writing. I use it here. I use it to sharpen sentences, to test an argument against its weakest point, to find the shorter way to say a thing. I am not shy about it, and I am not going to pretend otherwise.
I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. Wanting to be more productive is a good enough reason on its own. It needs no apology.
I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. I don’t consider myself a non-native speaker working in a second language he is not fluent in, and AI certainly is a real gift that scenario. AI also helps the fluent expert who simply wants leverage. Wanting to be more productive is a good enough reason on its own. It needs no apology.
The division of labor I keep is the same one every engineer keeps with code. I own the thinking and the judgment. AI helps with the production. And I check the result before it goes out. That last part is why what comes out is craft and not polished slop. I bring the intent. The tool lifts the execution. I verify the execution. The provenance is beside the point.
Even the publishing world, at its most protective, half-concedes this. One independent publisher, guarding the most human corner of writing there is, allowed that a genuinely meaningful work could earn a place on her list as long as it carried a clear note about how it was made. If that door opens even a crack for fiction, where human presence is the whole point, then for functional writing, refusing a useful tool is superstition dressed up as principle.
There is one real cost, and I will not wave it off. Writing is how you learn to think, and leaning on the tool can dull both the craft and the thinking behind it, not on any single piece, but in the writer, over time. That is the individual version of a problem I spend the book on, the slow erosion of a capability you stop exercising. It is a cost worth watching. It is also a different question from whether a given piece is slop, and it is answered the same way any skill is kept, by still doing the hard parts yourself. Which is exactly why I keep the thinking and the judgment, and use the tool for the production.
Own every word
There is one condition that makes all of this responsible, and it has nothing to do with which tool you used.
There is one condition that makes AI use responsible. Tool or no tool, the author is accountable for what goes out under their name, including any consequences.
You own every word. Tool or no tool, the author is accountable for what goes out under their name, including the consequences of anything they failed to check. A writer covered by an imprint recently shipped a book with quotes the AI had hallucinated. He had disclosed that he used the tool. He had not checked what it produced.
This is where the truth of the content comes back. The writing did its job. The quotes read cleanly and carried their point, so the goal was met and the prose was well made. But the quotes were still false. Whether a piece achieves its goal and whether its claims are true are two different verdicts, and the author owns both. The tool can lift the writing. It cannot carry the accountability.
This is also the distinction the whole detection industry misses. Provenance asks who produced the words. Accountability asks who answers for them. A detector chases the first. Everything that matters in business runs on the second. And the “I just used a tool” defense collapses. You cannot copyright what the AI wrote, so you may own less of the words than you think, while owning all of the liability for them. Less of the property. All of the responsibility. That is the deal, and it is the right one.
What this means past writing
This is not really about novels.
Almost all business writing is functional or somewhere in the middle. And leaders, faced with the flood, will reach for the same reflex publishing reached for. Ban the tool. Scan for the watermark. Make people prove they wrote it. It will not work. The marks come off, the detectors misfire, and the removers are already on GitHub.
Watch publishing to see the future of that approach. Literary agents turned into police. Certifications that rest on a promise. An entire trust-based industry being stress-tested by a volume it cannot inspect by hand. Detection and attestation are both attempts to replace trust, and neither scales against the flood.
What an organization actually needs is a chain of accountability. Someone has to answer for whether the work is good, and no trace of which tool touched it can supply that. The reason is the same one that runs through this whole piece. The standard for good work, in most writing and most judgment, cannot be fully specified in advance, so no detector and no rule can render the verdict for you. Detecting the tool is measurement. Judging whether the output serves its purpose and holds together is a harder thing, and a human one. A better detector will not get you there. What does is an architecture that keeps a person answerable for the result. That is what coherence means.
The last word
The compiler is why we forgive AI in code. It is a standard specified so completely that it settles whether the code works, cheaply, every time, and once that is settled we stop asking who wrote it. Most writing has no compiler. So deciding whether the work is any good stays with a person.
The watermark can tell you a tool was in the room. It cannot tell you whether anyone was thinking. That was never the tool’s job. It is yours.
Use the tool. Say so if you like. Stand behind every word. And let the work answer for itself.
My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.