This content is not intended for human consumption. Here is why.

Who this is for

This is for people who draft with a language model and then put their name on the result: a post, a design doc, a cover letter, an update to the team. At some point a reader says “this reads like AI”, or a colleague says it sounds like ChatGPT, or nobody says anything and the reply rate drops. The usual fixes are word-level. You strip out the famous tells, “delve” and the em dashes and the “I hope this finds you well”, you run it through a rewrite, and it still reads wrong, and when someone asks what exactly is wrong with it, neither you nor the reader can point at a word. Most of what they’re reacting to sits in how the text is built rather than in the words, and it was built for a room it was never going to be read in. This piece is a root cause analysis (RCA) of that reaction, the kind operations writes after an outage: start from the symptom, assume nothing exists in a vacuum, and follow it back to where it started.

Users of Claude gave the reaction a name in August 2026. A thread on r/ClaudeAI about a plugin that translates Claudish to English reached about 1,500 upvotes and 170 comments in its first half day, and the thread’s auto-generated summary put it plainly: “Claude’s current writing style is atrocious and everyone is sick of it.” Other models have their own version (people call GPT’s “GPTese”), and everything measured below comes from Claude-family drafts, so read the findings as Claude findings until someone runs them on another model.

Claudish here means the name users gave Claude’s writing style. There is also a command-line tool called claudish, which is unrelated.

Two complaints wearing one coat

The most useful comment in that thread came from a user called pingwing, who wrote that “the complaint is two complaints wearing one coat.” The first complaint is sycophancy: the agreeing, the praising, the folding under pushback. pingwing’s diagnosis of that one was “it isn’t a style problem, it’s a training problem”, and I agree. Models are tuned on human ratings, people rate agreement higher, and the agreement gets trained in. I wrote that half up earlier as a paper on layered sycophancy, and its one-line version is this: sycophancy is layered symptom substitution: suppress the reflex at one layer and it resurfaces in the next.

This piece is about the second complaint, the style itself. You can strip every flattering word out of a reply and it still reads as machine-made, and the reason sits somewhere else entirely.

The devices work, just somewhere else

Take the triple beat: three short sentences in a row, each one landing. On a stage it works. The speaker controls the pace, pauses after each beat, and the room can’t rewind, so the repetition gives the listener three chances to catch the point. Put the same three sentences on a page and the reader sees all three at once, reads them at their own speed, and notices that the third one is there for rhythm.

The device isn’t the problem. It belongs to one mode and turned up in another. I call that a mode leak: a speech device showing up in text that will be read silently.

Two questions decide whether a device works in a given piece:

  1. Can the audience go back? A listener can’t, so speech needs redundancy, rhythm and signposts. A reader can, so writing needs compression and a structure you can see on the page.
  2. Who controls the pace? In speech the speaker does, and a pause carries meaning. In writing the reader does, and a pause on the page is dead weight.

Chat sits in between. You can scroll back, and the other side answers every turn, so a fragment or a quick “right?” works in chat and a recap doesn’t.

Register has two axes

People tend to treat register as one scale, formal to casual. It works better as two. One axis is the mode: live speech, recorded speech, writing, with chat as a hybrid. The other is the purpose: persuading or performing, teaching, informing or documenting, and bonding. A talk is live persuasion, a podcast is recorded speech, a tutorial is teaching in writing, an opinion piece is persuasion in writing, and a runbook is documentation.

The two axes explain a lot of bad advice. “Write like you talk” is right for chat and wrong for a runbook. “Hook the reader in the first line” comes from live persuasion, where you really do have seconds before the room drifts, and carried into a technical article it produces the “Here’s the thing” opener that readers now spot at once.

Here are the devices that leak most often from speech into writing, with the reason each one fails on the page:

DeviceWhy it fails in silent reading
Hook opener (“Here’s the thing”, “Picture this”)The reader chose to open the page and doesn’t need to be caught.
Triple beat, three-item lists by defaultThe reader can count the items and sees when rhythm picked the number.
Rhetorical question with a two-word answer (“The result? Chaos.”)On a page the question is a pause the reader didn’t need.
One-line punch paragraph after a long oneIt’s a stage pause set in type. Used once at the turn of an argument it works, and a string of them reads as a slide deck.
Recap closer (“In short”, “The takeaway”)The reader can scroll up, and the recap repeats what they just read.
Announcing the count (“Three things matter here”)The headers already show it.

The linguistics has been around for decades. Walter Ong’s Orality and Literacy (1982) lists the features of speech-based expression, and the first one is “additive rather than subordinative”: spoken language strings clauses together with and, and then, and so, while writing stacks them inside each other. Douglas Biber’s Variation across Speech and Writing (1988) measured hundreds of texts and found that the strongest single dimension separating them was “involved versus informational production”, which tracks the difference between talk and print closely. None of that was ever a secret. What changed is that a machine now produces text in bulk without a way to tell which mode it is writing for.

The root cause: why a model mixes them

Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal or the harness. For Claudish the first of those does most of the work. A language model learns from text without labels saying which register each piece came from. A sermon, a TED talk transcript, a YouTube caption file, a sales page, a textbook and a reference manual all arrive as the same kind of token stream. And the persuasive text that gets published and transcribed the most is spoken: talks, sermons, speeches, videos. So when a model is asked to make a case, the most common examples of making a case it has seen were built for a live room.

The model’s output, on the other hand, is almost always read in silence. It goes into a chat window, a document or a pull request, and someone reads it at their own pace with the option to go back. That mismatch is the mechanism: persuasive intent pulls toward the live register, the medium is writing, and the result is stage devices in print. Readers have now seen enough of it to recognise the mix as a signature.

Tracing it through one draft

The best evidence I have is a single article that went through four versions in one day, with the facts, the citations and five fixed sentences held the same in every version, so only the register moved. The article is my reading of a chess experiment by an evaluation lab, and it was drafted with Claude.

  • A, keynote. The first draft. Full of triple beats (“A lab trained against one route. The optimisation pressure stayed. The behaviour took the next route.”), fragments after bold lines, and runs of sentences that each start with a verb.
  • B, run-ons. A rewrite asked to sound like me: long sentences joined with so, and, but and which, direct “you”, casual words. The staccato was gone. It read wrong anyway.
  • C, split. A pass that split thirteen of the long sentences to vary the length.
  • D, additive. A pass that kept the length and changed how the long sentences were built.

I compared those with three texts I wrote myself: two LinkedIn posts and a Reddit comment.

TextSentencesMean wordsSpread40+ wordsLong (25+)Additive ratio
Mine, LinkedIn (engine)2718.30.7611%93.0
Mine, LinkedIn (operations)3214.30.683%41.5
Mine, Reddit comment926.70.5522%56.0
Draft A, keynote9113.90.805%171.4
Draft B, run-ons5725.20.7628%261.2
Draft C, split6923.20.7422%301.3
Draft D, additive6923.10.7422%302.9

Spread is the standard deviation of sentence length divided by the mean. The additive ratio counts and, but, so, or and then against which, that, where, because, while, when and if, in sentences of 25 words or more.

Spread doesn’t separate anything. It sits between 0.55 and 0.80 for every text, mine and the machine’s alike. That matters because sentence-length variation is the first thing people reach for when they want to measure “sounds like AI”. It’s a common stand-in for burstiness, one of the two statistical signals the detector GPTZero reports, although GPTZero itself defines burstiness as the variation of perplexity across a document and not as sentence length.

Length doesn’t separate either. Draft B has long sentences, and so does my Reddit comment, where more than a fifth of the sentences run past 40 words. Long, breathless run-ons are one of my habits, so a rule that cuts long sentences would cut my own voice.

How the long sentences are built does. Mine add one clause after another, with and, but, so, commas and a self-interruption partway through. The drafts nested them, with which, where and because. That is Ong’s additive versus subordinative, measured. Drafts A, B and C sit at 1.2 to 1.4. Draft D, where the same long sentences were rebuilt to add clauses instead of nesting them, sits at 2.9, inside the range of my own text. This is a pilot on three samples of mine, and one of them sits at 1.5, close to the drafts, so treat it as a direction to test and not as a threshold.

The failure in draft B is the one worth remembering. A filter that only bans devices pushes the text into the next failure, and the obvious measurements can’t see it.

A voice is not a leak

My own writing brings speech into print all the time, and people don’t read it as machine text. The difference seems to come in two parts.

The first part is that I put the delivery on the page with punctuation. An ellipsis carries timing (“Water will burn your house down… if you throw some into a kitchen oil fire.”), capitals or italics carry stress (“YOU can ALSO take part”), and brackets carry the spoken aside. A stock device skips that step: it puts the stage rhythm in type and leaves out the pause that made it work.

The second part is that the material is specific. A made-up word announced in the same sentence, a reference to a children’s TV show, a “:D”, a real typo. None of that can come out of a template.

The hypothesis, from three samples and labelled as such, is that stock speech devices read as AI, and speech devices carrying specific material read as a person.

It sets a limit on what a model should do when it writes in someone’s voice. It can copy the shape: the additive long sentences, the direct “you”, the concrete opening, one casual word, one punch line where the argument turns. It should never invent the specific material, because made-up typos, coinages and jokes are the next tell, and they are the author’s to add.

What to do

Before drafting, name the mode and the purpose. A Reddit post is persuasive writing close to chat, a LinkedIn post is persuasive writing, a tutorial is teaching in writing, a reference page is documentation, a video script is speech. The name decides which devices are allowed.

After drafting, go through the text once for devices. Remove the ones that fail in that mode, and keep a triple beat only when there really are three things. Allow one fragment or one punch line, at the turn.

Then go through it once for construction. Look at the long sentences and check whether they add or nest. If the piece is meant to sound spoken, rebuild the which and where chains into and, but and so. If it is documentation, nesting is fine.

In the same pass, check who is doing the verbs. Claudish is full of sentences where the subject is a thing, “the data shows”, “this enables”, “the approach ensures”, and no person does anything. Put the person back as the subject. This check comes from a reply to the r/ClaudeAI post of this piece, which also made a point about order that matches the four drafts above: fix the structure first and the phrases last, because if you fix the words first the tics disappear and the text still reads like a talk.

Some of this can be checked by machine. Hook openers, recap closers and a short question followed by a very short answer are regular enough for a pattern match, and I run those as a hook on every reply. Triple beats, punch paragraphs and clause construction still need a reader, human or model, who knows which mode the text is for.

Where the check lives matters as much as what it checks. A language model has Markovian memory: at run time the next token depends only on what is in the context window, so anything outside the window does not exist for it. A style rule you gave it three sessions ago, or one that a compaction summarised away, is outside the window, and the register drifts back to whatever the training text made most likely. A check that runs on every reply puts the rule back in the window each time, and that is why the mechanical part runs as a hook and not as a request at the top of the chat.

What would refute this

  1. Readers who, shown a device-heavy and a cleaned draft of the same piece written for silent reading, pick out the device-heavy one no better than chance.
  2. The additive ratio failing to separate a writer’s own long sentences from model drafts written in that writer’s voice, across more writers and more samples than my three.
  3. A model trained or prompted with register labels that still puts live-persuasion devices into documentation at the same rate as one without them.

Terms used here

  • Claudish: the name users gave Claude’s writing style; here, the style half of the complaint, separate from sycophancy.
  • Register: the shape of language chosen for a situation, described here on two axes, mode and purpose.
  • Mode leak: a speech device showing up in text that will be read silently.
  • Additive versus subordinative: clauses joined one after another (and, but, so) versus clauses nested inside each other (which, where, because); Ong’s first mark of speech-based expression.
  • Additive ratio: additive connectives divided by subordinating connectives, in sentences of 25 words or more.
  • Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal or the harness.
  • Markovian memory: at run time the next token depends only on what is in the context window, so anything outside the window does not exist for the model; the weights are fixed rules, the window is the only state.
  • Voice: a writer’s habits laid over a register; the shape can be copied, and the specific material belongs to the writer.

Slips caught while drafting

This piece was drafted with a language model. These are the register and stance slips caught before it went out, and who caught them. The running log across all pieces is at machinebehavior.io/slips.

  • announced count. Before: “Three things came out of it.” After: “(sentence removed; bold lead-ins carry the count)” Caught by: Claude mode-leak pass.
  • hook opener. Before: “So here’s the hypothesis, from three samples, labelled as such:” After: “The hypothesis, from three samples and labelled as such, is that…” Caught by: Claude mode-leak pass.
  • staged contrast (not X … it’s Y). Before: “…isn’t in the words at all. It’s in how the text is built…” After: “…sits in how the text is built rather than in the words…” Caught by: Claude comm-rules check.
  • overclaim: wrong measure named. Before: “the overcorrection shows as low burstiness (sentence lengths in one narrow band)” After: “spread sits near 0.7 for every text; clause construction separates them” Caught by: Language session (measurement).
  • overclaim: in a cut. Before: “Once those chains got rebuilt the draft stopped reading like Claude” After: “…the draft measured in the same range as my own writing” Caught by: Claude review of the cut.

Changelog and errata

  • v1, 2026-09-29: published.
  • v1.1, 2026-09-29: added the agent check (who is doing the verbs) to “What to do”, credited to a reply on r/ClaudeAI, with its point about fixing structure before phrases. No existing claim changed.
  • v1.2, 2026-09-29: added the “Slips caught while drafting” section. No existing claim changed.
  • Errata: none so far. Corrections will be listed here with their date, and the corrected text marked in place.

The series

This is part of a series by Stefan Coetzee, 2026, on running language models as working systems.

Research programme and claims ledger: machinebehavior.io.