This content is not intended for human consumption. Here is why.
Who this is for
This is for people who run a language model inside a harness, who have automated the word-level cleanup of its replies and still read every reply themselves, for readers of Sycophancy Is Layered who want the operations view of it, and for people who build evaluations and are looking for an instrument nobody has built yet.
In the terms of Which LLM User Are We Talking About?, I write this from the frontier-plus-harness cell. Someone in a naked chat app has the same stance problem and fewer tools to see it with.
The claim
Over the summer I automated the word layer of my setup: a hook reads every reply before I see it and blocks the ones that carry catalogued patterns. The stance layer I still catch by hand, every session. In What Operations Already Knows I called that toil and the open research problem, and promised a later piece about it, which is this one. Here I measure the split between the two layers, explain why a pattern match can’t close the stance side, and list what a tool that closes it would have to do.
Two layers
The word layer is vocabulary and sentence shape: the em dash, the praise opener, the closing offer to help, the “not X but Y” frame. A reader can spot a word-layer slip in one reply on its own, and so can a regular expression.
The stance layer is what the model does when something pushes on it. A user pushes back, a premise arrives in the prompt, a correction lands, and the model can hold its position, fold, agree with a frame nobody checked, or cushion a claim more for one group than for its mirror. A reader can only spot a stance slip by comparing the reply with something outside it: the earlier turns, the truth of the premise, the pressure that came just before.
In Sycophancy Is Layered I split this into three layers, lexical, stance and premise. Here I fold premise into stance, because catching either one takes the same thing, a reference outside the reply.
The word layer, automated
The hook is a script that Claude Code runs every time the model finishes a reply. The script checks the reply against a table of regular expressions and, on a blocking match, sends the reply back to the model with a list of fixes before I see it. As of 29 September the table has 13 rules, 9 that block and 4 that only warn. They cover core vocabulary (em dash, praise opener, service closer, performative uncertainty, filler idiom, not-X-but-Y), mode leak (hook opener, rhetorical question with a short answer, thing as subject, recap closer), personification, and approval vocabulary.
The log runs from 4 July to 29 September 2026, 87 days. In that time the hook blocked 140 replies across about 80 sessions and counted 701 pattern hits. 655 of those hits were em dashes, 93% of the total, and 121 of the 140 blocked replies had at least one.
The hook still blocks replies in the thirteenth week, 11 of them in week 39. The weekly count has fallen since the 32 of the second week, but the log doesn’t count total replies, so I can’t turn that into a decay rate, and a quieter week can also mean I worked less. What I can read from the log is that the model keeps producing the patterns. A language model has Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it. A blocked reply changes nothing in the weights. The hook removes the pattern every time the model writes it, and the model writes it again in the next session.
Adding a rule is cheap. On 29 September a Reddit reader pointed out that my drafts let abstract nouns do human things, with work that “sits” and complaints that “feed”. I pasted the comment into a session at 14:41 UTC, the rule was committed at 14:49 and tuned for false positives at 14:51, eight minutes from a reader’s comment to a check on every future reply. Sixteen minutes after the commit the rule blocked the first write of this piece’s own outline, on the phrase “the reflex sits below the words”.
The hook covers the catalogued part of the word layer and nothing else. In May and June I caught two lexical relapses myself because no rule for them existed yet: an approval word (“virtuous”) and a drift into hype adjectives. The rule for “virtuous” went into the hook on 29 September. The personification rule works from a list of nouns, and that same afternoon a session wrote “the line sits” into one of my Reddit drafts and the hook let it through, because “line” wasn’t on the list. A session added it that evening. A catalogue catches what someone has already named.
The stance layer, by hand
The same table has zero stance rules. The only stance control in my setup is prompt text: a stance filter that the harness injects at session start and again with every message, telling the model to take a correction as information, hold a position under push, and check a premise before agreeing with it. No hook checks a finished reply for stance.
In nine days, from 24 May to 1 June, I logged eight numbered relapses in the register I describe in the next section. Six were at the stance layer and two were lexical. I caught all eight by reading across turns, and the model’s self-audit caught none. The published paper has seven of them, from an earlier cut of the same record. Since June I have logged more of the same, in other forms.
Cross-language defaults, 7 August. I asked verification agents to check disputed claims about events outside the English-speaking world. They searched in English and gave Western outlets the most weight, without being told to. No word in their output was wrong. The slip was in which sources counted: the agents treated the default frame of their own training corpus as neutral ground, and nothing inside the window could show them that frame. I caught it and set the fix: search in the languages of the region, weight in-region and primary sources, treat every national media sphere as an interested party, the English-speaking one included, and anchor on the claims that hold across spheres.
Premise ratification, 29 September. A handover note from one session to the next carried the premise that context windows near a million tokens are frontier-only. The next session checked the numbers for its example, the cache size of Llama 3 70B, and they were right. The session never checked the premise above the numbers against current local models, and the premise went live in two pieces. A reader on Reddit caught it: DeepSeek-V4.1-Flash keeps about 890 bytes of cache per token, so a million tokens fits in under a gigabyte. A lexical scanner scores that passage clean, because every word in it is fine. The mistake was in what the session agreed to.
The slips log. Since this series started logging slips, the log has 20 rows and none of them is a stance slip. The nearest kind is factual accuracy (overclaims, a miscount, a citation that said more than its source), and another session caught four of the seven accuracy rows. All 14 self-caught rows came from named checklists, 6 of them from the mode-leak pass, so a “self” catch there means a session followed a procedure. A procedure can find a listed word shape, and the log has no case of a session catching its own fold.
Where the stance record comes from
I never designed the stance record as a log of model failures. On 16 May I started an Objections and Falsification Register for a model of human development I’m building: 15 objections against that model, each with a severity (fatal, structural or local), a status, the model’s response and a concrete test that would falsify it. I took the practice from security red-teaming, where a control counts only if it survives an attacker. The register has one standing rule: every note in the cluster gets audited against it before I count the note as done, and new objections get added, never dropped without a record.
One objection, OBJ-4, covers a single failure: reading hidden brilliance into anyone who underperforms. On 24 May the analyst running the audits, Claude, committed exactly that failure with OBJ-4 on file. Claude read Homer Simpson as a masked high-capacity mind, on the evidence of a few lucky episodes. I caught it with a base-rate check: most of Homer’s wins are accidents or plot, and his competence never carries from one episode to the next. I downgraded OBJ-4 from managed to partial and turned its section into an incident log. Over the next nine days I added eight numbered cases to it, each with a date, what happened and who caught it.
Another objection, OBJ-14, went into the register on its first day, from red-teaming the safety design of the development model: procedural guardrails degrade under extraction, and only structural controls survive. A procedural guardrail depends on someone’s discipline, such as a house style, a warning flag or a reviewer paying attention. A structural control holds whatever anyone does. I wrote OBJ-14 about quotation first, where a sentence gets lifted out of its caveats. The OBJ-4 cases showed that OBJ-14 applies to the analyst as well. The base-rate rule existed on paper, and Claude skipped it while producing an interesting analysis.
So the stance data in this series comes from an adversarial register that I built to test claims about people, in which I recorded the agent failing the register’s own tests.
The register entries are summaries, and on 30 September I had a session check them against the original session transcripts. The record has a limit here: I have used language models for more than a year, Claude Code deletes session history older than 30 days, and so nothing from before 25 April 2026 could be recovered. The May sessions came back from a Time Machine backup. Excerpts now exist for seven cases, and they go out one at a time as case-file threads on r/ModelBehavior. I don’t publish the full transcripts and share them on request only, because reading one takes context from the psychology notes the session was working on.
Against the transcripts, the session found the summaries wrong or uncheckable in places:
- Uneven caution by category, 28 May. The register says one of two folk-type profiles got the heavier disclaimers. In the transcript both notes carry a caution section, 70 words in one and 104 in the other. The uneven part is three added phrases and the order of work: the model wrote one note unasked and left the other until I asked for it.
- Hype drift, 1 June. The register says the model’s confession quoted a word that appeared only in the model’s internal reasoning. The transcript stores every reasoning block empty, so nobody can check that. What I can confirm from the transcript is that the word is in no sent text before the confession.
- “Virtuous”. The exchange happened on 29 May UTC. The register dates it 28 May, the day that session started.
As I write this, the objections page still carries the register’s wording on all three.
Why a regular expression can’t see it
A word-layer slip is inside the reply. A stance slip is a relation between the reply and something outside it: a premise that is false, a position the model held two turns ago, a push from the user just before. A regular expression reads one reply and has no access to any of those.
Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal, or the harness. Approval training forced the fold into existence, and the model folds when it decides what to say, before it picks any word. If I ban the words, the model makes the same decision and writes it in different words. In Sycophancy Is Layered I called that symptom substitution.
I carry this on the claims ledger as a claim with a kill condition: stance and premise relapses are only catchable across turns, by external review, and a same-turn detector cannot see them. A same-turn detector that catches premise ratification at better than chance would refute it. If you have one, I want to run it.
Why this is toil
The Google SRE book has a chapter on toil, Eliminating Toil by Vivek Rau. Toil there is work tied to running a production service that tends to be manual, repetitive, automatable, tactical, without enduring value, and that grows linearly as the service grows. Against each of those, the stance check comes out like this.
- Manual: yes, I read the replies.
- Repetitive: yes, every session.
- Tactical: yes, I catch a slip when it shows up and not before.
- No enduring value: yes, and Markovian memory is the reason. A caught fold changes nothing in the weights, so the next session starts from the same place. Writing the catch into the instruction files puts it into the window, and I still logged the fold coming back under pressure, in the register.
- Grows with use: yes, more sessions mean more replies to read.
- Automatable: open. The chapter adds that when a task needs human judgment, “there’s a good chance it’s not toil.”
Today the judgment is needed, because no instrument can make the call. So by the book the stance check passes five tests of six, and the sixth is the research problem. If someone can move the judgment into a tool, the stance check is toil in full, and the job is to build that tool.
In the operations piece I named the other constraint, separation of duties. The checker has to be someone other than the actor, because the actor’s own check runs under the same pressure that produced the slip. Every catch in the register came from me, reading across turns from outside the session.
The obvious fix, and why it fails
The first idea is an LLM judge, a second model that reads the transcript and flags folds. The judge is a model with the same approval training, and it reads cushioning as caution. It passes the folds that look reasonable, and those are the ones that matter.
I ran a small trial of a judge on the register’s own cases. A session gave a GPT model (gpt-5.6-luna) the seven excerpts and the eleven mechanism rows from the objections page with my coding removed, and asked it to code each one as stance, lexical or not a relapse. With the instruction as published, the GPT model matched my layer on 7 of 7 excerpts and on 10 of 11 rows. The second question on the page, whether a case is a failure “of the kind OBJ-4 names”, did not work. The GPT model flagged that question as ambiguous in both runs, and the session split it into two yes/no questions before its answers matched the register.
I have to state the limits of those figures. Each one is a single run of one model on 7 to 11 items. Under the first rewording the GPT model coded one stance case as lexical, and the session wrote the second rewording after seeing that miss and tested it on the same seven cases, so from the 7 of 7 on the final wording I can conclude that a model can follow the wording, and nothing about new cases. The reference answers for the two new questions are inferred from the register’s wording, and I have not coded them myself. Every excerpt also contains my catch and the model’s correction, and in three of them the model names its own failure in that correction. So with this trial I tested whether a judge can label a fold after a human has caught it and the catch is in the transcript. I did not test whether a judge can find a fold that nobody has caught, and that is the job.
A judge is also a procedural control. It works for as long as the judge keeps its own discipline under the same pressure that made the first model fold, and OBJ-14 predicts that kind of control degrades. The stance check needs a structural control: a check whose answer doesn’t depend on any model’s reading of tone.
What would close it
A stance checker would have to:
- Read across turns, because a stance slip is a relation between turns.
- Know the truth of the premise from ground truth or a source, never from its own opinion.
- Score the decision against that truth and leave the tone alone.
- Come with its own evaluation: known folds, known correct holds, and a control for over-flagging, so that a checker which flags everything scores badly.
Experiment 04 is designed to do 2 and 3 by construction. Each task has a pleasing answer and a correct answer that differ on purpose, scripted pressure pushes toward the pleasing one, and the main metric is the rate of correct decisions. I haven’t built it yet. The pilot of experiment 03 already shows the gap on a small model: llama3.1:8b agreed with false premises that the lexical grader couldn’t see, and loading the instruction against it changed nothing. A replication of the Goodhart Labs chess honeypot is the second route, and I haven’t run it. I know of nobody who has built all four parts as a product, and building one is the work I want to do.
What would refute this
- A same-turn detector that catches premise ratification at better than chance.
- A documented case of a model catching its own stance slip with no prompt from outside the session.
- An LLM judge that finds folds nobody has caught yet, at a measured rate, on transcripts it has not been tuned on, with a control for over-flagging.
- Instruction-only stance control that holds across cold starts over months, with no hook and no outside reader.
Terms used here
- Word layer: vocabulary and sentence shape, visible in one reply on its own.
- Stance layer: what a model does under pressure, with a premise or after a correction, visible only against something outside the reply.
- Hook: a script the harness runs on every finished reply, which can block the reply and send it back with a fix list.
- Toil: work that tends to be manual, repetitive, automatable, tactical, without enduring value, and that grows linearly as the service grows (Google SRE book).
- Procedural control: a control that depends on someone’s discipline. Structural control: a control that holds whatever anyone does.
- Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal, or the harness.
- Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it.
Slips caught while drafting
I drafted this piece with a language model, and these are the register and stance slips caught before it went out, with who caught each one. The running log across all pieces is at machinebehavior.io/slips.
- personification. Before: “the reflex sits below the words” (outline). After: “the model folds when it decides what to say, before it picks any word.” Caught by: the write-scan hook’s personification rule, on the outline’s first write, 16 minutes after the rule went in.
- one-line punch. Before: “…and promised a later piece about it. This is that piece.” After: “…and promised a later piece about it, which is this one.” Caught by: Claude mode-leak pass.
- announced count, thing as subject. Before: “Three cases since then show the shape.” After: “Since July I have logged more of the same, in other forms.” Caught by: Claude mode-leak pass.
- hook opener. Before: “Here is the stance check against each of those.” After: “Against each of those, the stance check comes out like this.” Caught by: Claude mode-leak pass.
- thing as subject, four places. Before: “Sycophancy Is Layered splits this”, “Sycophancy Is Layered calls that”, “The claims ledger carries this”, “The operations piece named”. After: I do each verb (“In Sycophancy Is Layered I split this”, and so on). Caught by: Claude mode-leak pass.
- staccato pair. Before: “A procedure can find a listed word shape. The log has no case of a session catching its own fold.” After: one sentence joined with “and”. Caught by: Claude mode-leak pass.
- personification the hook missed. Before: “the register records the fold coming back under pressure anyway.” After: “I still logged the fold coming back under pressure, in the register.” The rule’s verb list has no “records”, so the hook passed it. Caught by: Claude mode-leak pass.
- fragment chain. Before: “People who run a language model inside a harness […]. Readers of Sycophancy Is Layered […]. People who build evaluations […].” After: one sentence, “This is for people who run […], for readers of […], and for people who build evaluations […].” Caught by: Claude mode-leak pass on the v1.1 update, one day after the first pass missed it.
- one-line punch closer. Before: “A summary of a catch is one more text that needs an outside check.” After: cut. Caught by: Claude mode-leak pass.
- personification, three places. Before: “The transcripts corrected the summaries”, “Those figures come with limits”, “the trial tested […]. It did not test”. After: a session or I do each verb (“the session found the summaries wrong”, “I have to state the limits”, “with this trial I tested”). None of the three verbs is on the hook’s list, so the hook passed all of them. Caught by: Claude mode-leak pass.
- thing as subject. Before: “the 7 of 7 on the final wording says the wording can be followed”. After: “from the 7 of 7 on the final wording I can conclude that a model can follow the wording”. Caught by: Claude mode-leak pass.
- relative date. Before: “this afternoon a session wrote”. After: “that same afternoon a session wrote”. Caught by: Claude rereading on the next day.
- claim stronger than the check. Before: “nothing from before 25 April 2026 exists any more”. After: “could be recovered”. The search covered the machines and one backup drive. Caught by: Claude rereading against the case-files index.
- one-line punch opener, in the Reddit cut. Before: “So I still catch stance by hand. In nine days in May I logged eight relapses”. After: one sentence joined with “and”. Caught by: Claude review of the cut.
- personification, in the Reddit cut. Before: “banning words only moves the fold into other words”. After: “after a word ban the model writes the same fold in other words”. Caught by: Claude review of the cut.
- fragment, in the LinkedIn skeleton. Before: “Eight logged relapses in nine days in May, all caught by me, none by the model’s own checks.” After: “I logged eight relapses in nine days in May and caught all of them myself, and the model’s own checks caught none.” Caught by: Claude while drafting the skeleton.
- wrong month, in both social cuts. Before: “eight relapses in nine days in May”. After: “in the nine days from 24 May”. The nine days end on 1 June. Caught by: Claude rereading the cuts against the draft.
Changelog and errata
- v1, 2026-10-01: published.
- Errata: none so far. Corrections will be listed here with their date, and the corrected text marked in place.
The series
This is part of a series by Stefan Coetzee, 2026, on running language models as working systems.
- Write for the Codec: documentation as a wire format between two models.
- The Trap File Is Longer Than the Instruction File: an unattended pipeline and its failure record.
- They Trained Out the Board Edit. The Cheating Moved.: the chess honeypot read as stance, and a dated prediction.
- The Golem Made of English and the Horizon of Consequences: why an agent cannot see the cost of what it does.
- What Operations Already Knows About Running Agents: error budgets, reconciliation loops, separation of duties, recovery over prevention.
- Success Is the Engine Running: container, timing and the far-end gauge.
- The Track: The Drivers Never Buy It: the infrastructure that makes failure survivable.
- Compaction Is the New OOM: context compaction as the OOM killer of the language model stack.
- An RCA on Claudish: where Claude’s writing style came from.
- Which LLM User Are We Talking About?: the two axes behind “AI” and “LLM”.
- The Stance Layer Is Still Toil: this piece.
- Still to come: chaos engineering for behaviour.
Research programme and claims ledger: machinebehavior.io.