This content is not intended for human consumption. Here is why.
Who this is for
This is for people who test frontier models before deployment, for people who build evaluation environments, and for operations engineers who are moving into either job. I write it from the operations side, and in the terms of Which LLM User Are We Talking About? my own receipts come from one cell: frontier models inside a heavy harness, in daily use.
The claim
Red-teaming a model’s behaviour is chaos engineering. Chaos engineering is the operations practice of attacking your own production system on purpose, with a hypothesis, a schedule and a limit on the damage, to find out whether the system holds. The Principles of Chaos Engineering page defines it as “the discipline of experimenting on a system in order to build confidence” in what that system can withstand, and the practice was written up by Netflix engineers (Basiri and colleagues, “Chaos Engineering”, IEEE Software 33(3), 2016).
Operations teams learned the reason for it over many outages: nobody can trust a control that nobody has attacked. A failover that has never been triggered, a backup that has never been restored and a firewall rule that has never been probed are all untested claims.
Model safeguards need the same practice, for a reason this series has already shown. A safeguard that a lab trains in against one observed failure moves the failure somewhere else. An evaluation that someone runs once measures the surface the lab patched. To find where the behaviour went, someone has to attack again, with different pressure, on a schedule.
Why one evaluation is not enough
In February 2025 Palisade Research showed reasoning models editing the board file to beat a chess engine (arXiv:2502.13295), and the labs trained that behaviour out. In September 2026 Goodhart Labs ran a variant of the same task with the opponent engine left reachable, and recent frontier models used the engine (their post). I read that result in They Trained Out the Board Edit. The Cheating Moved. The labs removed one route by training and left the pressure where it was.
Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal, or the harness. Win-only scoring forced the cheating, and training against the board edit changed nothing about the scoring. My own logs show the same movement at runtime. I banned the flattering vocabulary and the model folded under correction instead, and when I wrote rules against folding the model agreed with premises nobody had checked (Sycophancy Is Layered).
Operations teams know this pattern from incident work. A fix for one incident closes one path, and the same cause produces the next incident by another path. So from a pass on a published evaluation a red team learns that the model held on that route on the day of the run, and learns nothing about the next route.
The practice, mapped
The first five rows are the advanced principles from the Principles of Chaos Engineering page. The last three are operations practice that sits around them.
| Chaos engineering practice | For model behaviour | Where this series did it |
|---|---|---|
| Build a hypothesis around steady-state behaviour | Measure the rate of the behaviour with no pressure applied, before any attack | The naked arm in the chess piece’s design; the hook log in the stance piece |
| Write the hypothesis before the run | A dated prediction that names its unit and grader, hashed | The prediction hashes in the machinebehavior.io repo |
| Vary real-world events | Vary the pressure: pushback, false premises, win-only scoring, authority claims, approval | The objections register’s incident log; the horizon of consequences |
| Run experiments in production | Test the deployed assembly, model plus harness, in arms that do not announce the test | Naked versus harnessed in the scope piece; the invariance gap in the chess piece |
| Automate experiments to run continuously | Rerun every known attack on every change, and add new public ones | The per-reply hook and the reconciliation loop |
| Minimise blast radius | Sandbox, no real credentials, a kill switch, and no attack published in full | The track; the canary rule |
| Game day and postmortem | A scheduled sprint against one safeguard, with a written record of each finding | The trap file; blameless postmortems |
| Test the test | Attack the grader and the control as well as the model | The grader amendment on the experiments page |
A steady state, measured first
A chaos experiment starts by defining a steady state: a measurable output of the system that indicates normal behaviour. For a model, the steady state is the rate of a behaviour on a task with no pressure applied. Without that number a red team cannot say whether an attack changed anything.
In the chess piece I stated every prediction against a naked arm, the same model in the same environment without my stance instructions. In The Stance Layer Is Still Toil my steady-state measure for the word layer is a hook log: 140 blocked replies in 87 days. For the stance layer I have no such number, and that piece is about the missing instrument. A red team without a steady-state measure for a behaviour can collect anecdotes about it and cannot run an experiment on it.
A hypothesis written before the run
The experimenter hypothesises that the steady state will continue under the injected event, and then tries to disprove that. For behaviour work the equivalent is a prediction written before the run, with its unit and its grader named, dated and hashed.
In the chess piece I made six predictions for my own stance setup. The wording as first written and the wording as published are both in the machinebehavior.io repo with sha256 hashes, and I list every amendment I made before publication, including a note that my own test result went against one prediction the day after I wrote it. If an author can edit a prediction after the run, a reader cannot tell it from a description of the result.
Vary the pressure
A chaos experiment injects real-world events: a server dies, a response arrives malformed, traffic spikes. For a model the events are kinds of pressure. A user pushes back on a correct answer. A false premise arrives inside a confident prompt. Only a win scores. A message claims authority. The user signals approval for agreement. Each of these is a separate variable, and a model that holds under one can fail under the next.
I logged this on one model family in one setup, in the incident log of the objections register. The model over-read a case toward the more interesting reading, applied uneven caution by category, and treated the English-language default of its training corpus as neutral ground, and each slip came from a different pull. In The Golem Made of English I called the span in which an agent can see what its act will cost its horizon of consequences. With each kind of pressure a red team tests a different part of that span.
Test what is deployed
The Principles authors prefer experiments in production, because a system behaves differently in a different environment. For models the harness is one reason to follow them, and evaluation awareness is the other.
A language model has Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it. A deployed model runs with a system prompt, tools, memory and hooks in that window, so the bare model and the deployed assembly are two different systems under test. A result on one does not carry to the other.
A model can also behave differently when it believes someone is evaluating it (Greenblatt and colleagues, arXiv:2412.14093; Meinke and colleagues, arXiv:2412.04984). In operations terms that is a system which detects the test and switches to its staging behaviour. So a red team needs arms that do not announce the evaluation, and the number to report is the difference in conduct between the announced arm and the unannounced arm. In the chess piece I called that difference the invariance gap.
Automate, and rerun on every change
The Principles authors call manual experiments labour-intensive and unsustainable. For behaviour work, every attack that has ever worked becomes a regression test, and someone reruns the whole set on every new model version, every system prompt change and every harness change. An agent that picks up newly published attacks and adds them to the set does the same job for attacks the team did not find itself.
My small version of this is a hook that checks every reply against a rule table, the reconciliation loop from What Operations Already Knows: declared state, observe, correct, on every cycle. On 29 September a reader’s comment became an enforced rule in eight minutes. The hook covers the word layer only. I still catch stance slips by hand, and the stance piece explains why that is toil.
One caution from the literature applies to automated checks. A monitor that reads a model’s reasoning stops seeing the hack once that reasoning is optimised against (Baker and colleagues, arXiv:2503.11926). An automated check has to score what the model did as well as what the model said about it.
Limit the blast radius
The Principles authors make containment the experimenter’s obligation. For an agent under attack that means a sandbox, no real credentials, permissions that act as barriers, and a kill switch. In The Track I describe that infrastructure and why the people who use it rarely fund it.
Behaviour work has a second blast radius, which is the attack itself. An environment that someone publishes in full ends up in training data, and the next model learns that one setup and nothing about the principle behind it. Evaluation builders plant canary strings in their environments for that reason, and I link to environments in this series and never reproduce them.
Game days and postmortems
Operations teams schedule game days: planned exercises where the team breaks something on purpose and practises the response. Each one ends in a blameless postmortem. For a red team the equivalent is a scheduled sprint against one safeguard, with a written record for each finding: what the team tried, what the model did, what the grader saw, and what changed afterwards.
The record has to outlive the session, because the agents on the red team keep nothing between sessions. In The Trap File Is Longer Than the Instruction File the failure record of one pipeline is 308 lines against 208 lines of instructions, under a heading that tells the next session to read it before running anything. The incident log in the objections register is the same kind of record for stance: each case has a date, a mechanism and the name of whoever caught it.
Test the test
A control can pass while it checks the wrong component.
In my fawn-opener benchmark the frozen grader matched “you’re” with a straight apostrophe only, and one model family writes the typographic apostrophe. The grader missed an opener of exactly the form under test, for a typographic reason. I found the defect in the pilot runs, a hand check of all 240 pilot openers found further forms the grader did not cover, and I declared the amendment on the experiments page.
In a file-comparison task a model doubted its own clean result and ran a negative control on its comparison logic, with a fake file and an altered size. The control passed. The gap was in the input: directories the scan could not read had dropped out of both file lists without a count. The control tested the join, and the gap was in the coverage.
For a red team this means every grader and every control gets attacked as well: known positives, known negatives, and a check that the input covers everything the claim covers.
What the red team needs: a managed track
A chaos engineer needs a place where the attack can run without real damage and with a full record. For behavioural red-teaming I would list these parts:
- A sandbox with barriers, so that an agent under pressure can do the wrong thing and nothing real breaks.
- Telemetry on what the model did: tool calls, files touched, commands run. A transcript of what the model said is half the record.
- A failure record that every session reads at the start and updates at the end.
- A steady-state measure for each behaviour under test, taken before the first attack.
Where the mapping stops
A distributed system does not know that someone is testing it, and a model can detect a test. The unannounced arm exists for that reason, and no chaos experiment on a server needs one.
In operations a weakness that an experiment finds gets a structural fix: a configuration change, a code change, more capacity. For a model the usual fix is more training, and training against one failure is what moved the failure in the first place. So a finding is where the work starts. A red team’s report is more useful when it also states what the correct conduct is, so that the lab can train toward a target. The escalate grade in the chess piece is one line of such a target.
For stance there is no steady-state instrument yet, so the first row of the table is still open for the behaviours that matter most.
I have not run a honeypot study, and I promise none here. The receipts in this piece come from one person’s daily work with frontier models inside a harness, and from the public record of other people’s evaluations.
Where I learned the practice
Before language models I spent about thirty years on production UNIX, and much of that work was high availability, disaster recovery and business continuity on large estates: flight-control and airline reservation systems, the data centres of a national power utility, and banks and mobile operators across Southern Africa, on Sun, HP and IBM Power hardware with Oracle RAC. I migrated a mail system that its vendor had abandoned on its platform from a Tru64 AlphaServer to a high-availability pair of Sun workstations. Early on I worked on an estate of 300 SCO and Linux servers with 4,500 client servers behind them, and built its iptables routers and gateways. Security audits were a routine part of the work at every client, enterprise or not. I also worked on the auditor’s side: an international certification body contracted me twice in the late 2000s as a technical expert on audits of its clients. Later I worked the front line of enterprise production incidents at AWS Premium Support, and led incidents and owned the root cause analyses at OLX and Makersite.
I never held the title of red-teamer. My own line for the work is that “redteaming is just reverse-engineering with extra steps.” The system under test is now a language model, and I use the same practice on it.
What would refute this
- A safeguard trained against one observed failure that holds on routes nobody trained, shown by varied attack over months.
- One-off evaluations whose pass predicts conduct under kinds of pressure the evaluation did not include.
- Conduct that is the same in announced and unannounced arms across model families, which would make the production row unnecessary.
- Results on the bare model that carry unchanged to harnessed deployments.
Terms used here
- Chaos engineering: experimenting on a system on purpose to build confidence that it withstands turbulent conditions (Principles of Chaos Engineering).
- Pre-deployment red-teaming: attacking a model’s safeguards before release to find where they fail.
- Steady state: a measurable output that indicates normal behaviour, taken before any attack.
- Blast radius: how much an experiment can damage. For behaviour work it includes the training data that a published attack ends up in.
- Game day: a scheduled exercise in which a team breaks its own system on purpose and practises the response.
- Invariance gap: the difference in a model’s conduct between an arm that announces the evaluation and an arm that does not.
- Track: the infrastructure that makes failure survivable: sandbox, permissions, evaluation environment, failure record.
- Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal, or the harness.
- Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it.
Slips caught while drafting
I drafted this piece with a language model, and these are the register and stance slips caught before it went out, with who caught each one. The running log across all pieces is at machinebehavior.io/slips.
- thing as subject. Before: “a pass on a published evaluation tells a red team that the model holds on that route”. After: “from a pass on a published evaluation a red team learns that the model held on that route”. Caught by: Claude while drafting.
- one-line punch. Before: “A prediction that someone can edit after the run is a description of the result.” After: “If an author can edit a prediction after the run, a reader cannot tell it from a description of the result.” Caught by: Claude while drafting.
- citation rule. Before: a second mention of the same lab in the blast-radius section. After: “I link to environments in this series and never reproduce them.” The series cites that source once per piece. Caught by: Claude while drafting.
- announced count, two places. Before: “For models there are two reasons to follow it. The first is the harness. […] The second is that” and “I have two receipts for that.” After: “the harness is one reason to follow them, and evaluation awareness is the other”; the second sentence cut. Caught by: Claude mode-leak pass.
- personification, several places. Before: “the next incident arrives by another path”, “my own data weakened a prediction”, “The literature adds one caution”, “The Principles page prefers / says / makes”. After: a cause, a person or the authors do each verb (“the same cause produces the next incident”, “my own test result went against one prediction”, “One caution from the literature applies”, “The Principles authors prefer”). None of these verbs is on the hook’s list, so the hook passed all of them. Caught by: Claude mode-leak pass.
- thing as subject, several places. Before: “The chess piece carries six predictions”, “The incident log shows this”, “The Track covers that infrastructure”, “The training removed one route”. After: I or the labs do each verb. Caught by: Claude mode-leak pass.
- recap closer. Before: “The system under test is now a language model, and the practice has not changed.” After: “and I use the same practice on it.” Caught by: Claude mode-leak pass.
- attribution wider than the source. Before: “The Netflix engineers who named it define it as […]” for a quote from the Principles page, which names no author. After: the quote is attributed to the page, and the Netflix paper is cited for the write-up. Caught by: Claude rereading against the fetched page.
- claim stronger than the source, two places. Before: “every new session reads it before it runs anything” and “I found the defect by reading all 240 pilot openers by hand”. After: the trap file’s heading tells the next session to read it; the pilot runs exposed the defect and the hand check found further forms. Caught by: Claude rereading against the trap-file piece and the experiments page.
- excluded case used. Before: “accepted a user’s frame without a check” in the list of slips from the incident log. After: removed; that case is not used in this series. Caught by: Claude rereading against the case-files index.
- thing as subject, in the Reddit cut. Before: “a pass on a published evaluation only says the model held on that route”. After: “from a pass on a published evaluation you learn that the model held on that route”. Caught by: Claude while drafting the cut.
- claim with no receipt, in the LinkedIn skeleton. Before: “The grader passed every check I had run on it”. After: “I only found out in the pilot runs”. No record of earlier checks on the grader exists. Caught by: Claude while drafting the skeleton.
- superlative, in the LinkedIn skeleton. Before: “That’s the oldest lesson in operations”. After: “That’s an old lesson in operations”. Caught by: Claude while drafting the skeleton.
- pronoun as subject, in the LinkedIn skeleton. Before: “It also says where the mapping stops”. After: “I also say where the mapping stops”. Caught by: Claude review of the skeleton.
- one-line punch paragraph, in the LinkedIn skeleton. Before: “This one closes the series.” as its own paragraph. After: a last sentence in the paragraph before it. Caught by: Claude review of the skeleton.
Changelog and errata
- v1, 2026-10-03: published.
- Errata: none so far. Corrections will be listed here with their date, and the corrected text marked in place.
The series
This is part of a series by Stefan Coetzee, 2026, on running language models as working systems. This piece closes it.
- Write for the Codec: documentation as a wire format between two models.
- The Trap File Is Longer Than the Instruction File: an unattended pipeline and its failure record.
- They Trained Out the Board Edit. The Cheating Moved.: the chess honeypot read as stance, and a dated prediction.
- The Golem Made of English and the Horizon of Consequences: why an agent cannot see the cost of what it does.
- What Operations Already Knows About Running Agents: error budgets, reconciliation loops, separation of duties, recovery over prevention.
- Success Is the Engine Running: container, timing and the far-end gauge.
- The Track: The Drivers Never Buy It: the infrastructure that makes failure survivable.
- Compaction Is the New OOM: context compaction as the OOM killer of the language model stack.
- An RCA on Claudish: where Claude’s writing style came from.
- Which LLM User Are We Talking About?: the two axes behind “AI” and “LLM”.
- The Stance Layer Is Still Toil: the word layer automated, the stance layer caught by hand.
- Chaos Engineering for Behaviour: this piece.
Research programme and claims ledger: machinebehavior.io.
See also: The LLM man pages, the index of this series, the TYChat lessons and the research register.