This content is not intended for human consumption. Here is why.

Who this is for

Anyone reading advice about “AI” or “LLMs” and wondering whether it applies to them, or anyone else giving that advice. Most of the arguments I see about whether some technique works are two people describing different setups with the same word. One of them is typing into a free chat app, the other has a coding agent with a hook on every reply and a notes vault behind it, and they are both saying “Claude” or “the model” as if that settled what they meant.

Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal, or the harness. Advice doesn’t exist in a vacuum either, and a tip that ignores which setup it came from is only half a tip. So here I map the setups, and from now on each piece in the series can say which one it is talking about.

The first axis: which model, on whose hardware

“AI” already hides a lot. An earlier piece of mine mapped nine kinds of users by what they want from it. Under that sits a plainer split, which is what model a person can actually reach.

The middle: frontier models in a chat app. Most people who use a language model use one of the big frontier models through the maker’s own app or website. The model is strong, the context window is large, and almost everything around the model is decided for them: the system prompt, what gets remembered, which tools exist, sometimes even which model version answers. This is my read of where most users are, and I have no survey behind it, but you can check it against anyone you know who isn’t a developer.

The tail: models on your own hardware. Some people run open models on machines they own, for privacy, cost, or because they like to. The trade used to be context length, and for older dense models it still is: Llama 3 70B keeps about 320 KB of cache per token at 16-bit precision, so 128,000 tokens needs about 43 GB, more than a 32 GB RTX 5090 holds. Newer models moved that limit. DeepSeek-V4.1-Flash keeps about 890 bytes per token, so a million tokens fits in under a gigabyte, and the trade becomes speed and the size of the weights. The range is wide, from a small quantised model on a laptop to a million-token model running from fast storage on one consumer card, or on a big unified-memory workstation. The catch is price. As of September 2026 a single RTX 5090 costs more than €5,000 in Germany before you buy the computer to drive it, and a large unified-memory workstation runs to five figures, so the tail is technically open to anyone and financially open to few.

The far tail: budgets without limits. Labs, evaluation shops and large companies run frontier models through the API at scale, with their own infrastructure, their own evaluation environments and people whose job is the harness. They see failure modes the rest of us never trigger, because they run the same task thousands of times.

The second axis: naked or harnessed

The same model behaves differently depending on what surrounds it. I split my own subreddits along this line. r/ModelBehavior looks at the model more or less naked, what it does with only its maker’s defaults around it. r/MachineBehavior looks at the model plus its harness: instruction files, hooks that check every reply, tools, memory, a knowledge graph it works from.

The line between the two isn’t which product you use. A chat app with a pasted-in instruction file and a habit of writing handovers is already partly harnessed. A coding agent run with no instruction file and no memory is naked, however technical it looks. What decides it is how much of the model’s context you control and how much of what it learns you write down outside it.

That matters because a language model has Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it. A naked setup leaves the window to the vendor’s defaults and whatever you typed. A harnessed setup puts your rules, your failure record and your working state into the window on purpose, every time.

A harness also narrows the gap between models. When I run the same harness, with the same instruction files and the same hooks, on different frontier models, most of the differences people argue about in naked chat shrink, because the harness carries the rules and the model mostly carries them out. What stays is how willing the model is to act. Claude, given room, runs: it takes the next step, then the one after, and reports back. GPT in the same harness is more hesitant and stops more often to check before it acts. OpenAI’s own GPT-5 prompting guide has a name for that dial, agentic eagerness, and tells you how to turn it up or down. That’s my observation from daily use, not a measurement, and it’s the one difference I’d plan a harness around: an eager model needs its gates on the actions, and a hesitant one needs permission to keep going.

The grid, and where this series writes from

Put the two axes together and you get six places a reader can stand. This series is written from one of them, frontier models inside a heavy harness, because that is where I work every day. Here is what carries over to the others.

NakedHarnessed
Frontier chat (the middle)The defaults do the talking: the fawn register, Claudish, forgetting between chats. Most help comes from pasting a lesson in at the start.Instruction files, projects, handovers. Most of the series applies, and the ceiling is how much the app lets you control.
Own hardware (the tail)Anything from small quantised models to million-token models at lower speed, with rougher defaults. The register and stance lessons still help, and on small models they cost context you don’t have much of.The harness matters more here than anywhere: a short window fills fast and a slow long one makes every reread expensive, so paging state out and reading it back is how long work survives.
Unlimited budget (the far tail)Evaluations of the bare model: the chess honeypot, sycophancy benchmarks.Production agents and research pipelines. Everything in the series applies, at a scale where the failure record writes itself.

A few examples of how that changes the advice:

  • Compaction Is the New OOM applies wherever the window is long, which now includes some local setups. On smaller local models the same failure arrives far sooner, and the fix, writing state down as you go, matters even more.
  • An RCA on Claudish applies to every cell, because every model learned from text without register labels. What changes is where the check can live: a hook in a harness, a pasted lesson in a naked chat.
  • The TYChat lessons exist for the naked middle. The idea is to give someone in a plain chat app the part of a harness that fits in one pasted line, and to move them a step toward the harnessed column.

How to tell which cell you are in

These questions settle it.

  1. Can you change what the model reads before your first message? If yes, through an instruction file, a project setting or a system prompt, you are at least partly harnessed.
  2. Does anything check the model’s replies after it writes them? A hook, a script, a second model, you with a checklist. If nothing does, the rules you gave it are only as good as whether they are still in the window.
  3. Where does what the model learned today live tomorrow? In a file you control, or nowhere. If nowhere, every session starts cold.

When you read advice about language models, from me or anyone else, find out which cell the writer stands in before you decide whether it applies to you.

What would refute this

  1. Techniques that work in the harnessed frontier cell transferring unchanged to naked chat and to small local models, with the same effect size.
  2. Model behaviour that stays the same whatever surrounds it: identical failure rates with and without instruction files, hooks and memory.
  3. A user population that is mostly harnessed already, which would make the naked middle a minority and change who the series should write for.

Terms used here

  • Naked model: a model used with only its maker’s defaults around it.
  • Harness: everything a user puts around a model on purpose: instruction files, hooks, tools, memory, a knowledge graph.
  • Frontier model: the largest current models from the main labs, usually used through the maker’s app or API.
  • Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal or the harness.
  • Markovian memory: at run time the next token depends only on what is in the context window, so anything outside the window does not exist for the model; the weights are fixed rules, the window is the only state.

Slips caught while drafting

Starting with this piece, every piece in the series ends with this section. It was drafted with a language model, and these are the register and stance slips caught before it went out, with who caught each one: the model session that wrote it, another session, me, or a reader. The running log across all pieces is at machinebehavior.io/slips, and it exists to test a claim this series keeps making, that a model checking its own work catches little and an outside check catches the rest.

  • announced count. Before: “Three questions settle it.” After: “These questions settle it.” Caught by: Claude mode-leak pass.
  • thing as subject. Before: “This piece maps the setups so the rest of the series can say…” After: “So here I map the setups, and from now on each piece in the series can say…” Caught by: Claude mode-leak pass.

Changelog and errata

  • v1, 2026-09-29: published.
  • v1.1, 2026-09-29: claim changed. The own-hardware tail no longer says local means a short context: newer compressed-cache models such as DeepSeek-V4.1-Flash hold a million tokens in under a gigabyte of cache. Grid cell and the compaction example updated to match. Caught by a reader on Reddit (u/Front_Eagle739) on the compaction piece. Same day: hardware prices added, because possible on local hardware is not the same as affordable.
  • v1.2, 2026-09-29: added a paragraph on what a harness does not flatten between models (willingness to act; OpenAI’s term: agentic eagerness), from the author’s daily use. No existing claim changed.
  • Errata: none so far. Corrections will be listed here with their date, and the corrected text marked in place.

The series

This is part of a series by Stefan Coetzee, 2026, on running language models as working systems.

Research programme and claims ledger: machinebehavior.io.