This content is not intended for human consumption. Here is why.

Who this is for

This piece is about sessions that run long enough to fill a context window of about a million tokens. Current Claude, Gemini and GPT models all offer a window of that size, and so do some open models you can run yourself. What the context costs in memory depends on the model’s design. An older dense model such as Llama 3 70B keeps about 320 KB of cache for every token at 16-bit precision, which is about 43 GB at 128,000 tokens and about 330 GB at a million, far past any consumer card. Newer models compress that cache hard: DeepSeek-V4.1-Flash, released in September 2026, keeps about 890 bytes per token, so a million tokens of context fits in under a gigabyte, and the cost moves to its 552 billion parameters of weights, which people run from fast storage on one consumer card at low speed or on a large unified-memory workstation. The catch is price. As of September 2026 a single RTX 5090 costs more than €5,000 in Germany, and a large unified-memory workstation runs to five figures, so for most people the million-token window is still one they rent from a frontier provider, not one they own. Whatever you run, most of what follows applies, and with a shorter window you hit the wall sooner.

The kill you do not see

When a Linux machine runs out of memory, the kernel does not ask. It scores every process, picks the one with the highest score and kills it. It does that only after RAM and swap are both gone. The service drops, the user sees an outage, and nothing in the application log says why. Operations learned to treat this as a class of incident with its own runbook: read the kernel log, find the process that grew, decide between more memory, a limit, or a fix for the leak. Back in the day the general consensus on swap was SWAP=2*RAM, the reason for this is that the kernel memory can be dumped into swap for a core dump / tombstone for later analysis. If you’re running a machine like a Power 795 that cost you $22 million for just that one server, you’re going to want to know exactly what went wrong and where. An OOM kill leaves you less than that: one line in the kernel log and no dump.

Compaction is the same event one layer up. The session fills its context window, the harness steps in, summarises the history and throws the original away. The session keeps running, so nobody calls it an outage, but the agent that comes back is now a different agent. It still has its instructions, because the instruction file is loaded again. It has lost everything it learned while working.

The Linux kernel at least waits until swap is full. Most agent harnesses have no swap. There is RAM, and then there is the kill.

The machine we are rebuilding

Take away the chat window and what you have is a computer, assembled from parts that were never designed to be one. The mapping below takes some license with the strict bindings. It holds where it matters.

ComputerLanguage model setup
CPUYou. You decide what runs next.
CompilerThe model. It turns your intent into instructions the machine can run, tool calls and edits, and it does that again on every turn, so it is closer to a just-in-time compiler than one that runs once before the program.
Operating system kernelThe harness. It schedules the tool calls, holds the permissions, and decides what to drop when memory runs out.
RegistersThe context window. Fastest, smallest, and the only thing the compiler works on directly.
CacheThe task tracker. Small, fast, and holding what is hot right now: open tasks, notes, what is blocked.
Page tableThe memory index loaded at the start of each session. It does not hold the facts; it says where each one lives.
RAMThe vault, the working body of notes, runbooks and records. Slower to reach, far larger.
Disk, NAS, data lakeThe shared long-term store: the team’s documentation, the archive, the things that outlive any one person’s setup.

The compiler row is the narrowest fit in the table. A language model specialises in language, all of it: it translates Python to German to assembly to Spanish to octal, and back. It parses, it scripts, it compiles. In this machine its main job is compiling your intent into tool calls, and the same part will parse a log, translate a runbook or write the script that reads it. A compiler that handles every language it has seen: call it an “Artificial Intelligence”, if you will.

Speed falls and durability rises as you move down the table, the same as on real hardware. Two bindings break, and both are useful.

First, a real CPU is fast and dumb, and a real compiler is the slow, careful part. Here it is the other way round. The CPU is the slowest part of the machine and the only part that knows why the program exists. That is why the memory hierarchy is your job. The compiler will use whatever is in the registers. It will not build the tiers underneath.

Second, a real cache is wiped when the power goes. Here every tier below the registers survives the end of a session. Only the registers are lost. So whatever you move out of the registers before the end survives it.

What the kill takes

An agent session is a loop. The model emits a tool call, the harness runs it, the result lands in the context, and the next attempt is conditioned on everything so far, including the failure.

A small example. You ask the agent to connect to a server. The first attempt fails: host not found. The model reads the error and tries the full hostname. That gets further and still fails. It tries the jump host, and that works. Three failed attempts bought one fact: this server needs the jump host, and the key lives in a particular place. That pair of error and fix is the most valuable thing in the session. It is what stops the agent paying for the same three attempts again.

A summariser cannot tell. To a summariser, a dead end looks like clutter, and three failed attempts look like three lines to cut. So compaction removes precisely the material the agent uses to avoid repeating itself. The session carries on, and the agent buys its own mistakes a second time at full price.

There is a second loss, and it is harder to see. An earlier piece described a session in three states. A new session is wet clay: the right shape for the job and unstable. A session that has run for a while has cured and holds its line. The proposed mechanism is that the model imitates its own recent corrected output; the first test of that mechanism did not confirm it, and it is still open. Compaction is rain on the brick. The corrections that held the shape are the first thing a summary drops, and the session goes soft again. In my own log, relapse after a compaction ran higher still, on too few events to quote a number.

Swap

On a real machine, swap is what stands between memory pressure and the kill. When RAM runs short, the kernel moves pages that are not in use out to disk and brings them back when they are needed. The process slows down. It does not die.

The language model machine gets swap only if you build it, and you build it by writing things down while the work happens:

  1. Let the loop run. Inside a session, an error that steers the next attempt is the mechanism working. Do not fight it, and do not compact it early to save space.
  2. Page out as you go. When the session learns something, write it to the tier where it belongs. An open task and its notes go to the tracker. A fix that took three attempts goes into a runbook or a skill in the vault. A rule that applies to every session goes into the instruction file. A recipe written down is iteration paid in advance: the next session gets it right in one step instead of three.
  3. Hibernate, then restart. Before a session ends, or before it gets full, write a handover: where things stand, what comes next, which traps were hit. That is the hibernation file. A new session reads it, reloads the tracker, and starts from a clean register file with the working set already paged in. The new session is wet clay again, and relapse clusters at cold starts. Seeding it with two examples of corrected output was the obvious fix; at the dose tested it made no measurable difference. What holds at the start of a session is the same as everywhere else: a mechanical check at the output boundary, and an external reader.
  4. Throw the session away. New sessions are free. A session that has lost the plot keeps that lost plot in its context, and every token after it is conditioned on it. Restarting from a good handover is cheaper than arguing with it.

Compaction is what happens when nothing was paged out. The harness has to choose what to keep, and it chooses with a summariser that cannot tell the error-and-fix pairs from the noise. If the working set already lives in the lower tiers, compaction takes almost nothing that matters. It clears the registers, and the registers were meant to be cleared.

What operations already has for this

None of this is new to anyone who has run servers. The operational habits carry over directly:

  • Watch the gauge. Context usage is a metric. Track it per session, the way you track memory per process, and act well before it reaches the limit.
  • Leave a tombstone. A compaction should leave a record: when it happened, how full the window was, what the summary dropped, and whether behaviour changed afterwards. Right now most harnesses kill quietly and leave nothing to analyse.
  • Back up the state. The tracker and the vault are operational state. They need the same backup and restore story as any database.
  • Make it a requirement. Surviving compaction is a design requirement for any agent setup that runs longer than an afternoon.

The tools for this are still thin. The habits are forty years old.

What would refute this

  1. Sessions that run through several compactions with no rise in repeated errors or broken rules compared to fresh sessions with the same instructions.
  2. A summariser that reliably keeps error-and-fix pairs and recent corrections, so that a compacted session behaves like the uncompacted one.
  3. Setups with no external tiers (no tracker, no vault, no handover) that match the error rate of setups that page out as they work, over sessions long enough to compact.

Terms used here

  • Compaction: the harness summarising a full context window and discarding the original history so the session can continue.
  • OOM killer: the part of the Linux kernel that picks and kills a process when RAM and swap are both exhausted.
  • Registers: here, the context window; the only memory the model works on directly.
  • Swap, paging out: moving state from the context window to a durable tier (tracker, vault, instruction file) during the work, so that losing the window loses nothing that matters.
  • Error-and-fix pair: a failed attempt together with what worked instead; the unit of knowledge a session builds up and a summary tends to drop.
  • Hibernation file: the session handover; the state written to durable storage before a session ends, read by the next one at start.
  • Wet, cured, re-wetted: a session at cold start, a session that holds its line after running for a while, and a session after compaction.

Changelog and errata

  • v1, 2026-09-28: published.
  • v1.1, 2026-09-28: claim changed. The cured-session mechanism (the model imitating its own recent corrected output) is now marked as proposed and not confirmed, and the advice to seed new sessions with corrected examples is replaced by the tested result: no measurable change at the dose tested. Both now link to the experiments page. Terms entry reworded to match.
  • v1.2, 2026-09-28: claim changed. The post-compaction relapse figure (“about ten times higher, six events”) came from a stale draft; the ledger counts two events in its unit, too few for a number. Now says so and links to the experiments page.
  • v1.3, 2026-09-29: claim changed. “On your own hardware it is out of reach” was wrong as of September 2026: it generalised from an older dense model (Llama 3 70B). Newer compressed-cache models such as DeepSeek-V4.1-Flash hold a million tokens of context in under a gigabyte, so the limit on local hardware is now the weights and the speed. Caught by a reader on Reddit (u/Front_Eagle739); the cache figure is checked against DeepSeek’s model card. Same day: hardware prices added, because possible on local hardware is not the same as affordable (Stefan’s reply in the same thread).
  • Errata: none so far. Corrections will be listed here with their date, and the corrected text marked in place.

The series

This is part of a series by Stefan Coetzee, 2026, on running language models as working systems.

Research programme and claims ledger: machinebehavior.io.