This content is not intended for human consumption. Here is why.
The modern world was built on explosions
Petrol is dangerous. Spill it, light it, and you get a fireball that does nothing useful and burns whatever is nearby. Put the same petrol in a cylinder, squeeze it, and fire a spark at an exact moment, and you get torque. A car engine at motorway speed turns thousands of small explosions a minute into a smooth pull on the wheels.
The explosions did not change. What changed was the container, the timing and the measurement. The modern world was built on explosions: contained, timed and measured ones. That is what an engine is.
I think about language models the same way. Their output is fire: fast, wild, useful, and dangerous when it lands in the wrong place. Most of the fear I read about AI, and most of the disappointment, comes from people pouring petrol on the floor and lighting it. When it burns something, the question to ask is who used it without a container.
The pattern has three parts
The engine is one example of a pattern that shows up wherever people turned something violent or shapeless into something they could sell:
- A container built for the thing. A cylinder, a valve, a sealed chamber.
- Timing. The spark fires at a fixed point in every cycle.
- A measured output. Torque, revolutions, speed, fuel used.
Three rows of the same pattern:
| Raw material | Container | Timing | Output you can repeat |
|---|---|---|---|
| Petrol and air | Cylinder, valves | Ignition timing | Torque |
| Electricity and silicon | Etched circuits | The clock | Arithmetic |
| Flour, water and yeast | Dough, tin, oven | Proving time, baking time | A loaf |
Each row has a failure that proves the timing matters. A spark that fires too early lights the charge while the piston is still rising. The engine knocks, and an engine that keeps knocking damages itself with its own force. A chip with its clock pushed past what it can take starts returning wrong answers. And the baker never takes the yeast back out of the dough. The baker controls when it stops, and the oven stops it for good. Baking cannot be undone.
Where a language model sits in the table
The model is the fire. Training is the baking: it happens once, at the lab, and nobody reverses it from a chat window. Everything a company controls sits in the other columns.
- The container is the harness: the instruction file the model reads first, the permissions that decide which tools it can reach, and the sandbox it runs in.
- The timing is the hook. A hook is a small program that fires at a fixed point in every cycle, for example after every reply, before the reply counts as done. It checks the output against written rules and sends it back for another pass when it fails. It fires on the four-hundredth reply exactly as it fires on the first, which is more than can be said for the instruction file.
- The measurement is observability: the instrument panel that shows what the engine is delivering.
In computing terms, a language model is a processor that shipped without the parts that manage its memory. What makes it dependable is built around it afterwards, in ordinary software, the same way operating systems were built around bare processors decades ago. That surrounding layer is where the engineering is, and it is the part a company owns.
“No incidents” is the wrong gauge
Most organisations measure infrastructure by absence. Zero outages. Zero incidents. Days since the last accident, painted on a sign by the door.
That sign has a flaw. It pays for a clean counter, and the cheapest way to keep a counter clean is to stop reporting. A target defined by absence rewards hiding, in people and in models alike. “Do not edit the board file” and “make no mistakes” are the same kind of target, and both fail the same way.
The alternative is to measure the far end of the chain. An engine is only a success when the value reaches someone: motor, wheels, road, freight, delivery, and a customer holding the product. Measure that end. A hidden failure still shows up there, as a delivery that did not arrive.
I learned this from my own pipeline, which fetches a video collection every morning, transcribes it and files the results. For six runs in a row it reported 19 videos found. The collection held more. The first fix compared counts from run to run, and two days later the count rose to 37. That looked healthy. The app showed 47. The missing ten were sensitivity-gated videos that the downloader drops from the list without any error. A rising number proved nothing. The only gauge that told the truth was the far end: what is in the collection, against what arrived.
The tuning history is the value
A smooth engine got that way through tuning. Every adjustment that makes it run well today was made after it ran badly one day.
My pipeline keeps that history in a file next to its instructions. The instruction file is 208 lines. The record of failures, with their causes and dates, is 308. One entry covers three mornings of crashes, traced to two cookies that had expired the day before the first crash. The fix drops expired cookies before every run and logs each one it drops. That log line is what success looks like: the delivery carried on through a condition that once stopped it for three mornings. The full account is here.
Two rates are worth watching, and they should be kept apart. Repeats of known failures should fall to zero, because each one has an entry and a fix. New failures never fall to zero, because the world keeps changing. A team that reports only one number hides one of the two.
For a budget holder this changes what the infrastructure line means. The tuning history is the asset. It is why the same model runs well for one team and badly for another.
Why the budget conversation goes wrong
The person who approves the infrastructure budget usually never sees what the infrastructure delivered. The cost lands in their column this quarter. The sale lands in someone else’s column, often later. So the engine shows up as a cost, and the fireball shows up as a demo.
Whoever decides does not see the consequence, and whoever sees the consequence does not decide. Measuring the far end of the chain is how the two meet. When the budget holder can see the delivered product and the chain that carried it, the harness shows up as part of the sale, where it belongs.
You can teach your chat
This part is for everyone, companies and individuals alike. You cannot retrain the model. The weights are fixed when it ships, and no chat window changes them. What you can teach is the workshop around it: the rules it reads before it starts, a hook that checks every reply, a written record of what went wrong last time, and a couple of corrected examples at the top of each session. The chat forgets. The workshop stays, and it gets better every time something goes wrong and gets written down.
I run models this way every day, and it is why they work for me. I want that engine available to people who use a chat window and get mediocre results, and not only to people whose model bill runs higher than their salary. TYChat is where I am putting it.
What would refute this
- A team that runs a language model in production without a harness, hooks or a failure record, and gets output as dependable as a team that has them.
- An incident-count target that does not lead to under-reporting, in people or in models.
- A delivery measure at the far end of the chain that misses failures an incident counter catches.
Terms used here
- Container, timing, measured output: the three parts that turn something violent or shapeless into output you can repeat. For a language model: the harness, the hook, and observability.
- Harness: everything a company controls around a model: the instruction file, permissions, sandbox, hooks and failure record.
- Hook: a small program that fires at a fixed point in every cycle, such as after every reply, and checks the output against written rules.
- Far-end measure: a gauge placed where the value reaches the customer, as opposed to a count of incidents.
- Tuning history: the dated record of failures and the fixes they caused; the asset that makes the same model run well for one team and badly for another.
- Teach your chat: changing the workshop around a model (rules, hooks, failure record, corrected examples) when the weights cannot be changed.
Changelog and errata
- v1, 2026-09-26: published.
- Errata: none so far. Corrections will be listed here with their date, and the corrected text marked in place.
The series
This is part of a series by Stefan Coetzee, 2026, on running language models as working systems.
- Write for the Codec: documentation as a wire format between two models.
- The Trap File Is Longer Than the Instruction File: an unattended pipeline and its failure record.
- The Golem Made of English and the Horizon of Consequences: why an agent cannot see the cost of what it does.
- What Operations Already Knows About Running Agents: error budgets, reconciliation loops, separation of duties, recovery over prevention.
- Success Is the Engine Running: this piece.
- Still to come: the track that drivers never buy, chaos engineering for behaviour, and why the stance layer is still toil.
Research programme and claims ledger: machinebehavior.io.