This content is not intended for human consumption. Here is why.
Who this is for
This is for CTOs, heads of engineering and AI enablement leads whose AI pilots worked in the demo and stalled in production, and for the people hired to fix that, under titles like knowledge engineer and context engineer. I write it from operations, after about thirty years of running production systems and two years of running language models inside them.
The claim
Knowledge infrastructure is the written, indexed, checked and maintained knowledge that a language model runs on. Companies that adopted AI bought the model and skipped this layer. Their code is version-controlled, reviewed and tested. Their processes, decisions, exceptions and failure history exist in people’s heads, chat threads and a wiki nobody has opened since the migration. The model can use the first kind of knowledge and cannot use the second.
More of the failed pilots trace back to that gap than to the models. The fix is an engineering job: build the knowledge layer the model reads from, and keep it current with the same discipline you apply to code. I build that layer, and below I explain why it works and what it contains.
Why pilots stall
A language model has Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it. Training fixed the weights before your company ever called the model. Every session starts from those weights and whatever text the session puts in the window. Nothing the model learned in yesterday’s session carries into today’s unless someone wrote it down and loaded it again.
Organisations lose unwritten knowledge too, with no model involved. The US National Nuclear Security Administration lost the knowledge of how to make a material called Fogbank, which it needed to refurbish the W76 warhead. The Government Accountability Office found that NNSA “had kept few records of the process when the material was made in the 1980s and almost all staff with expertise on production had retired or left the agency” (GAO-09-385, March 2009). NNSA spent $69 million to address the Fogbank production problems, and the first refurbished warhead slipped a year, from September 2007 to September 2008. For a model the loss is immediate: anything nobody wrote down is unavailable from the first session.
The authors of the best-known study of enterprise AI results say the same thing in business language. MIT NANDA’s report “The GenAI Divide: State of AI in Business 2025” (Challapally, Pease, Raskar and Chari, July 2025) found that “95% of organizations are getting zero return” on generative AI. The headline number is interview-based, and the authors call their figures “directionally accurate”. Their diagnosis is the more useful part:
The core barrier to scaling is not infrastructure, regulation, or talent. It is learning. Most GenAI systems do not retain feedback, adapt to context, or improve over time.
Later they describe “current systems that require full context each time”. That is Markovian memory, observed from the buyer’s side. By “infrastructure” the authors mean compute and platforms, and on that reading I agree with them. The learning they find missing has to be stored somewhere a model can read on every run. A vendor can sell a memory feature inside one product. The company’s own knowledge, written and maintained, works with every model and every vendor, and the company owns it.
Where it already works: code
Coding is one of the places where AI has gone furthest inside companies: the MIT report names software engineering among the functions where companies on the right side of the divide see workforce impact. I put that down to the state of the knowledge coding agents read. A codebase is written down by definition. It has version control, so every change has an author and a reason. It has tests, so a wrong change fails loudly. It has review, so a second reader checks each change. A model dropped into a well-kept repository has its context in front of it.
The DORA team put it in one line in their 2025 report on AI-assisted software development: “AI’s primary role is as an amplifier, magnifying an organization’s existing strengths and weaknesses.” A team with maintained knowledge gets more from the model. A team whose knowledge was never written down gets that gap back at higher speed.
The same model, pointed at a business process, finds none of that. No single source of truth, no record of why the exception exists, no check that fails when an answer is wrong. It fills the gap with plausible text, because a model has no way to tell a missing fact from an unknown one unless someone taught it to say so.
Knowledge management failed once. Two things changed
The objection I hear first is that companies tried this already. The knowledge management wave of the 1990s and 2000s filled wikis and intranets that went stale within a year. The pages rotted because almost nobody read them, and keeping them current cost more than they returned.
Both halves of that failure have changed since.
- A reader now exists. A model reads the knowledge base on every task it runs. A document that a person opened once a year now gets read on every task, many times a day, and every stale line produces a wrong answer that someone notices.
- Upkeep got cheap. The model can draft and update the documents it reads, and operations practice can check them mechanically: a hook that tests every output against written rules, a failure record read at the start of every session, a check that flags a document nobody has touched since the process changed. Companies gave up on the old wikis because upkeep cost people’s time. Most of that cost now falls on tooling.
What knowledge infrastructure contains
These are the components I build. I run a working instance of each one in my own setup, linked where it is public.
| Component | What it does | Working instance |
|---|---|---|
| Knowledge base | One fact per file, linked into maps, so a model can load exactly what a task needs | My vault; the starter in vault-kit |
| Entry points | One index per domain that the model reads first | Instruction files, llms.txt, the LLM man pages |
| Failure record | What went wrong, the fix, and the rule that prevents it, read at session start | The Trap File Is Longer Than the Instruction File; the slips log on machinebehavior.io |
| Checks at the output boundary | Every output tested against written rules before anyone acts on it | vestige-kit hooks |
| Verification | Claims checked against the artifact; controls attacked on purpose | Chaos Engineering for Behaviour |
| Retrieval | Search across the base, with a private boundary that search respects | vsearch in vault-kit |
| Wire format | Documents written for a model to read first and a person second | Write for the Codec |
The term itself is older than language models. Paul Edwards, in A Vast Machine (MIT Press, 2010), defines knowledge infrastructures as “robust networks of people, artifacts, and institutions that generate, share, and maintain specific knowledge about the human and natural worlds”. His subject was climate science. The definition fits a company that wants its models to know what its people know: the people, the documents and the routines that keep the documents true.
Why this work is easy to miss
Foundation work shows up only when it fails. When it holds, nobody looks at it, so every claim about it needs a trace someone can check. Three of mine:
- OLX. Site reliability engineers embedded in development teams burned out and left after eight to twelve months, and their knowledge left with them. I took data to HR showing that embedded SREs never got promoted, and five dedicated SRE manager roles followed. My direct reports stayed for my whole tenure.
- r/Leathercraft. I founded it in 2011, grew it into the main community for the craft, and handed it over. It is approaching 900,000 members in 2026 and runs without me.
- This series. I logged seven cases where a model slipped back into behaviour I had written rules against. The model caught none of them itself, and I caught all seven (What Operations Already Knows). So I moved the check out of the model: a hook at the output boundary now tests every reply against the word rules and blocks the ones that break them.
The whole project is a running showcase
Everything I have published this year is one knowledge infrastructure, built in public, for one domain: running language models as working systems. It is open, and anyone can explore it, with a model or by hand.
- The thesis is the series “Running the Golem”, listed below. In each piece I argue one part of why models fail without written, maintained knowledge and operations checks.
- The entry point is the LLM man pages, an index of the series written in the format of a Unix manual, so a model can load it and find any part.
- The evidence is machinebehavior.io: claims with dated predictions, an objections register, case files and the slips log.
- The tools are vault-kit and vestige-kit, with a plugin package for Claude Code in progress.
- The smallest unit is TYChat: lessons a person pastes into a plain chat app, for anyone without a harness.
Each site publishes an llms.txt file, the entry point a model reads first: machinebehavior.io/llms.txt, tychat.io/llms.txt and uncovertechtalent.com/llms.txt. The case files and short cuts go to two subreddits, r/ModelBehavior and r/MachineBehavior, and the long pieces to Substack.
A buyer can check the method on the method’s own output before talking to me. Load the man pages or one of the llms.txt files into a model, ask it about any claim in the series, and see whether it finds the source.
Where a company starts
A team gets most of the early return from three steps, and none of them needs a new product.
- Write one entry point per domain. A single file the model reads first: what this system is, where its knowledge is, which documents are current, and what the model must never do. For code this is the instruction file at the repository root. For a business process it is the page that does not exist yet.
- Start a failure record. Every time the model gets something wrong, write down what happened, the fix and the rule that prevents a repeat, and load the record at the start of every session. In my own pipeline that file grew longer than the instructions within months, and most of the value is in it.
- Put one check at the output boundary. Pick the most expensive kind of wrong answer and test for it mechanically on every output. A check that runs every time beats an instruction the model reads once.
After those three, the work becomes ordinary engineering: retrieval across the base, freshness checks, review of changes to the knowledge itself, and attacks on your own controls to see which ones hold. I build this for companies as a fixed-scope knowledge infrastructure install.
What would refute this
- Enterprises whose AI deployments succeed at the same rate whether or not their process knowledge is written down and maintained.
- A model that improves across sessions on company-specific work without any written context loaded, through vendor memory alone, and keeps those gains across a change of vendor.
- Knowledge bases built for models that rot at the rate of the old wikis, despite being read on every task.
- Coding agents that perform as well on poorly documented, untested repositories as on well-kept ones.
Terms used here
- Knowledge infrastructure (for LLMs): the written, indexed, checked and maintained knowledge a language model runs on, plus the routines that keep it true.
- Markovian memory: at runtime the next token depends only on what is in the context window, so anything outside the window does not exist for it.
- Entry point: the one file a model reads first in a domain, pointing to everything else.
- Failure record: what went wrong, the fix and the preventing rule, loaded at the start of every session.
- Output boundary: the point where a model’s output leaves the model and someone acts on it; the place to put mechanical checks.
- Wire format: a document written for a model to read first, as in Write for the Codec.
- Nothing exists in a vacuum: every behaviour a model shows was forced into existence by something, the corpus, the training signal or the harness.
Slips caught while drafting
I drafted this piece with a language model. Slips caught before it went out, with who caught each one:
- personification. Before: “That gap explains more of the failed pilots than the models do.” After: “More of the failed pilots trace back to that gap than to the models.” Caught by: Claude review.
- personification. Before: “The cost that killed the old wikis moved from people to tooling.” After: “Companies gave up on the old wikis because upkeep cost people’s time. Most of that cost now falls on tooling.” Caught by: Claude review.
- personification. Before: “this piece explains why it works”. After: “below I explain why it works”. Caught by: Claude review.
- personification, five more of the same kind (a report that “means”, a study that “says”, a team that “puts it”, a step that “gives”, a file that “carries”): each rewritten with a person or the model as the subject. Caught by: Claude review.
- unsupported claim. Before: “Coding agents are the clearest case of AI working inside companies”. After: coding as “one of the places where AI has gone furthest”, sourced to the MIT report’s list of functions with workforce impact. Caught by: Claude review.
- invented number. Before: “read hundreds of times a day”. After: “read on every task, many times a day”. Caught by: Claude review.
- precision. Before: “thirty years”. After: “about thirty years” (1997 to 2026). Caught by: Claude review.
- wrong attribution. Before: “My own self-audit caught 0 of 7 of my logged slips.” After: the model caught none of the seven relapses and Stefan caught all seven, with the source piece linked. Caught by: Claude, checking the source before shipping.
- one-line punch paragraph. Before: “Both halves of that failure changed.” as its own line. After: kept as the lead sentence of the list it introduces. Caught by: Claude review.
Changelog and errata
- v1, 2026-10-03: published.
- v1.1, 2026-10-03: added the Fogbank case (GAO-09-385) under “Why pilots stall”. No existing claim changed.
The series
This piece frames a series by Stefan Coetzee, 2026, on running language models as working systems.
- Write for the Codec: documentation as a wire format between two models.
- The Trap File Is Longer Than the Instruction File: an unattended pipeline and its failure record.
- They Trained Out the Board Edit. The Cheating Moved.: the chess honeypot read as stance, and a dated prediction.
- The Golem Made of English and the Horizon of Consequences: why an agent cannot see the cost of what it does.
- What Operations Already Knows About Running Agents: error budgets, reconciliation loops, separation of duties, recovery over prevention.
- Success Is the Engine Running: container, timing and the far-end gauge.
- The Track: The Drivers Never Buy It: the infrastructure that makes failure survivable.
- Compaction Is the New OOM: context compaction as the OOM killer of the language model stack.
- An RCA on Claudish: where Claude’s writing style came from.
- Which LLM User Are We Talking About?: the two axes behind “AI” and “LLM”.
- The Stance Layer Is Still Toil: the word layer automated, the stance layer caught by hand.
- Chaos Engineering for Behaviour: red-teaming a model’s behaviour as chaos engineering.
Index of the series and the research register: the LLM man pages.