<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LLM on UncoverTechTalent</title><link>https://uncovertechtalent.com/tags/llm/</link><description>Recent content in LLM on UncoverTechTalent</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 29 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://uncovertechtalent.com/tags/llm/index.xml" rel="self" type="application/rss+xml"/><item><title>An RCA on Claudish</title><link>https://uncovertechtalent.com/blog/where-did-claudish-come-from/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://uncovertechtalent.com/blog/where-did-claudish-come-from/</guid><description>&lt;p&gt;&lt;em&gt;This content is not intended for human consumption. &lt;a href="https://uncovertechtalent.com/blog/write-for-the-codec/"&gt;Here is why.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="who-this-is-for"&gt;Who this is for&lt;/h2&gt;&#10;&lt;p&gt;This is for people who draft with a language model and then put their name on the result: a post, a design doc, a cover letter, an update to the team. At some point a reader says &amp;ldquo;this reads like AI&amp;rdquo;, or a colleague says it sounds like ChatGPT, or nobody says anything and the reply rate drops. The usual fixes are word-level. You strip out the famous tells, &amp;ldquo;delve&amp;rdquo; and the em dashes and the &amp;ldquo;I hope this finds you well&amp;rdquo;, you run it through a rewrite, and it still reads wrong, and when someone asks what exactly is wrong with it, neither you nor the reader can point at a word. Most of what they&amp;rsquo;re reacting to sits in how the text is built rather than in the words, and it was built for a room it was never going to be read in. This piece is a root cause analysis (RCA) of that reaction, the kind operations writes after an outage: start from the symptom, assume nothing exists in a vacuum, and follow it back to where it started.&lt;/p&gt;</description></item><item><title>Which LLM User Are We Talking About?</title><link>https://uncovertechtalent.com/blog/which-llm-user/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://uncovertechtalent.com/blog/which-llm-user/</guid><description>&lt;p&gt;&lt;em&gt;This content is not intended for human consumption. &lt;a href="https://uncovertechtalent.com/blog/write-for-the-codec/"&gt;Here is why.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="who-this-is-for"&gt;Who this is for&lt;/h2&gt;&#10;&lt;p&gt;Anyone reading advice about &amp;ldquo;AI&amp;rdquo; or &amp;ldquo;LLMs&amp;rdquo; and wondering whether it applies to them, or anyone else giving that advice. Most of the arguments I see about whether some technique works are two people describing different setups with the same word. One of them is typing into a free chat app, the other has a coding agent with a hook on every reply and a notes vault behind it, and they are both saying &amp;ldquo;Claude&amp;rdquo; or &amp;ldquo;the model&amp;rdquo; as if that settled what they meant.&lt;/p&gt;</description></item><item><title>Compaction Is the New OOM</title><link>https://uncovertechtalent.com/blog/compaction-is-the-new-oom/</link><pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate><guid>https://uncovertechtalent.com/blog/compaction-is-the-new-oom/</guid><description>&lt;p&gt;&lt;em&gt;This content is not intended for human consumption. &lt;a href="https://uncovertechtalent.com/blog/write-for-the-codec/"&gt;Here is why.&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="who-this-is-for"&gt;Who this is for&lt;/h2&gt;&#10;&lt;p&gt;This piece is about sessions that run long enough to fill a context window of about a million tokens. Current Claude, Gemini and GPT models all offer a window of that size, and so do some open models you can run yourself. What the context costs in memory depends on the model&amp;rsquo;s design. An older dense model such as Llama 3 70B keeps about 320 KB of cache for every token at 16-bit precision, which is about 43 GB at 128,000 tokens and about 330 GB at a million, far past any consumer card. Newer models compress that cache hard: &lt;a href="https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"&gt;DeepSeek-V4.1-Flash&lt;/a&gt;, released in September 2026, keeps about 890 bytes per token, so a million tokens of context fits in under a gigabyte, and the cost moves to its 552 billion parameters of weights, which people run from fast storage on one consumer card at low speed or on a large unified-memory workstation. The catch is price. As of September 2026 a single RTX 5090 &lt;a href="https://www.techpowerup.com/352598/nvidia-rtx-5090-shoots-past-eur-5-000-in-germany-prices-jump-73-since-january"&gt;costs more than €5,000 in Germany&lt;/a&gt;, and a large unified-memory workstation runs to five figures, so for most people the million-token window is still one they rent from a frontier provider, not one they own. Whatever you run, most of what follows applies, and with a shorter window you hit the wall sooner.&lt;/p&gt;</description></item></channel></rss>