Mark Levy/Writing/The Word Machine

A great stone hall lit by torches. A towering golem stands in the middle, its body made of countless tiny brass dials, holding a scroll with a blank where one word should be.

[ Floor 1 of 3 ] The monster itself

The Word Machine

How a large language model actually works, one guessing game at a time.

System
PromptParis is the capital of ___
CompletionFrance 92%
Welcome, crawler. This is the whole monster. Everything else on this floor is loot.

Every impressive thing a large language model does, from answering questions and drafting emails to writing code and passing exams, is one act, repeated: given some text, predict a plausible next word. That's it. The chat window, the personality, the apparent memory: all of it is built by running that single trick in a loop.

Two words you already use every day have exact meanings here. What you type is the prompt, and the full block of text the machine actually reads, your prompt plus the rules and chat history the app quietly adds, is the context. Every room below touches one or both, and each room ends with a note tying the mechanism back to the vocabulary of daily AI work.

One useful way to picture it: an un-deletion machine. Take real sentences, delete words, and train a circuit to restore them. Do that across roughly everything ever written, and the circuit gets eerily good at continuing text it has never seen.

This floor has six rooms. Every demo is a hand-built toy with illustrative numbers; the point is the mechanism, not the measurements. You don't need to know how to build an engine to drive well. You do need to know it isn't a horse.

Room 1 of 6The Cartographer's Table

Words become numbers

The machine never sees letters. Each word is converted to a long list of numbers, its coordinates in a huge space (real models use thousands of dimensions; this map flattens it to two). On day one those coordinates are random noise. Training slowly drags words that get used the same way to the same neighborhood.

Hover or tap a word to see its nearest neighbors, then press Day one to see the same words before training.

The cartographer's map: 2 of 12,288 dimensions shownhand-built toy
Hover a word…
Loot

Meaning is location. "Similar words" just means "nearby numbers." No definition of cat is stored anywhere, only where cat sits relative to everything else.

Back in town The map's real unit is the token, a word-piece (pharmacist arrives as pharma + cist; short common words are one piece). When a service bills per token, rates a model's speed in tokens per second, or caps a chat "in tokens," it is counting these pieces, the machine's raw unit.

Room 2 of 6The Golem

A trillion knobs, no moving parts

Inside, the machine is a circuit: numbers flow in, get multiplied and added according to adjustable knobs (the "parameters"; frontier models now run from hundreds of billions into the trillions), and scores flow out, one score per possible next word. No memory. No loops. Same input, same knobs, same output, every time. A golem, in the old sense: a body that does exactly what its inscription says, and nothing else.

Here's a three-knob golem trying to finish "The cat sat on the ___." Fresh from the workshop, its knobs are wrong; it likes moon. Tune the knobs until mat clears 70%, or press auto-train and watch a trainer do it.

Inside the golem: 3 of ~1,000,000,000,000 knobshand-built toy

"The cat sat on the moon"

GOLEM: MISCALIBRATED

Loot

You just trained a language model by hand. Real training does exactly this, nudging knobs so the right word scores higher, except with trillions of knobs, nudged a hair at a time, once per example, across trillions of words. No single knob means anything. The arrangement is the knowledge.

Back in town The model you pick from a dropdown (Claude, GPT, Gemini) is one specific, frozen arrangement of these knobs. A "model update" means the vendor re-trained and shipped a new arrangement. Nothing you do in a chat ever moves a knob.

Room 3 of 6The Sphinx's Riddles

School is a fill-in-the-blank quiz

Where do the training examples come from? Free, from any text: delete a word, and the original text itself is the answer key. No humans writing quizzes; the internet is the quiz. The sphinx at this door asks a few of them. After each guess, the machine's own candidate list is revealed.

The sphinx's riddles: 3 of ~1,000,000,000,000hand-built toy
Loot

The output is a bet, not a fact. Riddle 3 is the important one: when many continuations are common, no answer gets a high score. The machine's "confidence" measures typicality in text, never truth.

Back in town The text pile behind this quiz is the training data you hear about in the news, and the reason a model "knows" things at all: the answer key was, roughly, the internet.

Room 4 of 6The Scrying Pool

Attention: how "it" finds its owner

Predicting the next word means knowing which earlier words matter right now. That's the job of attention heads. Think of them as scrying eyes: each position in the text gets to look back at every other position and decide how much to care.

In the sentence below, what does "it" refer to? Flip the last word and watch where the eyes go.

Scrying from position 9hand-built toy, illustrative weights
Loot

Attention is a budget. Long-range coherence, like remembering your instructions from 40 lines ago, is attention doing its job. And tasks that need too many simultaneous look-backs (sort 300 unique items, track every letter in a long word) are exactly where the machine predictably falls apart.

Back in town Your rules, whether a system prompt, custom instructions, or a CLAUDE.md, are not settings inside the machine. They are text placed at the top of the context on every pass, held in force purely by these look-backs. That's why a very long chat can drift away from its rules: line one competes for attention with everything written after it.

Room 5 of 6The Endless Stair

Paragraphs are one word, repeated

The circuit outputs exactly one word. So how do you get an essay? Plumbing outside the machine: predict a word, staple it to the text, feed the whole thing back in, repeat. Every step of this stair is one word, and to take the next step the machine climbs the entire stair again from the bottom. Press the button and watch: the flash across every chip is the machine re-reading the whole transcript from scratch, every single pass.

The stair: one step per wordhand-built toy

Candidate next words

Press Predict next word to see the machine's bets.

Loot

The machine has no memory. The context is the memory. Your chat is just an ever-longer input re-read on every pass. That's also why long chats degrade: the re-read gets harder, and past a certain length the oldest text isn't read at all.

Back in town This loop is what the products call a chat. Everything fed in on one pass (rules, transcript, your latest prompt) is the context, and the ceiling on how much fits is the context window. When a long chat overflows it, the oldest text simply stops being fed in; nothing was "forgotten," it just isn't in the input anymore. And the checkbox above is the temperature dial: always-top-word is temperature zero, and sampling further down the bars is turning it up.

Room 6 of 6The Mimic's Den

Fabrication is the machine working as designed

Now the payoff. The machine's only job is plausible continuation. Nothing in the circuit checks truth; there is no wire for it. Every dungeon has a mimic: a monster that looks exactly like a treasure chest until you open it. This room has two chests. Ask about a real book and a book that doesn't exist, and the machine produces the same confident shape either way.

Two chestscanned outputs, for the shape

> Summarize the novel Moby-Dick.

Herman Melville's 1851 novel follows Ishmael aboard the whaler Pequod, whose captain, Ahab, is obsessed with the white whale that took his leg. The hunt destroys the ship; Ishmael alone survives to tell it.

MIMIC

> Summarize the novel The Lighthouse at Pelican Reef.

Marion Voss's 1974 novel follows a keeper's daughter who discovers her father has been falsifying the light's logbook for decades. A storm forces the truth ashore; the ending is often read as an allegory of institutional decay.

Every word above is fabricated: author, year, plot. There is no such book. It scores high because it's the shape a book summary takes.

You'll hear this failure called "hallucination," as if the machine mis-perceived something. It perceived nothing. Fabrication is the honest word: constructing plausible text is all the machine ever does. It's just that the most plausible continuation is usually true, and when it isn't, nothing inside notices.

Known traps on this floor
  • TRAP 1Plausible-wrong pits. Questions whose wrong answer is more common in text than the right one: trick riddles, subtly altered classics, "obvious" arithmetic. The typical answer wins the score, confidently.
  • TRAP 2Attention overload. Counting letters, sorting long unique lists, deep multi-step puzzles: anything needing more simultaneous look-backs than the attention budget covers.
  • TRAP 3No scratch pad. The circuit is one-way; it cannot pause, iterate, or check its own work. When a chat product does seem to iterate (running code, searching the web, reading your files) that's bolted-on tools plumbing around the machine, added precisely because the machine can't. What a tool finds gets pasted into the context so the next words anchor to it.
  • TRAP 4Frozen knobs. Training set the knobs and stopped. Anything that happened after (new facts, your last correction) reaches the machine only if it rides along in the context. A product's memory feature is exactly that plumbing: notes saved outside the machine, pasted back into the context of later chats.
Loot

Confidence carries no information. The tone is constant whether the content is Melville or Marion Voss. Judge the output the way you'd judge an unsigned memo: by checking it.

InventoryProduct words, machine parts

What you're carrying now

The vocabulary of daily AI work, the words on the box, all have exact addresses in the rooms you just walked:

Inventory: 11 items
  • promptThe text you hand over on this pass. To the machine it's just input, not privileged over anything else in the block.
  • contextEverything read on one pass: rules + chat history + your latest prompt. Re-read in full for every single word predicted (Room 5).
  • context windowThe hard ceiling on how much text fits in one pass. Text beyond it isn't "forgotten"; it's simply not in the input.
  • chatThe Room 5 loop with a growing transcript. The transcript is the only continuity there is.
  • memoryNotes the product saves outside the machine and pastes into future context. Talking to it never changes a knob.
  • rules · system promptText the product places at the top of the context, every pass. Enforced only by attention (Room 4), which is why long chats drift.
  • tokenThe word-piece from Room 1, the machine's raw unit. Limits, bills, and speed are all counted in these.
  • modelOne frozen arrangement of the Room 2 knobs. A version bump means re-trained knobs, shipped whole.
  • temperatureHow the loop picks from the candidate bars: zero always takes the top word; higher samples further down the list.
  • hallucinationRoom 6's fabrication, in cases where the most plausible continuation happens not to be true. The machine can't tell the difference.
  • tools · web searchPlumbing around the machine that fetches text (search results, file contents, code output) and pastes it into the context so the continuation anchors to something real.
Loot

Notice what every item has in common: except for model and token, they're all just ways of putting text into the context, or plumbing that does it for you. The machine only ever does the one thing from the top of this page.

Floor clearedWhat this buys you

Four skills unlocked

You now hold roughly the mental model a good driver has of a combustion engine: not enough to build one, plenty to avoid ruining one. Four rules fall straight out of the six rooms:

  • Feed it the truth and it shines. Summarize, rewrite, translate, draft, critique: transformations of text you supply are the machine's home turf, because the facts ride in with your input.
  • As a sole source of facts, it's a bettor, not a witness. Verify anything checkable: names, numbers, citations, anything you'd act on.
  • Ignore the tone. Fluency and confidence are properties of the training data, not of the answer.
  • Play to the circuit. Ask for steps and it does better; you're moving the iteration into the transcript, the only scratch pad it has. Expect stumbles on counting, long exact lists, and anything after its training date.

Not a mind, not magic: an un-deletion machine with trillions of well-tuned knobs, run in a loop. Used with that in mind, it's one of the best tools ever built for working with words.

System

Stairs down. Floor 2 is behind the glass: how your question becomes the start of a completion, what "inference" means, how thinking and effort work, and how images ride the same loop. Same monster, more plumbing.

Descend to Floor 2 Back to the map
Inspiration and further reading

The mental model here follows John Mount's "A Simplified Mental Model of LLMs" (Win-Vector): the un-deletion framing, knobs, look-backs, and the case for saying fabrication.

The teaching style, poking the simulation instead of reading the bullet list, follows Laurentiu Gabriel's "How I Use LLMs to Learn". The dungeon framing owes a debt to every LitRPG that ever announced a new achievement in a box. All demos on this page are illustrative toys, not real model outputs.