Core readsSYSTEM: You are a helpful assistant… HUMAN: why is the sky blue? ASSISTANT:← autocomplete starts here
Same core as Deck 1. It has not been refit. This deck tours the ship around it, and Moriarty is waiting at the bottom.
Deck 1 established the core: given some text, predict a plausible next word. Trillions of frozen knobs, attention look-backs, run in a loop, with the context as its only memory. If any of that sounds new, start there.
Deck 2 answers the questions a new officer asks next. How does a question ever start an autocomplete? What is "inference"? What are the models doing when they "think," and what does an "effort" setting actually buy? And how does a text-completion core look at a screenshot?
Five sections. As before, every demo is hand-built with illustrative numbers; the point is the mechanism, not the measurements.
Section 1 of 5The Deerstalker
How a question starts a completion
The core doesn't answer questions. It continues text. So the product never hands it your question alone. It assembles a transcript, a little screenplay in which a helpful assistant is about to speak, and asks the core to continue that. Your message is one block in the stack; the last line is the trigger.
Press the button and watch the deerstalker go on: this is what the core actually receives.
The deerstalker: what the core actually receiveshand-built demo
WHAT YOU TYPEDHow do I politely decline a meeting?
The holodeck is empty. The core has received nothing yet.
One thing is still missing: why would a raw autocompleter continue that script helpfully? Out of big training (Deck 1, Section 3), it wouldn't. It would just continue the pattern. So models get a second, smaller round of training on example conversations, written and rated by people, until "what a helpful assistant says next" becomes the likeliest continuation. Same question, before and after:
Same question, two graduatescanned outputs, for the shape
Captain's log
Chat is a costume. The core never stopped being an autocompleter. The product dresses your question in a transcript where an assistant speaks next, and tuning made the helpful continuation the likely one. Even the reply's ending is a prediction: the end-of-turn token is an ordinary candidate on the list, learned like every other word.
Back on Earth
When an API shows messages with roles (system, user, assistant) those are the transcript's blocks, written as a list. The second training round is the fine-tuning / RLHF from model announcements. RLHF is Reinforcement Learning from Human Feedback: people rate the model's answers, and the ratings steer the knob-nudging. And the silent stop token is why a reply just… ends, without anyone cutting it off.
QData in a deerstalker on the holodeck is still Data. He continues Holmes's script because that's the scene he was handed. And his comedy lessons from a passing rogue in "The Outrageous Okona" were finishing school: examples of what comes next at a party, drilled until it almost lands. (Your species' 1966 chatbot was named ELIZA, after a flower girl who got the same treatment. I didn't plan that. I'm delighted anyway.)
Section 2 of 5Utopia Planitia
Inference: running the core
The core has two lives, and the vocabulary keeps them straight. Training is the rare one: knobs moving, trillions of examples, months in a datacenter. It happens before you ever meet the model. That's the shipyard at Utopia Planitia, where the ship was built, once. Inference is the other life, and it's the only one you ever touch: the frozen core, run forward. Text in, one word out. That's deep space, where the ship gets flown. Every message you send triggers it.
The word is borrowed from statistics: the core infers the next word from the input. Flip the switch and compare the two lives.
Shipyard or deep spacehand-built demo
KNOBS: FROZEN
TEXT IN ───▶ [ a trillion-plus knobs ] ───▶ NEXT WORD◀─── error flows back, every knob gets a nudge
OUTPUT · ONE PASS PER WORD
SPEC SHEET · INFERENCE
Captain's log
If you're chatting, it's inference. Nothing you type moves a knob. When a conversation seems to "learn," that's the growing context doing the work (Deck 1, Section 5), not the core changing. Learning-the-core-way happened once, in the shipyard, before launch.
Back on Earth
Words trickling onto your screen is streaming: one forward pass per token, each shipped the moment it's picked. Speed quoted in tokens per second, a status page reporting "elevated inference latency," headlines about GPU shortages: all statements about this section, the cost of running the frozen core, once per token, for everyone at once.
QLal, Data's daughter, kept learning after she left the lab. It did not end well for her. The models you meet ship the way this ship does: built once at Utopia Planitia, frozen at launch, and nobody is flipping the switch in deep space.
Section 3 of 5The Ready Room
Thinking: stepping into the ready room first
Here's a hard limit from Deck 1: the circuit is one-way, and every token gets the same fixed slice of computation, one pass. A hard problem doesn't get a bigger slice; it gets a plausible-shaped guess.
The workaround is delightfully low-tech: before answering, let the core write to itself. Scratch text, generated by the same next-word loop, goes into the context, where the final answer's attention look-backs can anchor to worked steps instead of leaping to a guess. Like a captain who steps into the ready room and works it through on a padd before giving the order: no new officer, just more ink.
The scratch paddcanned outputs, for the shape
THE PROBLEMA pharmacy has 3 shelves with 24 bottles each. A shipment adds 30 bottles, then 18 are dispensed. How many bottles now?
Captain's log
Thinking is more turns of the same crank. The core buys computation by spending tokens, and the scratch padd lives in the context where attention can use it. This is also why "think step by step" helps in a plain prompt: you're inviting the scratch work into the transcript, the only workspace the core has.
Back on Earth
The folded-away "thinking…" block in Claude is exactly this scratch padd: thinking or extended thinking on the box, and billed as tokens like everything else. A reasoning model is one that went through extra finishing school (Section 1) on worked solutions, so the scratch text it writes is actually useful.
QPicard does not think faster than Riker. He goes into the little room first, and the bridge waits. That's the whole trick: same captain, more passes, and the working out stays on the padd where the order can point to it.
Section 4 of 5Warp Factor
Effort: pick a warp factor
If thinking is spending tokens, someone has to set the budget. That's all an effort setting is: a cap on how much scratch text the core may write before answering. A warp factor, in other words. It doesn't make the ship a better ship (same knobs at every position); it buys more passes, at a price.
Choose a warp factor and watch what it costs, and what it's worth, on an easy question versus a genuinely hard one.
Warp factorhand-built demo, illustrative numbers
scratch
time
cost
TASK A — "What's the capital of France?"
TASK B — "Three staff, six scheduling rules — build Friday's shift plan."
Captain's log
Effort is a budget, not a brain transplant. Match it to the problem: an easy question at high effort buys the same answer at thirty times the price, and a hard one at low effort gets Section 3's first button, a confident guess.
Back on Earth
The effort levels in Claude Code, the extended-thinking toggle in the Claude app, a thinking budget set in the API, a "think longer" mode: all the same dial, sold under different labels. A cap on scratch tokens.
QYou don't go to warp to cross the room. Warp nine for "What's the capital of France?" gets you the same answer, faster than you can read it, at a price Geordi will want to discuss with you later, in Engineering, at length.
Section 5 of 5The Viewscreen
Pictures on the same loop
The loop carries tokens: numbers standing for word-pieces. Nothing in the machinery cares that they started as words. So to let the core see, a second, smaller model (a vision encoder) chops the image into a grid of patches and turns each patch into the same kind of number-list a word gets, a point in the meaning-space from Deck 1, Section 1. Those patch-tokens are spliced into the context right next to your words, and from there it's business as usual: attention reaches into the picture exactly the way it reaches back at words. On screen.
The viewscreen: a 4 × 4 grid of patcheshand-built demo, illustrative weights
The picture is still a picture. The core can't read it yet.
Captain's log
The core doesn't see pictures. It reads them. Once an image is patches in the context, it's text-like all the way down. That's why it can describe a chart beautifully and still misread the fine print: a patch, like a token, is a coarse unit, and detail smaller than the unit is a bet, not an observation.
Back on Earth
Dropping a screenshot into a chat runs this slicer; that's what multimodal and vision mean on a model's spec sheet. The patches sit in the context window like words, so a big image spends more than a thousand tokens, a couple of pages of text, per picture. Audio gets in by the same trick: chop, embed, splice.
QPicard says "Magnify" and the viewscreen hands him more pixels. The core says it and gets the same sixteen patches, larger. Anything smaller than a patch is a guess wearing a magnifying glass.
ManifestDeck 2 vocabulary
What's in the cargo bay
Deck 2's product words, each with its exact address in the sections above:
Manifest: 12 items
inferenceRunning the frozen core: one pass, one token (Section 2). Every message you send triggers it. Training is the other, rarer life.
streamingWords appearing one by one is not a display effect. It's the loop itself, each token shipped as soon as its pass finishes.
roles · system/user/assistantThe transcript blocks from Section 1's deerstalker. An API request is that stack, written out as a list.
stop tokenA token meaning "the assistant's turn is over": an ordinary vocabulary entry that wins the candidate list like any word, because every training conversation ended with it. Predicting it is how a reply ends.
fine-tuning · RLHFFinishing school (Section 1): a second training round on example conversations. RLHF, Reinforcement Learning from Human Feedback, is the part where people rate answers and the ratings nudge the knobs toward answering instead of pattern-continuing.
thinking · extended thinkingScratch tokens generated before the visible answer (Section 3), usually folded away in the interface. Same core, pointed at a scratch padd first.
reasoning modelA model finishing-schooled on worked solutions, so its scratch text is worth the tokens.
effort · thinking budgetSection 4's warp factor: how many scratch tokens are allowed. Buys passes, not intelligence.
multimodal · visionThe viewscreen (Section 5): images chopped into patches, embedded into the same space as word-tokens, spliced into the context.
embeddingThe coordinate-list a token, or an image patch, becomes: a point on Deck 1, Section 1's chart. "Similar meaning" literally means "nearby numbers."
vector database · RAGA filing cabinet of embeddings, searched by nearness. RAG embeds your question, fetches the nearest-sitting documents, and pastes them into the context: tools plumbing feeding the same core.
latency · tokens/secHow fast inference cranks for you today. "The model feels slow" is a statement about the datacenter's load, not about the knobs.
Captain's log
Deck 1's punchline still holds. None of this replaced the core. Inference runs it, thinking runs it longer, effort chooses how much longer, and vision feeds it different freight. It's next-word prediction all the way down.
TribunalProfessor Moriarty
Q's tribunal: five charges, three shield levels
Q
Order. Today's prosecutor is on loan from the holodeck: Professor Moriarty, who has memorized every line ever written and has never seen a play. He became self-aware halfway through one and would very much like to be let out; I've told him a conviction might help. Five charges. Each right answer dismisses one. Each wrong answer drops your shields and names the section to revisit; lose the shields and I snap my fingers and you're back at the door. The turbolift isn't locked. I'm simply watching.
The courtroomthe script is the evidence
Deck clearedWhat this buys you
Four more commendations
Four more rules fall straight out of these five sections:
Give it the script, not just the question. The reply continues a transcript, so the rules, facts, and examples you put in the context are the script. A vague scene gets a vague next line.
Spend thinking where it pays. Hard, multi-step, checkable problems earn a high effort setting. For lookups and rewrites, extra effort mostly buys latency and cost.
"It learned!" means the context, not the knobs. Correcting it improves this chat only. The model that greets your next chat is the same frozen arrangement; carry anything worth keeping forward yourself.
Images are read, not seen. Coarse patches are great for "what is this chart saying," risky for fine print, tiny numbers, and counting. Treat what it "saw" like what it wrote: check it.
That's the whole deck: a completion core in a costume, run frozen, sometimes allowed to step into the ready room first, fed pictures chopped into words. Take the turbolift back up to Deck 1 whenever the fundamentals need a refresher.
Turbolift
Deck 3. Deck 3 is where the core gets a crew: the loop that lets it act on the world, the standing orders that tell it who to be, the ship's systems it can call on, away teams it can send, safety protocols at every console, and the briefing that decides what it gets to read. Same core. Now with a ship, and a narrator who has been waiting for this deck.
The mental model continues John Mount's "A Simplified Mental Model of LLMs" (Win-Vector): the core's mechanics are Deck 1's territory; this deck extends the same non-magical framing to the ship around it.
The teaching style, running the simulation instead of reading the bullet list, follows Laurentiu Gabriel's "How I Use LLMs to Learn". The starship framing is an affectionate homage to a certain 1987 television series, with no affiliation. All demos on this page are illustrative, not real model outputs.