Mark Levy/Writing/The Word Warp Drive

An 8-bit long-range sensor display: a rising pixel curve across a starfield, a tiny shuttle near its start and the starship at its current point, a figure in white robes pointing along the curve toward an amber blank where it leaves the screen.

[ Tier 3 of 3 ] Astrometrics & Temporal Mechanics: how far, how fast

The Word Warp Drive

Tier 3: scaling laws, shuttlecraft, robots, the definition problem, and what a captain does with a future they can't read.

Q
ScanLong-range sensors, aft to forward: fourteen years of one trick getting bigger.
The future is the next token in this sequence, and nobody has those knobs.

Welcome to Tier 3, the horizon. You know the ship: next-word warp core, answer costume, action crew, and a captain with better orders. Now the big question: how far and how fast does this go? The answer is shape, not certainty: three growth multipliers, smaller ships inheriting flagship capability, embodied systems lagging words for data reasons, and a definition problem that makes experts disagree by decades.

Numbers here are approximate as of 2026 and sourced below. Demos show shapes, not forecasts. Final section focuses on what captains can do across uncertain timelines.

QI once showed a captain three timelines. He asked which was real. Wrong question. He could act in all three, and that was the test.

Section 1 of 5The Warp Scale

Three multipliers, one curve

Speed comes from three multiplying trends. Hardware improved about 1.4x/year over decades and is slowing. Training compute spending has grown near 4x/year since 2010. Algorithms often deliver similar performance on much less compute, roughly 3x/year efficiency gains. Multiply them and you get today's curve.

Scaling laws made that spending rational: error fell predictably as compute, data, and knobs rose (Deck 1, Section 2). For users, equivalent capability has trended toward ~10x cheaper per year (task variance high). Exponentials still hit walls—power, fabs, data, or capital. Data is the one already being routed around: the readable internet is finite, so labs increasingly train on synthetic data, text one model writes for the next to learn from. The question is where first.

The exponent: pick a rate, pick a horizonillustrative numbers, as of 2026
Exponential (log scale)
1x1,000x1 million1 billion1 trillion1 quadrillion
If it were linear, same first-year gain

Captain's log

Three trends multiply, and fixed capability has been getting ~10x cheaper yearly. Hardware is the slow leg; spend + algorithms carry more load. Plan for curve shape, not infinite extrapolation.

Back on Earth Scaling laws are smooth error-vs-compute lines; compute-optimal (Chinchilla) sets data-per-knob ratio. Training is measured in FLOPs and purchased as huge clusters, making power and chip supply key constraints. Synthetic data pushes the data wall back, but needs hard filtering—models fed their own unfiltered output degrade. The bitter lesson: general methods plus more compute often beat hand-crafted tricks.
QThis era's warp scale approaches ten and never quite arrives. Every exponential finds an asymptote. The interesting question is where it bends.

Section 2 of 5Shuttlecraft

Last year's flagship fits in a shuttle

Flagships set the ceiling, but most work does not need ceiling performance. As costs fall, yesterday's flagship capability fits smaller hulls. Four drivers: distillation, quantization, mixture of experts, and specialization.

Command rule: route each task to the smallest vessel that clears the bar, and reserve flagships for high-stakes judgment. On-device shuttles are instant/private, runabouts are cheap/fast, flagship handles irreversible missions.

The hangar: six tasks, four hullshand-built demo, illustrative prices
Pick a hull for every task, then check the manifest.
Captain's log

Route to the smallest ship that clears the bar. Capability trickles down quickly into smaller, cheaper hulls. Keep flagship capacity for judgment-heavy work.

Back on Earth Small language models, distillation, quantization, MoE, and on-device/edge models are the hull set. A router picks hull by task; fine-tuned specialists often beat larger generalists on narrow jobs.
QYou do not deploy a flagship for groceries. Distillation is flagship behavior in shuttle form—Soong needed decades, labs need weekends.

Section 3 of 5Exocomps

Words came first because words were on the internet

Moravec's paradox still bites: symbolic tasks that feel hard to humans can be easier for machines, while toddler-level perception and dexterity remain difficult. Evolution spent far longer on sensing and motor control than abstraction. LLMs learned words first because text data was abundant; there is no internet-scale touch dataset.

Robotics must generate data via simulation, teleoperation, and video. Then similar next-token machinery predicts motor actions (VLA models). It works, but on a slower clock because physical examples consume real time. Self-driving's path from demos to paid rides took about two decades.

The Moravec sort: which is harder for a machine?hand-built demo
Four pairs. In each, pick the one you think is harder for a machine today.
Captain's log

Bodies lag words mostly for data reasons. Perception/dexterity have weaker datasets and slower feedback loops. The same guessing framework is extending to actions, just slower.

Back on Earth VLA models pair camera input with motor output. Sim-to-real means simulation training that transfers to hardware. Teleoperation supplies demonstrations. Humanoids bet on reusable human-shaped data. Self-driving and warehouse robots are current revenue-scale embodiment.
QExocomps were built for conduits, not sonnets. Crawling through hardware was the hard part all along.

Section 4 of 5The Measure of a Machine

AGI is a definition wearing a date

The frontier is jagged: strong legal prose, weak letter counting; exam wins, weak freshness without context. Benchmarks saturate and get replaced. Even Turing-test claims triggered argument, not consensus.

Forecasts diverge because AGI lacks one definition. Every date implies a chosen bar. Some bars are already met, some contested, some clearly unmet with frozen knobs (Deck 1, Section 2). Better signals than headlines: task horizon and price trends.

Pick a definition: six bars, three verdictsas of 2026; the verdicts move every year
Verdict

Choose a definition. Each one carries its own date.

Six definitions, three verdicts. The date in any headline depends on which card the writer was holding.
Captain's log

Every AGI date hides a definition. Frontier is jagged, benchmarks move. Watch task horizon and cost; ask which definition card a forecast uses.

Back on Earth AGI, ASI, and transformative AI are different bars. Benchmarks track progress until saturation. Jagged frontier means superhuman in some tasks, weak in others. Task horizon tracks unattended job length at target reliability; forecasting platforms aggregate still-divergent estimates.
QMaddox demanded a sentience definition; Picard demanded consequences. The judge ruled on consequences. Your species will do the same with "general," and that is usually wiser.

Section 5 of 5All Good Things

You prepare the same way in every timeline

Three plausible futures dominate discussion. Plateau: progress bends and mostly gets cheaper/reliable. Steady climb: current doubling continues and routine work shifts to agents. Fast takeoff: systems improve systems and institutions lag. Reasonable people hold each; nobody knows which wins.

Stoic move: separate control from uncertainty. You can't set the curve; you can set habits that work in all timelines—define done, verify evidence, use small checked steps, hold human boundaries, know your domain, route to smallest sufficient model, and keep practicing.

Three timelines, one planhand-built demo

Known hazards on this deck
  • HAZARD 1Headline drift. Mistaking demos and press releases for trend measurements.
  • HAZARD 2Benchmark theatre. Saturated or contaminated tests, or scores without task/cost context.
  • HAZARD 3Forecast worship. Treating one date as truth instead of one definition-weighted scenario.
  • HAZARD 4Waiting. Delaying skill-building until certainty appears.
Captain's log

The plan survives the forecast. Controllables stay similar across timelines; delegation dial changes. Watch curves, keep learning.

Back on Earth AI-native workflows are Deck 4 done daily. Human 3.0 names augmentation that amplifies people. Continuous learning is non-delegable: model knobs are frozen, yours are not.
QIn the finale, the captain finally joins poker night. I told him the trial never ends. I never said you could not enjoy the game.

ManifestThe future vocabulary

What's in the cargo bay

Words that appear in future-of-AI arguments, with deck addresses:

Manifest: 14 items
  • Moore's lawHardware trend, now the slowest multiplier in Section 1.
  • scaling lawPredictable error curve versus compute, data, and knobs.
  • compute-optimalData-per-knob recipe (Chinchilla) for fixed-compute training.
  • algorithmic efficiencyBetter recipes delivering similar results on less compute.
  • synthetic dataModel-generated training text, used as readable human data runs short.
  • distillationBig model teaching small model (Section 2).
  • quantizationFewer bits per knob, large memory savings.
  • mixture of expertsLarge net, partial activation per token.
  • on-deviceLocal models: private, instant, lower capability ceiling.
  • Moravec's paradoxReasoning easier than perception/dexterity for machines.
  • VLA modelVision-language-action next-token control for bodies.
  • jagged frontierSuperhuman in some tasks, weak in others.
  • task horizonUnattended task length at target reliability.
  • AGI · ASI · transformative AIDifferent bars; forecast dates depend on chosen bar.
Captain's log

Every item here is a trend, hull, or definition. None changes the core's function. Deck 1's guessing game is still what's being scaled and debated.

TribunalQ

Final tribunal: five charges, three shield levels

Q
Order. I ran out of useful villains, so this tribunal's prosecutor is me. Future-watching is my specialty. Five charges, one per section. Right answers dismiss one; wrong answers drop shields and point to review. Lose all shields and I snap you back to the door. Win, and I may consider probationary Continuum candidacy. Control yourself.
The courtroomthe trial never ends

Mission completeWhat this buys you

Four commendations

Four rules hold across all timelines:

  • Read curves, not headlines. Multipliers, cost trends, and task horizon are sourced measurements; demos are bets paying off.
  • Route to the smallest sufficient ship. Capability flows downhill; reserve flagship use for judgment-heavy work.
  • Bodies move slower than words. Robotics progress is data-constrained and on a slower clock.
  • Prepare similarly in every timeline. Specify, verify, iterate small, keep human boundaries, know domain, keep learning.

That's the full ship: one guessing core, answer costume, action crew, and command habits riding an uncontrollable curve. You cannot steer the curve; you can steer the ship.

Turbolift

Last deck. No Deck 6. The trial never ends, but the tour does. Revisit Deck 1 for mechanism and Deck 4 for next-year execution.

Back to the ship Start again on Deck 1
Colophon

Most of the writing across all five decks was drafted by Claude Fable 5.1, then edited, cut, and fact-checked by Mark Levy, who owns every claim that survived.

A confession, since you made it to the end and have therefore earned one. The words you just read were largely produced by exactly the sort of machine they describe. It built you a starship, walked you through its own guts, wrote my dialogue, and then handed me the gavel so it could be judged in my voice. I have watched entire species crawl out of the ocean with less nerve than that. Your captain supplied the argument, the deletions, and the willingness to be wrong in public, which remains the part no model can hand you. Dismissed.
Inspiration and further reading

Section 1 trend numbers follow Epoch AI, plus scaling-law work by Kaplan et al. and Hoffmann et al. (Chinchilla); fixed-capability price decline references a16z's "LLMflation". Task-horizon doubling comes from METR. Section 3 references Moravec's paradox.

Numbers are approximate as of 2026 and will age; shape claims should age slower. Deck 1 covers core mechanics via John Mount's "A Simplified Mental Model of LLMs". Starship framing is affectionate homage to a certain 1987 series, with no affiliation. Demos are illustrative, not forecasts.