Mark Levy/Writing/The Word Warp Drive

[ Tier 3 of 3 ] Astrometrics & Temporal Mechanics: how far, how fast
Tier 3: scaling laws, shuttlecraft, robots, the definition problem, and what a captain does with a future they can't read.
New here? This final tier stands mostly on its own, but Deck 1 (mechanism) and Deck 4 (command) sharpen the frame.
Welcome to Tier 3, the horizon. You know the ship: next-word warp core, answer costume, action crew, and a captain with better orders. Now the big question: how far and how fast does this go? The answer is shape, not certainty: three growth multipliers, smaller ships inheriting flagship capability, embodied systems lagging words for data reasons, and a definition problem that makes experts disagree by decades.
Numbers here are approximate as of 2026 and sourced below. Demos show shapes, not forecasts. Final section focuses on what captains can do across uncertain timelines.
Section 1 of 5The Warp Scale
Speed comes from three multiplying trends. Hardware improved about 1.4x/year over decades and is slowing. Training compute spending has grown near 4x/year since 2010. Algorithms often deliver similar performance on much less compute, roughly 3x/year efficiency gains. Multiply them and you get today's curve.
Scaling laws made that spending rational: error fell predictably as compute, data, and knobs rose (Deck 1, Section 2). For users, equivalent capability has trended toward ~10x cheaper per year (task variance high). Exponentials still hit walls—power, fabs, data, or capital. Data is the one already being routed around: the readable internet is finite, so labs increasingly train on synthetic data, text one model writes for the next to learn from. The question is where first.
Three trends multiply, and fixed capability has been getting ~10x cheaper yearly. Hardware is the slow leg; spend + algorithms carry more load. Plan for curve shape, not infinite extrapolation.
Section 2 of 5Shuttlecraft
Flagships set the ceiling, but most work does not need ceiling performance. As costs fall, yesterday's flagship capability fits smaller hulls. Four drivers: distillation, quantization, mixture of experts, and specialization.
Command rule: route each task to the smallest vessel that clears the bar, and reserve flagships for high-stakes judgment. On-device shuttles are instant/private, runabouts are cheap/fast, flagship handles irreversible missions.
Route to the smallest ship that clears the bar. Capability trickles down quickly into smaller, cheaper hulls. Keep flagship capacity for judgment-heavy work.
Section 3 of 5Exocomps
Moravec's paradox still bites: symbolic tasks that feel hard to humans can be easier for machines, while toddler-level perception and dexterity remain difficult. Evolution spent far longer on sensing and motor control than abstraction. LLMs learned words first because text data was abundant; there is no internet-scale touch dataset.
Robotics must generate data via simulation, teleoperation, and video. Then similar next-token machinery predicts motor actions (VLA models). It works, but on a slower clock because physical examples consume real time. Self-driving's path from demos to paid rides took about two decades.
Bodies lag words mostly for data reasons. Perception/dexterity have weaker datasets and slower feedback loops. The same guessing framework is extending to actions, just slower.
Section 4 of 5The Measure of a Machine
The frontier is jagged: strong legal prose, weak letter counting; exam wins, weak freshness without context. Benchmarks saturate and get replaced. Even Turing-test claims triggered argument, not consensus.
Forecasts diverge because AGI lacks one definition. Every date implies a chosen bar. Some bars are already met, some contested, some clearly unmet with frozen knobs (Deck 1, Section 2). Better signals than headlines: task horizon and price trends.
Choose a definition. Each one carries its own date.
Every AGI date hides a definition. Frontier is jagged, benchmarks move. Watch task horizon and cost; ask which definition card a forecast uses.
Section 5 of 5All Good Things
Three plausible futures dominate discussion. Plateau: progress bends and mostly gets cheaper/reliable. Steady climb: current doubling continues and routine work shifts to agents. Fast takeoff: systems improve systems and institutions lag. Reasonable people hold each; nobody knows which wins.
Stoic move: separate control from uncertainty. You can't set the curve; you can set habits that work in all timelines—define done, verify evidence, use small checked steps, hold human boundaries, know your domain, route to smallest sufficient model, and keep practicing.
The plan survives the forecast. Controllables stay similar across timelines; delegation dial changes. Watch curves, keep learning.
ManifestThe future vocabulary
Words that appear in future-of-AI arguments, with deck addresses:
Every item here is a trend, hull, or definition. None changes the core's function. Deck 1's guessing game is still what's being scaled and debated.
TribunalQ
Mission completeWhat this buys you
Four rules hold across all timelines:
That's the full ship: one guessing core, answer costume, action crew, and command habits riding an uncontrollable curve. You cannot steer the curve; you can steer the ship.
Last deck. No Deck 6. The trial never ends, but the tour does. Revisit Deck 1 for mechanism and Deck 4 for next-year execution.
Back to the ship Start again on Deck 1Most of the writing across all five decks was drafted by Claude Fable 5.1, then edited, cut, and fact-checked by Mark Levy, who owns every claim that survived.
Section 1 trend numbers follow Epoch AI, plus scaling-law work by Kaplan et al. and Hoffmann et al. (Chinchilla); fixed-capability price decline references a16z's "LLMflation". Task-horizon doubling comes from METR. Section 3 references Moravec's paradox.
Numbers are approximate as of 2026 and will age; shape claims should age slower. Deck 1 covers core mechanics via John Mount's "A Simplified Mental Model of LLMs". Starship framing is affectionate homage to a certain 1987 series, with no affiliation. Demos are illustrative, not forecasts.