You decide what to run.
SpacePilot decides how and where.

This laptop. The box under the desk. A dock you rent for nine minutes. A managed API. One fleet — and every job goes to whichever machine is already holding the model warm.

$pip install spacepilot
then open 127.0.0.1

No account. No API key. No port opened. Nothing to sign up for, because there is nothing to sign up to.

wan 2.1 t2v one 5 s clip chosen your laptop warm · 0 s · $0.00 your GPU box cold +14 s · $0.00 a rented GPU cold +92 s · $0.31 — bills
One job, three ways to run it. Gold is the route the pilot picked; every number here is real and repeated in the cards below.
scroll
The whole product, in three cards

The right machine is the one
already holding the model.

A model already sitting in memory answers now, for nothing. The same model on better silicon spends a minute arriving. Pick by hand and you are guessing at invisible state — what is loaded, where, right now. SpacePilot watches it, and that is the whole product.

your laptopmodel warm
Apple M1 Max
24.96 GB usable
0sto first frame
$0.00loaded already
chosen
your GPU boxcold
RTX 5090
32 GB VRAM
14scold start
$0.00free, but cold
faster silicon, still loses
a rented cloud GPUmetered
L40S · 48 GB
us-east-1
92scold start
$0.31and it bills
last resort

One real job, from a real fleet. The GPU box is faster silicon and free to run, and it still loses — fourteen seconds of nothing happening is worse than zero. Sort by sticker price and you would have picked wrong.

score = price + (cold ? load_seconds × value_of_latency : 0)

That is the entire ranking. value_of_latency is a dial you own: at zero it buys the cheapest seconds available, turned up it stops making you wait. There is no second scoring pass you cannot see, and no sponsored tier. "Best" is ranked against constraints you set: cost, speed, privacy, reliability.

The gap, measured

Open models win adoption.
Production is another story.

Mozilla's State of Open Source AI (July 2026, surveying 1,410 developers) found that 79% of developers adding AI use open models — but only 53% of their teams get them to production. The report's own conclusion: the gap is tooling and trust, not resources. That gap is what SpacePilot exists for.

ADOPT REACH PRODUCTION 79% 53% open 71% 63% closed the gap: tooling and trust not resources
With company size, closed climbs 54% → 73%. Open barely moves: 53% → 57%. Scale rules out a resources explanation.
Mozilla, State of Open Source AI (stateofopensource.ai) · Jul 2026 · n=1,410
serve a 70B fine-tune a LoRA embed 40,000 docs transcribe an archive generate 500 images run an agent sandbox

One scheduler, any AI workload. Models are one kind of cargo, not the whole abstraction.

Fifteen seconds, start to measured

This is the whole
first run.

No tenant, no onboarding wizard, no trial. It asks permission, reads the machine, tells you what fits, stows one model, and then measures it — and the number at the end is the only one on the screen we produced ourselves. What follows is a real capture, on our machine.

127.0.0.1 — first run, recorded end to end
SpacePilot first run: consent, hardware probe showing 24.96 GB usable, the model compatibility matrix, a download, a model coming aboard, and a first measured run at 4.6 times realtime.
01Ask before reading anything
0224.96 GB, not the 32 on the box
03What fits, before you download
04One model stowed
054.6× realtime, measured here

Recorded from the running interface at 1280×720. Nothing in those frames is a mockup, and nothing in them is a number we did not take.

The uncomfortable slide

Every number tells you
how we know it.

2of 34 flown — on our fleet, today

Most numbers in any registry are other people's arithmetic, and we would rather say so than let you find out later. Provenance sits at the same weight as the number it qualifies — never a tooltip, never a footnote, never grey. Your registry starts empty and fills with your own measurements.

flown

We ran it. On this ship. On a date you can read. Two of thirty-four, and the count only goes up by working.

on paper

Their number, cited and dated. Sometimes a spec sheet, sometimes bytes summed from the Hub — honest arithmetic, still not a measurement.

unflown

Nobody has run it here. A rule applied to a parameter count, and the note names the rule. Most of the registry, and we are not hiding it.

Flown · speech 4.6× realtime

Kokoro 82M through ONNX. 12.5 seconds of speech from one sentence in 2.7 seconds of wall clock, M1 Max, solo — nothing else running.

Flown · image 10.52 GB

FLUX.2 klein-4B at 4-bit through MFLUX — the working set we measured, not a figure derived from parameter count.

And when a call fails, you get the failure. Not a plausible number wearing its clothes.

The fleet

Ships, docks, providers.
One fleet, three lanes.

Machines you own find each other over your own tailnet and keep their models warm. Machines you rent are docks — a meter, an ID, a running clock. Managed APIs are a third lane, priced per token, first-class rather than a fallback. No port is opened. We never hold your identity, and we never see your work.

laptop
M-series · 32 GB
2 models warm
gpu-tower
RTX 5090 · 32 GB
1 model warm
spare-mac
M2 · 16 GB
asleep · not gone — work waits
rented-gpu
L40S · 48 GB · rented
not running · $0.00/hr
providers
managed APIs · per-token
metered lanes

A real fleet — yours will have your names. Identity and transport belong to Tailscale and Cloudflare, deliberately. Two things we will never be good at, and two things we therefore refuse to store. That is why this page has no login button — not as a growth tactic, as an architecture.

The decade this is built for

Weights cross every border.
Chips need visas.

Cargo moves freely

The most-downloaded open model family passed two billion downloads this year. Weights are freight: anyone, anywhere, can carry the same models you do.

Compute has borders

Export rules flipped four times in eighteen months. Committed GPU contracts rose ~40% in six months while spot stayed flat. Where a thing runs — and what the run costs right now — is the game.

More models, more kinds of compute, and more of the demand coming from agents. Whoever holds the live map owns the decision layer.

Where it is today

Early, and saying so
where you can see it.

Works now

Probing, and the real ceiling rather than the advertised one. Compatibility across 34 variants. Runtime install that shows every package it would move before it moves one. Speech, locally, at 4.6× realtime. An MCP server, so an agent can fly this without ever opening a window — and a WebMCP proof-of-concept already running on Chrome’s origin trial.

Not yet

Video and image generation on your own silicon. Scheduling across more than one ship at a time. You will not find any of it above, written as though it already worked. Resisting that is the product.

One command

Find out what your own
machines can actually do.

$pip install spacepilot
or get the launch note