You decide what to run.
SpacePilot decides how and where.
This laptop. The box under the desk. A dock you rent for nine minutes. A managed API. One fleet — and every job goes to whichever machine is already holding the model warm.
No account. No API key. No port opened. Nothing to sign up for, because there is nothing to sign up to.
A model already sitting in memory answers now, for nothing. The same model on better silicon spends a minute arriving. Pick by hand and you are guessing at invisible state — what is loaded, where, right now. SpacePilot watches it, and that is the whole product.
One real job, from a real fleet. The GPU box is faster silicon and free to run, and it still loses — fourteen seconds of nothing happening is worse than zero. Sort by sticker price and you would have picked wrong.
That is the entire ranking. value_of_latency is a dial you own: at zero it buys the cheapest seconds available, turned up it stops making you wait. There is no second scoring pass you cannot see, and no sponsored tier. "Best" is ranked against constraints you set: cost, speed, privacy, reliability.
Mozilla's State of Open Source AI (July 2026, surveying 1,410 developers) found that 79% of developers adding AI use open models — but only 53% of their teams get them to production. The report's own conclusion: the gap is tooling and trust, not resources. That gap is what SpacePilot exists for.
One scheduler, any AI workload. Models are one kind of cargo, not the whole abstraction.
No tenant, no onboarding wizard, no trial. It asks permission, reads the machine, tells you what fits, stows one model, and then measures it — and the number at the end is the only one on the screen we produced ourselves. What follows is a real capture, on our machine.
Recorded from the running interface at 1280×720. Nothing in those frames is a mockup, and nothing in them is a number we did not take.
Most numbers in any registry are other people's arithmetic, and we would rather say so than let you find out later. Provenance sits at the same weight as the number it qualifies — never a tooltip, never a footnote, never grey. Your registry starts empty and fills with your own measurements.
We ran it. On this ship. On a date you can read. Two of thirty-four, and the count only goes up by working.
Their number, cited and dated. Sometimes a spec sheet, sometimes bytes summed from the Hub — honest arithmetic, still not a measurement.
Nobody has run it here. A rule applied to a parameter count, and the note names the rule. Most of the registry, and we are not hiding it.
Kokoro 82M through ONNX. 12.5 seconds of speech from one sentence in 2.7 seconds of wall clock, M1 Max, solo — nothing else running.
FLUX.2 klein-4B at 4-bit through MFLUX — the working set we measured, not a figure derived from parameter count.
And when a call fails, you get the failure. Not a plausible number wearing its clothes.
Machines you own find each other over your own tailnet and keep their models warm. Machines you rent are docks — a meter, an ID, a running clock. Managed APIs are a third lane, priced per token, first-class rather than a fallback. No port is opened. We never hold your identity, and we never see your work.
A real fleet — yours will have your names. Identity and transport belong to Tailscale and Cloudflare, deliberately. Two things we will never be good at, and two things we therefore refuse to store. That is why this page has no login button — not as a growth tactic, as an architecture.
The most-downloaded open model family passed two billion downloads this year. Weights are freight: anyone, anywhere, can carry the same models you do.
Export rules flipped four times in eighteen months. Committed GPU contracts rose ~40% in six months while spot stayed flat. Where a thing runs — and what the run costs right now — is the game.
More models, more kinds of compute, and more of the demand coming from agents. Whoever holds the live map owns the decision layer.
Probing, and the real ceiling rather than the advertised one. Compatibility across 34 variants. Runtime install that shows every package it would move before it moves one. Speech, locally, at 4.6× realtime. An MCP server, so an agent can fly this without ever opening a window — and a WebMCP proof-of-concept already running on Chrome’s origin trial.
Video and image generation on your own silicon. Scheduling across more than one ship at a time. You will not find any of it above, written as though it already worked. Resisting that is the product.