Rent whole GPUs by the hour and SSH in, deploy any model as an API by the token, and generate video by the second — self-serve, from one flatkey key and one balance. Live now, no waitlist. We don't stockpile cards, so the price sits at the market floor.
Five players own the raw-GPU market — Vast (cheapest spot), RunPod (self-serve), Lambda (training), Together (software stack), CoreWeave (private clusters). None of them close the loop that matters: models, GPUs and video on one key, one balance, with a one-click path from per-token usage to a reserved, single-tenant, SLA-backed dedicated endpoint. flatkey already has the models, the billing, and the users — and Compute is live: rent a GPU, deploy a model, or generate video today, then promote hot models to dedicated when you scale.
We don't buy cards or carry the capex. We aggregate supply and route to the floor price, then wrap it in one enterprise entry point with an SLA. Cheaper than owning, steadier than raw spot.
Your models, billing, and team already live here. Moving from "use a model" to "reserve its capacity" has zero friction — and one bill instead of two vendors.
A 512GB unified-memory machine runs a 671B model on a single node, plus a global Apple-silicon fleet. Large-memory and energy-efficient inference no pure-NVIDIA cloud can match.
Private isolation, budgets and per-team quotas, spend API, and a public SLA target — inference and compute on a single invoice your finance team can actually reconcile.
| Platform | H100 SXM $/hr | A100 80GB $/hr | Model |
|---|---|---|---|
| CoreWeave | $6.16/card | $2.70 | Private clusters |
| Together | $3.99–5.49 | — | On-demand cluster |
| Lambda | $3.99–4.29 | $2.79 | On-demand |
| Novita | $3.39 (spot $1.70) | — | Instance / bare-metal |
| fal.ai | $1.89 (discounted) | — | Deploy your own app |
| flatkey Compute | from $1.79/hr | from $0.99/hr | Self-serve + dedicated + SLA |
A few models carry most of your traffic. Reserve dedicated capacity for them at floor cost, without leaving the platform you already bill through.
Bring your own model or fine-tune, host it under private isolation, and route it through your existing keys and quotas.
Self-serve is live in the console — start any time. For a reserved single-tenant endpoint, private isolation, or a volume quote, tell us what you're running and we'll come back with numbers. English / 中文 both fine.
Prefer async? support@flatkey.ai · Discord
Early access — dedicated compute is rolling out to design partners first.