NEW · flatkey Compute · live now

All your compute,
one key.

Rent whole GPUs by the hour and SSH in, deploy any model as an API by the token, and generate video by the second — self-serve, from one flatkey key and one balance. Live now, no waitlist. We don't stockpile cards, so the price sits at the market floor.

Open Compute — it's live → Talk to sales
RTX 4090 from $0.72/hrSSH & JupyterModels by the tokenVideo by the secondOne key · one balance
The gap

The compute market is crowded. The loop isn't.

Five players own the raw-GPU market — Vast (cheapest spot), RunPod (self-serve), Lambda (training), Together (software stack), CoreWeave (private clusters). None of them close the loop that matters: models, GPUs and video on one key, one balance, with a one-click path from per-token usage to a reserved, single-tenant, SLA-backed dedicated endpoint. flatkey already has the models, the billing, and the users — and Compute is live: rent a GPU, deploy a model, or generate video today, then promote hot models to dedicated when you scale.

How it works

Rent, deploy, or generate — self-serve

Rent a GPUWhole card by the hour, SSH & Jupyter in
Deploy a modelAny model as an API, by the token
Generate videoBy the second, whitelabel
One balanceAll billed from your flatkey wallet
Scale to dedicatedPromote hot models to reserved SLA endpoints
Why flatkey

Not another GPU cloud

01

Asset-light — the OpenRouter of compute

We don't buy cards or carry the capex. We aggregate supply and route to the floor price, then wrap it in one enterprise entry point with an SLA. Cheaper than owning, steadier than raw spot.

02

Inference ↔ compute, closed loop

Your models, billing, and team already live here. Moving from "use a model" to "reserve its capacity" has zero friction — and one bill instead of two vendors.

03

Apple Silicon large-memory fleet

A 512GB unified-memory machine runs a 671B model on a single node, plus a global Apple-silicon fleet. Large-memory and energy-efficient inference no pure-NVIDIA cloud can match.

04

Enterprise-grade, unified billing

Private isolation, budgets and per-team quotas, spend API, and a public SLA target — inference and compute on a single invoice your finance team can actually reconcile.

Pricing anchor

Priced at the market floor

PlatformH100 SXM $/hrA100 80GB $/hrModel
CoreWeave$6.16/card$2.70Private clusters
Together$3.99–5.49On-demand cluster
Lambda$3.99–4.29$2.79On-demand
Novita$3.39 (spot $1.70)Instance / bare-metal
fal.ai$1.89 (discounted)Deploy your own app
flatkey Computefrom $1.79/hrfrom $0.99/hrSelf-serve + dedicated + SLA
Self-serve today: rent an RTX 4090 from $0.72/hr with SSH, or run models by the token and video by the second — all from one balance. Backend supply is sourced by price across aggregated marketplaces, so a dedicated H100 endpoint still lands below Lambda / CoreWeave / Novita on-demand, while adding an enterprise SLA, private isolation, and unified billing that raw spot can't. Margin comes from routing and metering, not from carrying idle hardware. Anchor prices reflect public rates on 2026-07-17; self-serve rates indicative and shown live in the console.
Who it's for

Built for teams running inference at scale

Agent & AI-app companies

A few models carry most of your traffic. Reserve dedicated capacity for them at floor cost, without leaving the platform you already bill through.

Private deployment & compliance

Bring your own model or fine-tune, host it under private isolation, and route it through your existing keys and quotas.

Need dedicated capacity or private deployment?

Self-serve is live in the console — start any time. For a reserved single-tenant endpoint, private isolation, or a volume quote, tell us what you're running and we'll come back with numbers. English / 中文 both fine.

Prefer async? support@flatkey.ai · Discord
Early access — dedicated compute is rolling out to design partners first.