AI gateway pricing is easy to underestimate because the visible model rate is only one part of the bill.
The lowest advertised token price does not automatically produce the lowest operating cost. Buyers also need to account for gateway fees, subscription allowances, credit purchase charges, retries, fallback traffic, observability, quota controls, and the time required to reconcile usage across providers.
This guide gives procurement, engineering, and finance teams a practical way to compare those costs. It also explains how Flatkey's current plans fit teams that want one API key, one base URL, centralized usage visibility, and access to text, image, and video models.
Pricing note: Product prices and plan limits in this guide were checked on July 27, 2026. Pricing changes frequently. Confirm current terms on each vendor's official pricing page before making a purchase decision.
The Short Answer: Compare Total Gateway Economics
A useful AI gateway pricing comparison should answer seven questions:
- What do the underlying models cost for the workloads you actually run?
- Does the gateway add a markup, credit purchase fee, subscription, or usage-based platform charge?
- Is provider usage included in the plan, passed through, or billed separately through your own provider keys?
- What happens when the application retries a request or falls back to another model?
- Can you set quotas before an experiment, team, or environment overspends?
- Can engineering and finance reconcile the same request-level evidence?
- Does the plan match your traffic pattern, including short bursts and media generation?
The comparison should end with an effective cost per successful task, not merely cost per million tokens.
effective cost per successful task =
(model usage + gateway fees + retry/fallback cost + operating overhead)
/ accepted task outcomes
That formula works for a support answer, a document extraction, an image, a video clip, or any other product outcome.
The Four Common AI Gateway Pricing Models
AI gateways package costs in different ways. Before comparing prices, identify which model each vendor uses.
| Pricing model | What you pay | Best fit | Main comparison risk |
|---|---|---|---|
| Subscription with included usage | A fixed monthly plan with defined text or media allowances | Teams that want a predictable entry point and bundled access | Allowances, short-term caps, and overage behavior may matter more than the headline subscription |
| Prepaid credits with a purchase fee | Model usage plus a percentage or fixed fee when credits are purchased | Teams that value broad access without separate provider accounts | The funding fee compounds as usage grows |
| Pass-through model rates | Provider rates with no stated gateway markup, sometimes paired with platform or hosting charges | Teams that want transparent model costs | Observability, deployment, data transfer, or other platform costs may sit elsewhere |
| SaaS gateway subscription plus usage | A monthly platform fee with log, request, seat, or feature limits | Teams bringing their own provider accounts and buying control-plane software | Provider bills and gateway bills must be combined to see total cost |
There is no universally cheapest model. The best choice depends on traffic volume, request size, modalities, provider mix, operational maturity, and how much control the team needs.
A Current Market Snapshot
Official pricing pages illustrate why a category comparison cannot use one column labeled “price.”
| Gateway | Public pricing structure checked July 27, 2026 | What a buyer should verify |
|---|---|---|
| Flatkey | Monthly Go, Pro, and Max subscriptions include access to all listed models, defined text-model usage, and media credits; Enterprise uses custom terms | Allowance fit, short-term caps, media-credit consumption, and enterprise invoicing or routing terms |
| Cloudflare AI Gateway | Core gateway features are available without a gateway fee; using Cloudflare-managed provider access through Unified Billing adds a percentage fee on purchased credits | Whether you will use Unified Billing or your own provider keys, plus any related Workers or platform usage |
| Vercel AI Gateway | Model access is presented at provider list price with no markup; accounts receive monthly AI Gateway credit and can buy additional credits | Whether the surrounding Vercel plan, observability, deployment, or team needs add costs outside model usage |
| OpenRouter | Prepaid credits include a purchase fee; bring-your-own-key requests have a separate free allowance and then a platform fee | Credit funding frequency, BYOK request volume, provider routing, and payment-method costs |
| Portkey | A monthly production plan includes defined log volume, with higher-volume and enterprise options | Provider spend, log overages, retention, seats, and enterprise feature requirements |
| Helicone | A monthly Pro plan includes platform features and a usage allowance, with usage-based charges and custom enterprise terms | Provider spend, request volume, retention, seats, and which features require a higher plan |
These are not identical products. Some sell managed model access, some sell a control plane for your own provider accounts, and some combine both. A fair comparison must first decide whether the buyer wants one commercial relationship for model access or software layered over existing provider relationships.
Cost Layer 1: Underlying Model Usage
Start with the workload, not the vendor homepage.
For text models, estimate:
- input tokens;
- cached input tokens when the model supports them;
- output tokens;
- requests per user action;
- expected context growth;
- percentage of requests routed to premium models.
Image, audio, and video models may use different units such as generated images, seconds, minutes, resolution tiers, or jobs. Do not normalize these workloads into request counts alone.
A useful monthly model-cost estimate is:
monthly model cost =
Σ(requests × input units × input rate)
+ Σ(requests × output units × output rate)
+ media generation charges
If you have not yet measured real workloads, use three scenarios: normal, growth, and incident. The incident case should include longer prompts, repeated calls, and premium-model fallbacks.
For a deeper forecasting workflow, see AI API spend forecasting.
Cost Layer 2: Gateway Fees, Funding Fees, and Subscriptions
Next, translate the vendor's commercial model into the same monthly unit.
For a subscription model:
gateway cost = monthly subscription + overages + add-ons
For a credit purchase model:
gateway cost = model credits purchased × purchase-fee percentage
For a SaaS control plane:
gateway cost = platform plan + request/log overages + seats + retention or enterprise add-ons
For bring-your-own-key setups, do not treat the gateway plan as the entire bill. Add every provider invoice to the platform charge. Also include the operating cost of maintaining provider accounts, payment methods, limits, keys, and regional access.
Cost Layer 3: Retries and Fallback Routing
A request that reaches the user may contain several billable attempts.
For example:
- The primary model times out after processing input.
- The client retries the same model.
- The gateway falls back to a second model.
- The application accepts the final response.
The user sees one outcome, while the billing system may record three requests.
Ask each vendor whether request logs show the full chain: original request, provider response, retry, fallback, final model, units consumed, and cost. Then calculate:
retry amplification = total billable attempts / accepted outcomes
Even a modest increase in retry amplification can erase a model-rate advantage. This is why routing reliability and cost visibility belong in the same buying decision.
Cost Layer 4: Quotas and Spend Controls
Cost controls have economic value only if they work before the money is spent.
Evaluate whether the gateway lets you:
- create separate keys for products, teams, customers, or environments;
- set key-level or team-level quotas;
- see consumption against those quotas;
- restrict access to expensive models;
- identify sudden changes in traffic or model mix;
- review recharge and balance history;
- disable or rotate a key without changing every provider integration.
A provider rate that is 5% lower may not be cheaper if a shared key allows an uncontrolled test to consume a large production budget.
Use AI API quota limits to build the first set of team and environment guardrails.
Cost Layer 5: Finance and Reconciliation Work
Gateway pricing should also account for the monthly effort required to explain the bill.
Separate provider accounts can create multiple invoices, currencies, credit balances, payment methods, and export formats. Engineering may know which model served a request while finance sees only account-level totals.
A centralized gateway can reduce that work when its dashboard connects:
- API key or workload owner;
- model and route;
- input, output, cache, or media units;
- request status;
- retry and fallback activity;
- balance or recharge record;
- cost for a defined period.
If dashboard quality is a major buying criterion, use the unified AI billing dashboard evaluation guide to run a request-to-invoice test.
Flatkey Pricing: Which Plan Fits Which Workload?
Flatkey packages managed multi-model access into monthly plans. The current Flatkey pricing page lists Go, Pro, Max, and Enterprise options.
| Plan | Public monthly price | Included text-model usage | Included media allowance | Practical fit |
|---|---|---|---|---|
| Go | $10/month | Up to $45 of model usage per month | 300 media credits | Individual builders and light daily use |
| Pro | $30/month | Up to $90 of model usage per month | 1,200 media credits | Daily development and higher-frequency requests |
| Max | $100/month | Up to $300 of model usage per month | 5,000 media credits | Production workloads and heavier media use |
| Enterprise | Custom | Committed-volume and custom terms | Custom | Larger usage, procurement, invoicing, custom routing, or team controls |
The self-serve plans also publish short-term usage caps. Review those caps against your burst pattern, not only your monthly estimate. A team with a predictable monthly total but a concentrated launch-day spike may need a different plan from a team with steady background traffic.
Choose Go When
- one developer or a small prototype needs broad model access;
- usage is light and distributed;
- the team wants one key without opening several provider accounts;
- media generation is occasional.
Choose Pro When
- development traffic runs every day;
- prompt testing or automation creates more frequent calls;
- the team needs more room for image or video experiments;
- Go's short-term caps are too tight for the expected workflow.
Choose Max When
- the application is moving into production;
- several workloads share the gateway;
- media usage is material;
- the team wants a larger included allowance before discussing custom terms.
Choose Enterprise When
- usage is large enough for committed-volume economics;
- procurement needs invoicing or negotiated terms;
- the team needs custom routing or organizational controls;
- a signed commercial arrangement matters more than self-serve simplicity.
Flatkey's model is most relevant for buyers who prefer centralized access and billing over maintaining separate provider accounts. Teams that already have negotiated provider contracts should compare those direct rates plus gateway software costs against a managed-access plan.
A Buyer Scorecard for AI Gateway Pricing
Use this scorecard with the same workload sample for every vendor.
| Decision area | Evidence to request | Pass condition |
|---|---|---|
| Model cost | Rate card and request-level usage | Input, output, cache, and media units are visible |
| Gateway charge | Subscription, funding fee, markup, or overage schedule | Every non-model charge can be expressed monthly |
| Included usage | Allowance, expiry, cap, and overage terms | The expected workload fits both monthly and short-term limits |
| Retry economics | Attempt chain and final outcome | Duplicate and fallback costs can be measured |
| Quotas | Key, team, and environment controls | A test workload can be capped before overspend |
| Reconciliation | Dashboard, export, recharge, and invoice evidence | Engineering and finance totals match for one period |
| Portability | Base URL, SDK compatibility, and key migration steps | A model or route can change without an application rewrite |
| Support | Incident path and service commitments | Response expectations match production risk |
Score each row from 0 to 2:
- 0: unavailable or unclear;
- 1: available with manual work or plan restrictions;
- 2: available, testable, and included in the intended plan.
Do not select a gateway solely because it wins the model-rate row. A strong buying decision should also pass retry, quota, reconciliation, and portability tests.
Run a Seven-Day Pricing Proof Before You Commit
The cleanest comparison is a short proof using representative traffic.
- Select one text workflow and one media workflow if media matters.
- Send the same traffic distribution through each shortlisted setup.
- Record accepted outcomes, not just requests.
- Measure input, output, cache, media, retry, and fallback units.
- Add every platform, funding, subscription, and provider charge.
- Reconcile the dashboard total with exports and balances.
- Compare effective cost per successful task and operator time.
Also document what would change at 3× and 10× volume. Some pricing models are attractive at evaluation scale but create a different cost curve in production.
Frequently Asked Questions
Is an AI gateway more expensive than calling providers directly?
Not necessarily. A gateway can add a subscription, funding fee, or platform charge, but it may also reduce provider-account overhead, integration work, uncontrolled usage, and failed-request cost. Compare total operating cost for the same accepted outcomes.
What is the most important AI gateway pricing metric?
Use effective cost per successful task. It combines model usage, gateway charges, retries, fallbacks, and operating overhead in a unit the product team understands.
Should I use prepaid gateway credits or bring my own provider keys?
Prepaid managed access is simpler when you want one commercial relationship and broad model access. Bring your own keys can be attractive when you already have negotiated provider terms, but it preserves separate accounts, limits, invoices, and key management.
How should I compare text, image, and video costs?
Keep each workload's native billing units, then normalize the result to an accepted business outcome. Request counts alone are not sufficient.
Why do quota controls belong in a pricing comparison?
Because an avoidable overspend is part of total cost. Quotas, key separation, and usage visibility help contain experiments, compromised keys, unexpected loops, and sudden model-mix changes.
Compare Flatkey Plans Against Your Real Workload
Start with your expected text usage, media usage, burst pattern, and need for procurement controls. Then review Flatkey pricing to compare Go, Pro, Max, and Enterprise against that workload.
If the priority is one API key, one OpenAI-compatible base URL, managed model access, and centralized cost controls, Flatkey gives the buying team a concrete plan structure to evaluate without assembling separate provider accounts first.



