October 9, 2026 · 8 min read · finops.qa

AI Gateway for Cost Control: LiteLLM vs Portkey vs Kong vs Cloudflare vs Bifrost (2026)

AI gateway cost control compared: budgets, per-key and team limits, caching, routing and spend attribution in LiteLLM, Portkey, Kong, Cloudflare and Bifrost.

AI Gateway for Cost Control: LiteLLM vs Portkey vs Kong vs Cloudflare vs Bifrost (2026)

Quick answer

For AI gateway cost control in 2026, LiteLLM and Bifrost give you the deepest budget hierarchies in open source, Cloudflare AI Gateway is the fastest managed option now that it has dollar spend limits, and Portkey and Kong are strongest when you are buying an Enterprise platform anyway. The deciding factor is rarely features on a slide. It is which cost controls sit behind an Enterprise licence, and whether the gateway can attribute spend at the level your chargeback model needs.

We wrote in our AI cost audit guide that LLM spend should be instrumented at the gateway. This post is the follow-up question: which gateway? Every capability below was checked against vendor documentation and pricing pages on 9 October 2026. Where a doc did not say whether a feature is open source or paid, we say so rather than guess.

Why put cost control in a gateway at all?

Provider dashboards tell you what you spent. They do not tell you which team, feature or customer spent it, and they cannot stop a request before it happens. A gateway can do four things nothing else in the stack does:

  • Attribute every call to a key, user, team or tag before it reaches the provider.
  • Enforce a budget or rate limit in real time, not at month end.
  • Cache repeated requests so you do not pay for the same answer twice.
  • Route work to cheaper models or providers when quality allows.

The catch is that the gateway only sees traffic that goes through it. Developers calling a provider SDK directly with their own key are invisible, which is why our Claude Code and Cursor cost guide recommends a gateway or OpenTelemetry for anyone reaching models through Bedrock, Vertex or Foundry.

How do the five gateways compare on cost control?

CapabilityLiteLLMPortkeyKong AI GatewayCloudflare AI GatewayBifrost
DeploymentSelf-hosted proxy (MIT core); Enterprise licenceManaged SaaS; open-source gateway; EnterpriseSelf-hosted or KonnectManaged (Cloudflare edge)Self-hosted (Apache 2.0); Enterprise
Dollar budgetsGlobal, team, team member, user, key, end customerCost and token limits per provider integration (Enterprise and select Pro)Cost-based limits via AI Rate Limiting Advanced (Enterprise)Spend limits by model, provider or custom metadata (open beta, all plans)Customer, team, virtual key and provider config
Budget resetAny duration (for example 30d)None, weekly or monthlyWindow per limitFixed or sliding windowsMinutes to yearly, optional calendar alignment
Rate limitsRPM, TPM, parallel requests per key, user, team, customerGranular rate limits (Enterprise)Token or cost per consumer, group, model, providerRequests per windowRequests and tokens per virtual key and provider
CachingExact (Redis, S3, GCS) and semantic (Redis, Valkey, Qdrant)Simple cache from Production plan; semantic on EnterpriseSemantic cache plugin (Enterprise)Exact matchExact and semantic (Redis/Valkey, Weaviate, Qdrant, Pinecone)
Cost-aware routingCost-based routing strategy, fallbacksConditional routing and fallbacksLowest-usage balancing by tokens or cost (Enterprise)Dynamic routes, Auto Router, fallback after spend limitBudget-aware routing; providers excluded when over limit
Spend attributionKey, user, team, end user; tags and spend reports are EnterpriseMetadata on all plans; automatic user attribution on EnterpriseToken and cost metrics via logs, OpenTelemetry, KonnectCustom metadata in logs and analytics; custom costsVirtual key, team, customer hierarchy

Sources: LiteLLM budgets, LiteLLM spend tracking, Portkey budget limits, Portkey pricing, Kong AI Rate Limiting Advanced, Kong AI Proxy Advanced, Cloudflare AI Gateway features, Cloudflare spend limits changelog, Bifrost budgets and limits.

When is LiteLLM the right choice?

LiteLLM is the default for teams that want to self-host and control the data path. Its budget model is the broadest we found in an open-source core: a global proxy budget, team budgets, per-member budgets inside a team, user budgets, key budgets and end-customer budgets, each with an optional budget_duration reset. When a key crosses its max_budget, requests fail. Routing includes a cost-based-routing strategy that picks the lowest-cost healthy deployment.

Two things to plan for. Budgets need a Postgres database; without one, the global budget check is skipped. And several chargeback features are marked Enterprise in the docs: tag-based spend tracking, the spend report endpoint that groups spend by team or customer, and per-model budgets on a key. If your chargeback runs on tags such as feature or cost centre, price the Enterprise licence in from day one.

When is Portkey the right choice?

Portkey is a managed gateway with an open-source core, and since May 2026 it belongs to Palo Alto Networks, which closed the acquisition and plans to make it the gateway layer of its Prisma AIRS platform. That is good news if you want security and gateway from one vendor, and a reason to ask about roadmap and contract terms if you do not.

On cost, Portkey budget limits support both a dollar cap and a token cap per provider integration, with weekly or monthly resets and email alerts at a threshold. The docs say the feature is available on the Enterprise plan and to select Pro customers. A few details matter for FinOps: limits only count requests made after they are set, a limit cannot be edited once set (you duplicate the provider instead), and spend only counts for models Portkey has pricing for.

When is Kong AI Gateway the right choice?

Kong AI Gateway makes sense if Kong already runs your API traffic. Its AI Rate Limiting Advanced plugin supports token limits and, since version 3.8, cost-based limits computed from input and output token prices you configure. Policies can match a consumer, a consumer group, a model or a provider, so you can give one team a cost ceiling on one model. AI Proxy Advanced adds a lowest-usage balancing algorithm that routes on token counts or cost, plus semantic and priority routing.

The trade-off: the rate limiting, advanced proxy and semantic cache plugins are all documented as part of the AI Gateway Enterprise offering. Cost-based limits also depend on prices you enter yourself, so someone owns keeping them current.

When is Cloudflare AI Gateway the right choice?

Cloudflare AI Gateway is the quickest to stand up. Core features, including analytics, caching and rate limiting, are free according to its pricing page. In June 2026 Cloudflare added spend limits: dollar budgets scoped by model, provider or custom metadata such as user or team, with fixed or sliding windows, working with both Unified Billing and your own provider keys. When a budget runs out, requests are blocked by default, or a dynamic route can send them to a cheaper fallback model.

The limits to know: caching is exact-match, not semantic. Spend limits only work for models with known pricing. And spend limits launched in open beta, so treat them as a guardrail you test, not a contract.

When is Bifrost the right choice?

Bifrost, from Maxim AI, is an Apache 2.0 gateway written in Go. Its governance model is hierarchical: budgets on customers, teams, virtual keys and provider configs, and a request only proceeds if every level has budget left. The same cost is deducted at each level, which maps neatly onto a chargeback hierarchy. If one provider under a virtual key exceeds its limit, that provider is dropped from routing while others keep serving. Resets run from a minute to a year, with an option to align budgets to calendar periods in UTC.

Customer scoping is an Enterprise feature, and the docs point to Enterprise for high availability across nodes. Performance figures for Bifrost come mostly from Maxim’s own benchmarks, so test with your traffic.

Other gateways, including TrueFoundry and MuleSoft, also appear in 2026 shortlists. We left them out of the table because we could not verify their cost-control features to the same depth from public docs.

How do you choose the right gateway for chargeback?

Start from your allocation model, not the feature list:

  1. Write down the unit of chargeback. Team, product feature, customer or cost centre. If it is a tag rather than a key, check whether tag attribution is open source or paid in your shortlist.
  2. Map the hierarchy. If budgets must roll up (key to team to business unit), favour gateways that enforce at every level, such as LiteLLM or Bifrost.
  3. Decide hard stop or graceful fallback. Blocking at the limit protects the budget; routing to a cheaper model protects the workflow. Cloudflare and Bifrost both document fallback behaviour.
  4. Check pricing coverage. Every gateway in this list only counts spend for models it can price. Custom or fine-tuned models need custom prices, or they show up as zero.
  5. Reconcile monthly. Compare gateway spend to the provider invoice. Our showback vs chargeback guide explains why to start with showback until the gap is small.

Prompt caching at the provider and caching at the gateway are different levers, and they stack. Our LLM API cost optimization guide covers the provider side.

Where to start

If you are not running a gateway, pick the one that matches where you already operate: Cloudflare if you live on its edge, Kong if it runs your APIs, LiteLLM or Bifrost if you want to self-host. Put your three most expensive applications behind it first, give each a budget, and reconcile one month of gateway data against the invoice before you charge anyone back.

If you want help with that, our fixed-scope gateway selection and chargeback implementation engagement, delivered through our AI and GPU cost governance service, scores your shortlist against your providers and allocation model, configures budgets and attribution in the gateway you choose, and tests that limits actually fire. If your AWS side is next, see our AWS FinOps Agent readiness checklist. Talk to us before the next invoice does the talking.

Frequently Asked Questions

What is an AI gateway and why does it matter for cost?

An AI gateway is a proxy that sits between your applications and model providers. Every call passes through it, so it is the one place where you can attach an owner to each request, enforce a budget before the request is sent, cache repeated answers and route cheap work to cheaper models. Without one, provider invoices give you a total and very little attribution.

Which AI gateway has the best budget controls?

On documented features, LiteLLM and Bifrost have the deepest budget hierarchies in their open-source builds, with budgets on keys, users or teams and configurable resets. Cloudflare AI Gateway added dollar spend limits in June 2026, in open beta on all plans. Portkey budget limits are an Enterprise feature, and Kong's cost-based limits need AI Gateway Enterprise.

Is LiteLLM free for cost tracking?

The core of LiteLLM is MIT-licensed and tracks spend per key, user and team, with budgets and rate limits on each. Some chargeback features are Enterprise-only according to the docs: tag-based spend tracking, the spend report endpoint used to charge teams and customers, and per-model budgets on a key. Budgets also need a Postgres database to be enforced.

Does Cloudflare AI Gateway support semantic caching?

No, not according to the current docs. Cloudflare AI Gateway caching serves identical requests from Cloudflare's cache, which is exact matching. If you need semantic caching, LiteLLM supports Redis, Valkey and Qdrant semantic caches, Bifrost supports several vector stores, Kong has an Enterprise semantic cache plugin, and Portkey lists semantic caching on its Enterprise plan.

Should we build chargeback on gateway data or on the provider invoice?

Both. Use the gateway as the attribution layer, because it knows which key, team and feature made each call, and use the provider invoice as the source of truth for the total. Reconcile the two monthly. The gap between them is attribution loss, and it tells you which traffic is bypassing the gateway or which models it is pricing incorrectly.

Get Your FinOps Defect Score

Book a free 30-minute cloud cost review. We will identify your top three FinOps gaps and give you a preliminary Defect Score - no pitch, no obligation.

Every engagement is scoped by our principal architect, Adrian Vale: 20+ years in production engineering, 40+ professional certifications. Meet Adrian

Talk to an Expert