Best open source AI gateways and LLM routers

7 minUpdated:
Best open source AI gateways and LLM routers

LiteLLM is the most common open source AI gateway, exposing many providers behind one OpenAI-compatible API with keys, budgets and fallbacks. Portkey’s gateway is a lightweight TypeScript alternative. Teams already on Kong or Envoy can add AI plugins instead of running a separate proxy.

What is an AI gateway?

An AI gateway is a proxy between your applications and LLM providers. Apps send requests to one endpoint in one format; the gateway translates, routes and records them.

It centralises concerns that otherwise get copied into every service: API keys, retries, provider fallbacks, rate limits, spend tracking and logging. It also makes switching models a config change instead of a code change.

An LLM router is a narrower idea: choosing which model handles a request, based on rules, cost or a classifier. Many gateways include simple routing, while dedicated routers experiment with learned routing.

Think of it as the same role an API gateway plays for microservices, adapted to the quirks of model providers: streaming responses, token-based billing and fast-changing APIs.

Which open source AI gateways are worth considering?

OpenRouter itself is a hosted service, not open source, but its model of one key and one API for many models is what most self-hosted gateways replicate.

ProjectLanguageSelf-hostingStrengthsBest forTrade-off
LiteLLMPythonYes, proxy server or SDKWide provider coverage, virtual keys, budgets, fallbacksTeams standardising on an OpenAI-compatible APISome enterprise features sit outside the open source tier; check the licence file
Portkey gatewayTypeScriptYes, Node or edge runtimesSmall footprint, routing configs, retriesLow-latency proxy close to the appRicher observability lives in the hosted product
Envoy AI GatewayGo, EnvoyYes, KubernetesBuilds on Envoy Gateway traffic managementPlatform teams already using EnvoyKubernetes-centric, heavier to operate
Kong AI pluginsLua, KongYesAI proxy plugins inside an existing API gatewayOrganisations already running KongAdvanced AI plugins may require paid tiers
RouteLLMPythonYesLearned routing between a strong and a weak modelResearch into cost-aware routingA framework, not a full gateway

Why is LiteLLM the default choice?

LiteLLM does two things. As a Python SDK it lets code call many providers through one completion interface. As a proxy server it exposes an OpenAI-compatible endpoint, so any client or tool that speaks that format can use Anthropic, Bedrock, Vertex, Azure, local Ollama models and others.

The proxy adds virtual keys per team or project, spend limits, fallbacks when a provider errors, and logging callbacks into observability tools. For many companies it is the fastest way to give every internal app governed model access.

Run it with a database for keys and spend tracking, and put it behind your normal authentication and TLS setup.

Its breadth is also its main maintenance burden. With so many providers supported, check the release notes for the providers you rely on and test them after each upgrade.

When is a lighter gateway better?

If you only need retries, fallbacks and a unified format, a small TypeScript gateway such as Portkey’s can run at the edge or as a sidecar with little overhead. It suits JavaScript-heavy stacks and latency-sensitive paths.

If your platform team already manages Envoy or Kong, adding AI features there keeps one control plane for all traffic, one set of auth rules and one place to monitor.

A lighter gateway also means fewer features to misconfigure. If you do not need per-team budgets, the simpler option is easier to reason about when something goes wrong.

How to choose an AI gateway

  • List the providers and local models you need today and those you expect next year.
  • Decide whether you need multi-tenant features: per-team keys, budgets, usage reports.
  • Check streaming, tool calling and structured output support for each provider you use, not only plain chat.
  • Measure added latency on your own traffic, including streaming time to first token.
  • Confirm how logs are stored, since prompts may contain customer data.
  • Read the licence and note which features are in the open source edition versus paid tiers.

Where AI gateways cause trouble

  • Lowest-common-denominator APIs that hide provider-specific features such as prompt caching or extended reasoning settings.
  • Silent fallbacks to a weaker model that change output quality without anyone noticing.
  • A single gateway instance becoming a new single point of failure for every AI feature.
  • Logging full prompts and responses into a system with weaker access controls than your main database.
  • Version drift: provider APIs change quickly, so an outdated gateway breaks new features.

Do you need a gateway at all?

A single app calling one provider does not need one. The provider SDK plus a retry wrapper is simpler and exposes every feature.

A gateway earns its keep once you have several apps, several teams or several providers, or once finance asks who spent what. At that point, central keys and budgets pay back the extra moving part.

RepoLoot’s catalog lists gateway and proxy projects with notes on difficulty and what you can build on them, which helps when you want to extend one rather than just deploy it.

If you are unsure, start without one but keep all model calls behind a single internal module. Swapping that module for a gateway client later is then a small, contained change.

What features should an AI gateway have?

Gateways vary widely, and the marketing pages blur the differences. These are the capabilities that tend to matter once more than one team depends on the gateway.

FeatureWhy it mattersWhat to test
OpenAI-compatible endpointExisting clients and tools work unchangedChat, streaming, tool calls, embeddings
Virtual keysRevoke one team without rotating provider keysCreate, limit and revoke a key
Budgets and rate limitsStops runaway spend and abuseHit a limit and check the error
Fallbacks and retriesSurvives provider outagesForce a provider error mid-stream
CachingCuts cost on repeated promptsConfirm cache keys include model and parameters
Logging hooksFeeds observability and auditsCheck redaction of sensitive fields

How should you deploy a self-hosted gateway?

Run at least two instances behind a load balancer so a restart does not take every AI feature down. Keep state, such as keys and spend records, in a managed database rather than on the container.

Place the gateway close to your applications, in the same region or cluster, to keep the extra hop short. Put it behind your normal authentication so only internal services and approved users can reach it.

Pin versions and upgrade deliberately. Test streaming and tool calling for each provider after every upgrade, because those paths break more often than plain chat.

How do gateways help control LLM costs?

Central logging shows who spends what, which is usually the first surprise. Budgets per key turn that visibility into limits, and rule-based routing sends simple tasks, such as classification, to cheaper models.

Response caching helps for repeated identical requests, such as the same system prompt with the same question. It does little for open-ended chat, so measure hit rates before counting on savings.

Which gateway fits which team?

Whichever you choose, keep application code talking to a standard API format. That keeps the gateway replaceable if your needs change.

  • Solo developer or single app: skip the gateway and use the provider SDK directly.
  • Startup with a few services and two or three providers: LiteLLM proxy with a small Postgres database.
  • Latency-sensitive JavaScript stack or edge deployment: a lightweight TypeScript gateway such as Portkey’s.
  • Platform team running Kubernetes and Envoy: Envoy AI Gateway, so AI traffic follows existing policies.
  • Enterprise already on Kong: Kong’s AI plugins, reviewing which ones need a paid tier.

Frequently asked questions

Is LiteLLM free to use?
The core SDK and proxy are open source and can be self-hosted at no licence cost. Some enterprise features, such as advanced single sign-on or certain admin capabilities, are offered under separate terms. Check the repository’s licence file and documentation for the current split.
Is there an open source alternative to OpenRouter?
OpenRouter is a hosted marketplace, so there is no exact clone, but LiteLLM proxy gives you the same pattern on your own servers: one OpenAI-compatible endpoint in front of many providers. You still need your own accounts and keys with each provider.
Does an AI gateway add latency?
Some, because every request passes through an extra hop and some processing. For most chat and agent workloads the overhead is small compared with model generation time. Measure it on your own infrastructure, especially time to first token when streaming.
What is the difference between an AI gateway and an LLM router?
A gateway is infrastructure: unified API, keys, limits, logging and fallbacks. A router is a decision layer that picks a model per request, for example sending easy prompts to a cheaper model. Many gateways include rule-based routing; learned routing is a separate, more experimental area.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides