Skip to content

Gateway API

Gateway

One OpenAI-compatible endpoint in front of 5 model providers. The answer comes back in the shape you already parse; the routing decision and the price come back with it, in headers.

Chat completions is a single POST endpoint. It authenticates the caller, scores the request for complexity, picks a model, relays to that provider, prices the tokens the provider reports, and writes a ledger row and a trace row before it answers. What you get back is an OpenAI-shaped completion plus a set of x-lobstack-* headers describing every one of those decisions.

There is nothing to install. If your code already talks to the OpenAI Chat Completions API, point its base URL at Lobstack and change the key.


Endpoints

EndpointWhat it does
POST /chat/completionsChat, routed and receipted. Chat Completions
POST /embeddingsOpenAI embedding models, metered on input tokens. Embeddings
POST /searchWeb search results, priced per search. Web search
GET /modelsThe catalog, chat and embedding models. Models & providers

Paths are relative to https://www.lobstack.ai/api/gateway/v1.


One request

curl https://www.lobstack.ai/api/gateway/v1/chat/completions \
  -H "Authorization: Bearer $LOBSTACK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{ "role": "user", "content": "What time zone is Lisbon in during summer?" }]
  }'
The host is www.lobstack.ai, and the www mattersThe bare apex redirects, and RFC 9110 requires a client to drop the Authorization header when a redirect crosses hosts — curl, requests, httpx and therefore the OpenAI SDK, Go, Java and PowerShell all do. A base URL on the apex turns a valid key into a 401 that reads like an invalid one, so copy the base URL above rather than retyping the domain. Quickstart has the longer version.

What comes back

The body is an OpenAI chat completion: choices[0].message.content, a finish_reason, and a usage block with prompt, completion and total tokens. The model field carries the Lobstack key that served the request, which is not necessarily the one you asked for.

The headers are the part you do not get elsewhere. This is the response to the request above, sent with "model": "auto":

Response headershttp
x-lobstack-request-id: 8f2b1e0c-4d5a-4b91-9c3e-6a7f0d2b5c41
x-lobstack-model: gemini-3.1-flash-lite
x-lobstack-tier: nano
x-lobstack-complexity: 5
x-lobstack-savings-pct: 98
x-lobstack-routed: true
x-lobstack-mode: managed
x-lobstack-principal: api_key
x-lobstack-cost-usd: 0.000123
x-lobstack-savings-usd: 0.004014
x-lobstack-baseline-model: claude-fable-5-1
x-lobstack-baseline-reason: plan_ceiling
x-lobstack-baseline-usd: 0.004137
x-lobstack-priced: true
x-lobstack-metered: true
x-lobstack-quota-meter: spend
x-lobstack-quota-allowance-usd: 10.000000
x-lobstack-quota-spent-usd: 1.284000
x-lobstack-quota-remaining-usd: 8.716000

The request said auto, so nobody named a model, and x-lobstack-baseline-reason says so: plan_ceiling means the saving is measured against the most expensive model this plan may reach — what the caller would have paid had they asked for the best one. That is a real comparison and it is not the same claim as named, which measures against a model the caller actually asked for. Read the reason before you render the figure. Every header is listed on Metering & cost.

x-lobstack-quota-meter names which allowance the counters beside it describe. A plan that includes dollars of model spend reports dollars and no request counters; the BYOK plan (no longer sold) and the older messages-per-month tiers report counts instead. The three meters are set out on Pricing & plans.

x-lobstack-request-id is allocated before anything can fail, so even a 401 comes back with one. It is the id of the trace row, which means quoting it in a support thread turns a description into a lookup.


The 5 providers

One credential reaches all of them. Lobstack holds the provider keys and routing can cross vendors freely. On a paid plan you can add your own key for any of them in Settings › Provider keys; a call to that provider then runs on your key, is not marked up, and does not spend the allowance.

ProviderCurrent modelsKeys in registry
AnthropicClaude Fable 5.1, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.56
OpenAIGPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna9
GoogleGemini 3.1 Pro, Gemini 3.8 Flash, Gemini 3.1 Flash Lite5
xAIGrok 4.6, Grok 4.5, Grok 4.3, Grok Build 0.16
DeepSeekDeepSeek V4 Pro, DeepSeek V4 Flash4

The larger number counts previous-generation keys kept alive so stored configurations keep resolving. The full catalog, with prices and context windows, is on Models & providers.


When to send auto, and when to name a model

Send "auto" when the traffic is mixed and you care more about the bill than about which engine answers. The router scores each request 0–100 and takes the lowest tier that clears the score, bounded by your plan.

Name a model when the output has to be reproducible, when you are comparing engines, or when a downstream evaluation is pinned to one vendor. Naming a model does not force it: it sets the ceiling. Ask for Opus 5 and send a one-line greeting and the request still goes to the nano tier. If you need a specific model on every request regardless of complexity, pin the model and read x-lobstack-routed to see when it was not used.

Routing is not the differentiatorOpenRouter, LiteLLM and Requesty all route by complexity, and LiteLLM does it free and open source. The unusual part is the receipt: the score, the tier, the model and the price on the response itself, and a null saving where there is nothing honest to measure. You can run the router on your own prompt without spending a token on the page that runs the router on your own prompt.

The rest of this section

Chat Completions

The endpoint in full: request body, response shape, streaming, tool calls.

Embeddings

OpenAI-compatible embeddings: the request, the two models, and how they are metered.

Web search

A query in, titles, links and snippets out: the request, the per-search price, and the errors.

Models & providers

Every key the Gateway serves, its provider, tier, context window and price.

Nex

The routing algorithm in full: the seven signals and their points, the five tier thresholds, plan ceilings, what it does not do, and how to reproduce any decision.

Metering & cost

Every cost header, the labelled baseline, the three quota meters, and what a ledger row records.

Errors & retries

Six error classes, their status codes, and which of them are worth retrying.

SDK

Using the OpenAI SDK unchanged, and reading the Lobstack headers through it.

Lobstack

An AI team that asks before it acts, and an API with a receipt on every call.

© LobstackXLinkedInGitHub