Gateway API
Concepts
Ten terms the rest of these pages use without stopping to define them. One short section each.
Most of Lobstack is one request path: a credential arrives, it resolves to a context, the context bounds which models the router may reach, the router picks one, a provider answers, and two rows are written. Everything below is a name for a step in that sentence.
API keys and scopes
A Lobstack key is lsk_<env>_<selector><secret>: an environment marker (live or test), eight hex characters of selector, and forty-eight hex characters of secret. The selector is not secret. It is stored as the key's prefix and is what makes a key identifiable in a list or a log line after issue.
What the database holds is the SHA-256 of the whole key and nothing else. Verification is one indexed read on the prefix followed by a constant-time hash comparison. A key you have lost cannot be recovered, only revoked and replaced.
| Scope | Grants |
|---|---|
inference | Calling Gateway inference endpoints, including chat completions, embeddings and web search. |
usage:read | Reading the usage and cost API. |
agents:read | Vestigial. See Legacy, below. |
agents:write | Vestigial. See Legacy, below. |
Absence of a scope is a denial. New keys get inference and usage:read. A key may optionally expire; that is set at creation and cannot be added later.
The Gateway request path
The primary chat endpoint is POST /api/gateway/v1/chat/completions. Its request id is allocated before authentication, so even a rejected request has an id you can quote back. Embeddings and web search share authentication, quota and receipt behavior, but do not use this chat-routing path.
Step 5 happens whether or not step 4 succeeded. A failed round still writes a ledger row with zero tokens and a trace row carrying the error class, which is why error rate is computable from the tables alone.
Model keys and provider model ids
You address models by a Lobstack model key, such as claude-sonnet-5 or gemini-3.1-pro. Each key maps to a provider and to that provider's own model id, which is often different: gemini-3.1-pro sends gemini-3.1-pro-preview, gpt-5.6 sends gpt-5.6-sol.
Some keys are aliases that deliberately point somewhere else: an older key resolves to the model that replaced it, so grok-3 is served by grok-4.3 and a configuration that still names it keeps working. The registry holds 30 keys in all — 17 you can ask for today across 5 providers, and 13 previous-generation keys kept so an older agent row still resolves and still prices. A separate table of non-canonical spellings resolves onto those, and resolves for pricing as well as for routing. Every one carries a price per million tokens. Four DeepSeek entries are flagged priceUnverified because their USD list prices could not be confirmed against a first-party page.
Tiers and complexity
Every prompt is scored 0–100 by a heuristic that looks at length, code fences, multi-step phrasing, analysis verbs and conversation depth. The score selects a tier, and the tier selects a model. A request with tools has a standard tier floor: tool definitions do not raise the score, but they keep the request off a tier that is too weak for dependable function calling.
| Tier | Handles complexity up to | First choice on “auto” |
|---|---|---|
| nano | 20 | gemini-3.1-flash-lite |
| small | 40 | claude-haiku-4-5 |
| standard | 70 | claude-sonnet-5 |
| premium | 90 | gemini-3.1-pro |
| flagship | 100 | claude-opus-5 |
standard to DeepSeek V4 Flash instead of Claude Sonnet 5 would be roughly twelve times cheaper and a different answer, so the order inside each tier is a quality preference set by hand.Plan ceilings
A ceiling is the highest tier a request may reach. Two things can lower it. The plan sets one: Free is capped at standard, and every paid plan reaches flagship. Naming a model sets the other, because asking for a standard model caps the request at standard even on a plan that could go higher.
The effective ceiling is the lower of the two. A hard prompt on Free does not get a flagship model; it gets the best standard model and a tier header that says so. Nothing fails, and nothing silently costs more.
Subscriptions bought on the older messages-per-month tiers keep their own ceilings, which are held in a separate table and are not being migrated.
Plans, allowances and meters
A plan buys access, a ceiling, and an allowance. The allowance is denominated in dollars of model spend on every plan on sale, priced at the same rates every response reports. A message was never a unit of cost: the same word covered a forty-token ping and a two-hundred-thousand-token context, so an allowance denominated in messages bounded nothing.
| Meter | Who is on it | The unit |
|---|---|---|
spend | Every plan that includes model spend | Dollars, at Lobstack rates |
requests | The BYOK plan, no longer sold | Requests per calendar month |
legacy | Subscriptions on the older messages tiers | Messages, counted as they always were |
Every response names the meter it was measured against in x-lobstack-quota-meter, because the counters mean different things on each. The plans themselves, and what each includes, are on Pricing & plans.
Metering, and the two prices
Each round writes a ledger row that carries its own arithmetic: token counts, the per-million prices that were in force at the time, and the resulting cost. Prices are frozen into the row, so editing the catalog later cannot move a historical invoice. A model missing from the catalog is priced as null, never as zero, and the response says so via x-lobstack-priced: false.
A request has two costs and they are not the same number. cost_usd is what the customer owes: the provider's list price times 1.25 on managed traffic, and pass-through on calls that ran on your own provider key and on direct, where you already paid your provider. provider_cost_usd is what we paid, it is recorded on the row, and it never leaves the server — no header, no stream frame, no API response carries it.
The baseline, and why it has a reason
A baseline is the model a saving is measured against, and it always arrives with baseline_reason saying why that model was chosen. named means you asked for it and the router served something else; the saving is a measurement against your own request. plan_ceiling means you sent "auto" and named nothing, so the comparison is against the most expensive model your plan may reach.
Both are real subtractions between two real prices, and they are not the same claim. A plan-ceiling comparison is also the most flattering one available to us, which is exactly why it is labelled. Where neither applies the baseline is null and no saving is reported.
x-lobstack-savings-pct is a routing heuristic derived from tier cost multipliers, and it is present on Chat Completions responses. x-lobstack-savings-usd is a measured dollar delta and is present only when a baseline exists. They are not two views of the same figure and adding them together means nothing.Principals
The principal is which door a caller came through. It is recorded on every trace row, because "which credential caused this spend" is the first question anyone asks when a bill moves.
| Principal | Credential | Notes |
|---|---|---|
api_key | Authorization: Bearer lsk_… | Org-owned, scoped, revocable. Quota is enforced. |
gateway_token | Authorization: Bearer per-agent token | Legacy agent-VM credential. Quota is enforced. |
agent_secret | x-agent-secret + x-agent-id | Legacy platform fallback for a VM whose token is missing or stale. Quota is enforced. |
An API key bound to an agent row inherits that row's model, mode and provider key. An unbound key is org-level: managed mode, the org's plan tier, and no agent on the usage row. Every key minted today is the second kind.
Allowances are enforced for every credential. API keys are the supported way for a new caller to reach the Gateway; the agent credentials above remain only for legacy callers.
Modes
The mode says who paid the provider. It is on every ledger row and in x-lobstack-mode.
| Mode | Whose provider key | Effect on routing |
|---|---|---|
managed | Lobstack's | Routes freely across providers, restricted to those whose key is actually configured. |
byok | Yours, added in Settings › Provider keys on a paid plan | Adds that provider to what the router can reach. Your key is sent only to the provider that issued it. |
direct | Yours, outside the Gateway | A legacy runtime calling a provider itself. Recorded, never billed by us. |
The Gateway first works out which providers a key exists for — Lobstack's, plus any your organization added — and routes only among those. The mode is then set per call by whose key the chosen provider runs on. Without that check, "auto" could select an unkeyed provider and every request would fail with a 503 raised one layer too late.
Legacy
Lobstack used to run agents in hosted VMs. That product was retired on 11 September 2026; these notes stay for anyone who used it. A runtime was a long-running process on an agent's dedicated VM that held a credential and called the Gateway on its behalf. No machine is created now, and none can be.
The word survives in a grouping in the Console and in agent_id on a trace row; both are empty for an API key. The agents:* scopes guarded routes that went with the runtime, and stay listed so existing keys keep their grants. Binding a new key to an agent row is refused.
plan_tier, a key bound to it before binding was retired still inherits from it, and Gateway authentication reads it on every call — so it is the billing record, not a machine record. Words like server_tierand region still appear on it and no longer describe anything.