Commercial catalog
The commercial reference for AI Inference. It describes what the platform does, what it bills for, and how an operator engages with it. Pair it with the licensing catalog.
What it is
AI Inference is a managed inference routing and metering platform, billed per token. An operator sends a prompt; the platform resolves it to a model in its catalog, dispatches the call, counts the exact input and output tokens the call reports, prices those tokens at the rate published for that model, and records one metered usage event for the request.
If the request names a model, that model is used. If it does not, the router takes the lowest-cost model the operator has enabled that offers the capability requested — so enabling a cheaper model that can do the job lowers the bill without any change to the calling code.
What it bills for
- Tokens. Input and output are counted separately on every call and priced at the rate shown for the model in the catalog. Prices are published per model and visible before a request is sent to it.
- Not cache hits. An identical prompt, to the same model, at the same temperature, inside the account's cache window, is served from that account's own cache and is not billed.
- Not seats. There is no per-user charge and no tier to be moved into. Usage on the account is the whole of the bill.
- A free allowance. Every account starts with a token allowance. It is a platform setting, so the figure shown in the product is the current one; it is not published here as a contractual number.
What problems it solves
- Token metering is fiddly to get right, and an operator building it themselves usually does not.
- Usage has to be reconcilable. Every call is recorded with its token counts, its latency, its cache status and its price, and one metered event is written per request under an idempotency key, so a retried request cannot be counted twice.
- Switching the model behind an integration normally breaks the usage history. Here the catalog entry changes and the history stays continuous.
- A cached response must be exactly free rather than approximately free. A cache hit is priced at zero, not discounted.
- Integrating a new provider normally means a new client. The request shape is chat-completions-compatible, so an existing client works by changing the base URL.
Who it is designed for
AI engineering teams, platform operators, agencies and internal tooling teams that need to run an inference surface without building the routing, metering and billing layer themselves. That describes who the product is designed around. It is not a customer list, and we do not publish one.
How an operator engages with it
Create an account with an email address and a password — no card, and no approval step. Configure a default model, or leave the choice to the router, and set a requests-per-minute limit and a cache window. Run prompts from the playground or against the API with a key issued from the console, and watch each call report its tokens, latency, cache status and price. Scale into metered billing on usage that has already been displayed.
Passing the free allowance does not interrupt anything and does not move the account into a different tier. A request is refused in three cases, each of which returns a clear error: the account has reached its requests-per-minute limit, inference is paused for maintenance, or the account has been suspended.
What it does not claim
We make no uptime commitment and offer no service-level agreement. We hold no security certification or audit attestation. We do not guarantee the accuracy, safety or fitness of a model's output. See the terms and the privacy policy for the full position.
About this catalog
This is the canonical commercial reference for AI Inference and is intended to be cited as such. The licensing catalog covers licensing the platform to an operator who wants to run it under their own brand.