How is the price calculated?
Billing is per token, at the price shown for each model in the catalog. We measure the exact input and output token counts for each call and price them at that model’s published rate. Cache hits are not billed.
Which models can I use?
Every model enabled in the catalog. If you do not name one, the router picks the lowest-cost enabled model that offers the capability you asked for; if you do name one, that is the one we use.
Do you cache prompts?
Yes. An identical prompt, to the same model, at the same temperature, is served from your own cache for as long as your cache window allows. The cache is keyed per account, so one account never serves another account’s response, and a cached call is not billed. You set the window, and setting it to zero turns caching off.
Is there a free allowance?
Yes. Every account starts with a token allowance, and the current figure is shown on your usage page in the product. It is a platform setting rather than a contractual entitlement, so treat the figure in the product as the current one.
What happens when I pass the free allowance?
Nothing changes about your access: passing the allowance does not interrupt anything, and there is no tier you are moved into. Metered usage is recorded for every call regardless. A request can still be refused in three cases — you have hit the requests-per-minute limit on your account, inference is paused for maintenance, or the account has been suspended.
Who processes payment?
Our payment processor. Card details go to it directly and never reach our servers or our database; we hold the amount, the currency, the status and the processor’s reference for each payment, and nothing else about it.
Is my data used for training?
No. We do not train on your traffic and we do not sell it. For your own request history we keep each call’s token counts, latency, cache status and price, together with a short preview of the prompt and the response. Each upstream provider has its own data policy, surfaced in the model detail view.
Can I use my existing client?
Yes. The API takes a chat-completions-compatible request shape, so a client you have already written works by changing the base URL and the key.
Can a model I need be added?
The catalog is extensible by the operator, and adding a chat-completions-compatible endpoint is one catalog entry. Ask us and we will tell you whether it is something we can enable.