Metered per token · cache hits not billed · no seat charge

The whole catalog. One endpoint.
Pay per token.

Send a request without naming a model and AI Inference routes it to the cheapest model you have enabled that matches the capability you asked for. Every call reports its exact input and output token counts, its latency and its price, as it happens.

  • Billed on exact token counts
  • A free allowance to start
  • No card needed to sign up
  • No seat charge

What the platform actually does

Three mechanisms. Each one is described here as it is implemented, and no further.

The cheapest eligible model wins

Name a model and we use that one. Leave it out and the router takes the lowest-cost model you have enabled that offers the capability you asked for — so adding a cheaper model to the catalog lowers your bill without a code change.

Token-accurate billing

Input and output tokens are counted separately for every call and priced in integer micro-dollars at the rate published for the model that served it. One usage record per request, keyed so that a retry cannot be billed twice.

One endpoint for all of it

A chat-completions-compatible API shape, so an existing client works by changing the base URL. Switching the model behind it does not change your integration.

How billing works

Metered per token. The price for each model is shown in the catalog, in the product, before you send a request to it.

Per token
Granularity

Input and output are counted separately on every call, from the counts the call reports.

Not billed
Cache hits

An identical prompt to the same model at the same temperature, inside your cache window, is served from your cache at no charge.

One
API shape

A chat-completions-compatible endpoint. An existing client works by changing the base URL.

No seat
Per-user charge

You are billed for the usage on the account, not for the number of people who have access to it.

Starter
$0

A free token allowance on every new account. The current figure is shown on your usage page.

  • Every model enabled in the catalog
  • Playground and request history
  • Per-call token counts and price
  • No card required

Built for

Who the product is designed for. This is not a list of customers, and we do not publish one.

  • AI engineering teams
  • Platform operators
  • Agencies and consultancies
  • Internal tooling teams

From first prompt to first invoice

Four steps, none of which involves talking to us.

Step 01

Create an account

An email address and a password. No card is needed, and nothing has to be approved by us first.

Step 02

Configure

Pick a default model or let the router choose. Set your requests-per-minute limit and how long a cached response stays valid for you.

Step 03

Run

Send prompts from the playground or against the API with a key you issue yourself. Each call reports its token counts, its latency, whether it was a cache hit, and its price.

Step 04

Scale

Metered billing on the usage you have already been shown. You pay for the tokens you used — there is no seat to buy and no tier to be moved into.

Reference catalogs

Two reference documents: one for licensing the platform, one for the commercial picture. They describe the same product this page describes, at more length.

Who builds it

AI Inference is built and operated by Ventureship. What this page describes is what the product does today. We hold no security certification and make no uptime commitment, and we would rather under-describe the product than put a number in front of you that you cannot check in the console.

Frequently asked questions

How is the price calculated?

Billing is per token, at the price shown for each model in the catalog. We measure the exact input and output token counts for each call and price them at that model’s published rate. Cache hits are not billed.

Which models can I use?

Every model enabled in the catalog. If you do not name one, the router picks the lowest-cost enabled model that offers the capability you asked for; if you do name one, that is the one we use.

Do you cache prompts?

Yes. An identical prompt, to the same model, at the same temperature, is served from your own cache for as long as your cache window allows. The cache is keyed per account, so one account never serves another account’s response, and a cached call is not billed. You set the window, and setting it to zero turns caching off.

Is there a free allowance?

Yes. Every account starts with a token allowance, and the current figure is shown on your usage page in the product. It is a platform setting rather than a contractual entitlement, so treat the figure in the product as the current one.

What happens when I pass the free allowance?

Nothing changes about your access: passing the allowance does not interrupt anything, and there is no tier you are moved into. Metered usage is recorded for every call regardless. A request can still be refused in three cases — you have hit the requests-per-minute limit on your account, inference is paused for maintenance, or the account has been suspended.

Who processes payment?

Our payment processor. Card details go to it directly and never reach our servers or our database; we hold the amount, the currency, the status and the processor’s reference for each payment, and nothing else about it.

Is my data used for training?

No. We do not train on your traffic and we do not sell it. For your own request history we keep each call’s token counts, latency, cache status and price, together with a short preview of the prompt and the response. Each upstream provider has its own data policy, surfaced in the model detail view.

Can I use my existing client?

Yes. The API takes a chat-completions-compatible request shape, so a client you have already written works by changing the base URL and the key.

Can a model I need be added?

The catalog is extensible by the operator, and adding a chat-completions-compatible endpoint is one catalog entry. Ask us and we will tell you whether it is something we can enable.

Ship inference. Stop metering by hand.

A free allowance to start, no card to sign up, and a bill that is the tokens you used.