Skip to main content

Pricing

Pay for the requests you run, and nothing else

Every model is priced in the unit that matches what it does, charged against a prepaid wallet in US dollars. There is no subscription, no monthly minimum and no plan to cancel.

Rates are read from the live catalogue on every request.

Open LLMs bills per request from a prepaid US-dollar wallet. Chat and vision models are priced per million tokens, speech-to-text per minute of audio, and video per second of output. The minimum top-up is US$1.00 and there is no subscription.

Per million tokens

Billed on the combined tokens of your prompt and the model's reply.

Open LLMs rates — per million tokens
Row labelTaskContextRate
GLM-4.5-AirChat33K tokens$1.50 / 1M tokens
Wan2.1-VACE VideoVideo$50.00 / 1M tokens
Qwen-3.5Vision33K tokens$0.15 / 1M tokens
Qwen-2.5Chat33K tokens$0.15 / 1M tokens

Per minute of audio

Open LLMs rates — per minute of audio
Row labelTaskContextRate
Whisper Large v3Speech to text$0.004 / minute

How am I charged?

Prepaid, in US dollars
You top up a wallet and requests are metered against it. The wallet is credited the full amount you pay.
Minimum top-up US$1.00
There is no monthly minimum and no subscription to cancel.
Metered per request
You are charged for the work you actually run, in each model's own unit.
A call needs a positive balance
An empty wallet is why an otherwise correct request is refused.
Note:

The policy is the authority

This page describes the rates. The terms that govern them — refunds, metering, rounding and what happens to an unspent balance — are in the billing policy and the refund policy.

Why is this cheaper than a closed model?

Because the weights are open. There is no licence fee priced into a request — the cost is the machine that ran it, and the models here are chosen to be efficient on the hardware that serves them. That is also the honest limit of the claim: on a task where a frontier closed model is genuinely better, the cheaper request is not the better one. Compare on the model catalogue and pick per workload.