Pricing
Pay for the requests you run, and nothing else
Every model is priced in the unit that matches what it does, charged against a prepaid wallet in US dollars. There is no subscription, no monthly minimum and no plan to cancel.
Rates are read from the live catalogue on every request.
Open LLMs bills per request from a prepaid US-dollar wallet. Chat and vision models are priced per million tokens, speech-to-text per minute of audio, and video per second of output. The minimum top-up is US$1.00 and there is no subscription.
Per million tokens
Billed on the combined tokens of your prompt and the model's reply.
| Row label | Task | Context | Rate |
|---|---|---|---|
| GLM-4.5-Air | Chat | 33K tokens | $1.50 / 1M tokens |
| Wan2.1-VACE Video | Video | — | $50.00 / 1M tokens |
| Qwen-3.5 | Vision | 33K tokens | $0.15 / 1M tokens |
| Qwen-2.5 | Chat | 33K tokens | $0.15 / 1M tokens |
Per minute of audio
| Row label | Task | Context | Rate |
|---|---|---|---|
| Whisper Large v3 | Speech to text | — | $0.004 / minute |
How am I charged?
- Prepaid, in US dollars
- You top up a wallet and requests are metered against it. The wallet is credited the full amount you pay.
- Minimum top-up US$1.00
- There is no monthly minimum and no subscription to cancel.
- Metered per request
- You are charged for the work you actually run, in each model's own unit.
- A call needs a positive balance
- An empty wallet is why an otherwise correct request is refused.
The policy is the authority
Why is this cheaper than a closed model?
Because the weights are open. There is no licence fee priced into a request — the cost is the machine that ran it, and the models here are chosen to be efficient on the hardware that serves them. That is also the honest limit of the claim: on a task where a frontier closed model is genuinely better, the cheaper request is not the better one. Compare on the model catalogue and pick per workload.