Skip to main content
Prices are calculated separately for input, cached input, and output tokens.
Every new OverControl API account includes free quota to test the API before adding credit.
All prices below are per 1 million tokens (MTok).

Model pricing

Lume 2 pricing will be added here when its public API pricing is available.

How token pricing works

Your API cost depends on the number and type of tokens processed.
  • Input tokens are the tokens sent to the model, including instructions, messages, and other request context.
  • Cached input tokens are reusable input tokens served at the model’s lower cached-input rate.
  • Output tokens are the tokens generated by the model in its response.
Your total request cost is the sum of these three components.

Example

Suppose a request to Lume 3 uses:
  • 100,000 input tokens;
  • no cached input;
  • 10,000 output tokens.
Using EUR pricing:
For high-volume applications, model choice and cached input can have a significant effect on total API cost.

Choose the right model

Pricing is only one part of model selection. For most applications, start with Lume 3 and evaluate other models when you need greater capability, lower cost, or multimodal support. Compare Lume models

Enterprise

For larger production workloads, OverControl offers enterprise options including dedicated infrastructure, regional data hosting, and custom capacity.

Contact sales

Discuss enterprise workloads, infrastructure requirements, and custom capacity with the OverControl team.

Next steps

Models

Compare Lume models and choose the right model for your workload.

Quickstart

Create an API key and make your first request.

API dashboard

Manage your API access and review your usage.