Skip to main content
OverControl prices text requests by the model and by the tokens in the call. Prices below are per million tokens. Cached input is a portion of the input, billed at the lower cached rate. Tokens spent creating a cache are billed at the ordinary input rate. There is no separate cache-write price.
Included text requests scale with the plan. Plus is 5× the Free allowance, and Pro is 20×. Chat Completions and Messages share that allowance for the UTC calendar month. Listing models, and calls to the Files API, do not spend it.

Model pricing

You are billed in euros. Context windows and output caps are on Models. Lume 3 matches Lume 3.5 on context, output, and price. Lume 3 Max matches Lume 3.5 Max.

How a request is priced

The cost of a request comes from three kinds of tokens.
  • Input tokens are what you send: instructions, messages, and the rest of the request.
  • Cached input tokens are the part of that input served at the model’s lower cached rate. They are already inside the input count.
  • Output tokens are the tokens the model writes in its reply.
On Chat Completions, usage.prompt_tokens is the full input. usage.prompt_tokens_details.cached_tokens appears when the cached count is greater than zero. Suppose a Lume 3.5 request uses 100,000 input tokens, none of them cached, and 2,000 output tokens:
If 50,000 of those input tokens are cached, the input charge is 50,000 at €0.30 and 50,000 at €0.06, which is €0.018. Output stays €0.003. The request comes to €0.021. On Messages, input_tokens and cache_creation_input_tokens are charged at the input rate. cache_read_input_tokens is charged at the cached rate. Output uses output_tokens. The euro rates are the same ones in the table. Caching can show up on a response, and there is no request field that turns it on. cache_control does not create a cache.
For a busy application, the model and any cached input change the bill more than small differences in prompt wording.

Included requests

Chat Completions and Messages draw from one allowance each UTC calendar month. How much you get depends on the plan, as a multiple of Free. Once that allowance is used, further text requests are billed in euros at the rates above. If a request cannot be billed, you’ll get HTTP 429. OverControl does not publish a requests-per-minute cap, and these responses do not carry x-ratelimit headers. On Chat Completions the body looks like this. error.type is rate_limit_error.
Messages uses the same words, in its own shape. error.type is rate_limit_error, and there is no error.code.
Both responses include X-Request-Id. Messages also sends request-id. Usage for the current month is in the API dashboard.

Next

Models

Compare the Lume models and pick one for the workload.

Quickstart

Create a key and make a first request.