Included text requests scale with the plan. Plus is 5× the Free allowance, and Pro is 20×. Chat Completions and Messages share that allowance for the UTC calendar month. Listing models, and calls to the Files API, do not spend it.
Model pricing
- EUR
- USD
How a request is priced
The cost of a request comes from three kinds of tokens.- Input tokens are what you send: instructions, messages, and the rest of the request.
- Cached input tokens are the part of that input served at the model’s lower cached rate. They are already inside the input count.
- Output tokens are the tokens the model writes in its reply.
usage.prompt_tokens is the full input. usage.prompt_tokens_details.cached_tokens appears when the cached count is greater than zero.
Suppose a Lume 3.5 request uses 100,000 input tokens, none of them cached, and 2,000 output tokens:
input_tokens and cache_creation_input_tokens are charged at the input rate. cache_read_input_tokens is charged at the cached rate. Output uses output_tokens. The euro rates are the same ones in the table. Caching can show up on a response, and there is no request field that turns it on. cache_control does not create a cache.
Included requests
Chat Completions and Messages draw from one allowance each UTC calendar month. How much you get depends on the plan, as a multiple of Free.
Once that allowance is used, further text requests are billed in euros at the rates above.
If a request cannot be billed, you’ll get HTTP 429. OverControl does not publish a requests-per-minute cap, and these responses do not carry
x-ratelimit headers.
On Chat Completions the body looks like this.
error.type is rate_limit_error.
error.type is rate_limit_error, and there is no error.code.
X-Request-Id. Messages also sends request-id.
Usage for the current month is in the API dashboard.
Next
Models
Compare the Lume models and pick one for the workload.
Quickstart
Create a key and make a first request.

