Skip to main content
OverControl serves a family of Lume models through one API. You choose a model for the context, the output length, the kind of input, and the price you want. Changing the model ID leaves the base URL and your API key as they are.
For most applications, start with Lume 3.5. Lume 3.5 Max is the higher-priced text model, at the same 400,000 token context and 16,384 token output cap. Lume 2 is the lower rate, for work that fits in 200,000 tokens of context and 8,192 tokens of output. Lume Omni is the model that can also take an image, audio, or a PDF.
Lume 3 (lume-3) and Lume 3 Max (lume-3-max) are still accepted. They match Lume 3.5 and Lume 3.5 Max on context, output, and euro price.

Choose a model

Lume 3.5

The model these guides use. A strong fit for general text, coding, analysis, and application workloads.

Lume 3.5 Max

The same context and output cap as Lume 3.5, at twice the token rate. A fit for harder reasoning, coding, and multi-step work.

Lume 2

The lowest text rate. A fit for lightweight, high-volume work inside a 200,000 token context window.

Lume Omni

Text plus image, audio, and PDF input on Chat Completions. Messages can send it images.

Comparison

Rates for these models are on Pricing. The ID is the string you put in model. GET /v1/models returns these IDs in this order when the key has no allowlist. That response includes id, object, created, and owned_by.

Context window

The context window is how much room the model has for one request: your instructions, the messages, and the reply. What you send, plus the output you ask for, needs to fit inside it.
A larger window does not by itself make a better answer. What helps is sending the context that belongs to the task.
On Chat Completions you can set max_tokens, or max_completion_tokens, but not both in the same request. If you ask for more than the maximum in the table, OverControl brings the value down to that maximum. If you set neither, the cap is the model’s own maximum: 8,192 tokens on Lume 2, and 16,384 on the others. On Messages you always send max_tokens. It has to be a whole number from 1 up to that same maximum. Anything higher comes back as HTTP 400, with the message 'max_tokens' exceeds the model limit. If the input and that output cap are larger than the context window, the call is also HTTP 400. On Chat Completions the message names the model:
Messages says the same thing in its own shape. The response includes X-Request-Id and request-id.
If you name a model OverControl does not serve, Chat Completions returns HTTP 400 with Unsupported model 'not-a-model'. Messages returns HTTP 404 with Model not found. Both bodies are in the quickstart.

Images, audio, and PDFs

Lume Omni is the model that can take more than text. On Chat Completions, an image can be GIF, JPEG, PNG, or WebP, up to 20 MiB. Audio can be MP3 or WAV, up to 25 MiB. A PDF can be up to 25 MiB. Managing files uploads one, and chat with files asks about it. Messages can send images to Lume Omni. It will not accept audio or PDF blocks. The text-only models will not accept image, audio, or PDF input on either API.

List the models your key can call

You can ask the API which models your key is allowed to use. GET /v1/models needs the models:read scope. The call is authenticated, and you are not billed for it. If the key has an allowlist, the list contains only those models.
Each object has an id, object, created, and owned_by. A real response lists every model this key can call, in the same order as the table. The sample shows that shape with two of them. If the key is missing or not recognized, you’ll get HTTP 401:
The response includes X-Request-Id. A key without models:read gets HTTP 403 and the message Forbidden. Authentication shows that body.

Next

Quickstart

Send a first request with Lume 3.5.

Pricing

See token rates in euros and dollars, and how included requests scale by plan.