Skip to main content
A text request carries the conversation so far, and the model answers the latest question. Chat Completions and Messages both work this way, with the same API key and the same model IDs. The examples use lume-3.5. Each call stands on its own. To continue, send the earlier messages again, with the new question at the end.

Chat Completions

POST /v1/chat/completions needs the scope chat:completions. Send your key as Authorization: Bearer $OVERCONTROL_API_KEY, or as x-api-key. Authentication covers scopes and what a rejected key looks like. A system or developer message is the instruction from your application. A user message is the person’s question. An assistant message is a reply from an earlier turn, which you include so the model can see it. tool and function messages are for tools. The example asks for a short follow-up. max_tokens is 256, so the call reserves a short reply on lume-3.5.
The new reply is choices[0].message.content. finish_reason is stop when the model finishes inside max_tokens. The counts below are an illustration.
The response includes X-Request-Id. temperature runs from 0 to 2. You set the length of the reply with max_tokens or max_completion_tokens. On lume-3.5 the ceiling is 16,384 tokens. A larger value is brought down to that ceiling. If you leave the cap out, the request reserves the full 16,384. stop is one string, or up to four, where the reply should end. The messages and that cap need to fit in the 400,000 token context window. When they do not, you’ll get HTTP 400:
Models has the window and the output cap for every ID.

Messages

POST /v1/messages needs the scope messages:create and the header anthropic-version: 2023-06-01. The instruction goes in the top-level system field. messages holds the turns: the person’s question, the assistant reply, and the new question. max_tokens is required. The official Python SDK adds /v1/messages itself, so its base URL is https://api.overcontrolgroup.com.
The new reply is in content, as a text block. stop_reason is end_turn when the model finishes on its own. The counts below are an illustration.
Messages responses include X-Request-Id and request-id. temperature and top_p each run from 0 to 1. On lume-3.5, max_tokens has to be from 1 through 16,384. Above that, the response is HTTP 400 and the message is 'max_tokens' exceeds the model limit. stop_sequences takes up to four strings, and each one can be at most 1,024 bytes. The messages and max_tokens need to fit in the 400,000 token context window. When they do not, you’ll get HTTP 400. The body has no error.code:

Next

Streaming

Read the reply as it is written.

Tools

Let the model call a function in your application.

Structured outputs

Ask Chat Completions for JSON.

Models

Compare context windows and output caps.