> ## Documentation Index
> Fetch the complete documentation index at: https://docs.overcontrolgroup.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate text

> Send a conversation to a Lume model and read the next reply.

A text request carries the conversation so far, and the model answers the latest question. Chat Completions and Messages both work this way, with the same API key and the same model IDs. The examples use `lume-3.5`.

Each call stands on its own. To continue, send the earlier messages again, with the new question at the end.

## Chat Completions

`POST /v1/chat/completions` needs the scope `chat:completions`. Send your key as `Authorization: Bearer $OVERCONTROL_API_KEY`, or as `x-api-key`. [Authentication](/authentication) covers scopes and what a rejected key looks like.

A `system` or `developer` message is the instruction from your application. A `user` message is the person's question. An `assistant` message is a reply from an earlier turn, which you include so the model can see it. `tool` and `function` messages are for [tools](/text/tools).

The example asks for a short follow-up. `max_tokens` is 256, so the call reserves a short reply on `lume-3.5`.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.overcontrolgroup.com/v1/chat/completions \
    -H "Authorization: Bearer $OVERCONTROL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "lume-3.5",
      "max_tokens": 256,
      "messages": [
        {"role": "system", "content": "Answer in one sentence."},
        {"role": "user", "content": "What is the capital of Portugal?"},
        {"role": "assistant", "content": "The capital of Portugal is Lisbon."},
        {"role": "user", "content": "Which river runs through it?"}
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      api_key=os.environ["OVERCONTROL_API_KEY"],
      base_url="https://api.overcontrolgroup.com/v1",
  )

  response = client.chat.completions.create(
      model="lume-3.5",
      max_tokens=256,
      messages=[
          {"role": "system", "content": "Answer in one sentence."},
          {"role": "user", "content": "What is the capital of Portugal?"},
          {"role": "assistant", "content": "The capital of Portugal is Lisbon."},
          {"role": "user", "content": "Which river runs through it?"},
      ],
  )

  print(response.choices[0].message.content)
  ```

  ```javascript Node.js theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    apiKey: process.env.OVERCONTROL_API_KEY,
    baseURL: "https://api.overcontrolgroup.com/v1",
  });

  const response = await client.chat.completions.create({
    model: "lume-3.5",
    max_tokens: 256,
    messages: [
      { role: "system", content: "Answer in one sentence." },
      { role: "user", content: "What is the capital of Portugal?" },
      { role: "assistant", content: "The capital of Portugal is Lisbon." },
      { role: "user", content: "Which river runs through it?" },
    ],
  });

  console.log(response.choices[0].message.content);
  ```
</CodeGroup>

The new reply is `choices[0].message.content`. `finish_reason` is `stop` when the model finishes inside `max_tokens`. The counts below are an illustration.

```json theme={null}
{
  "id": "chatcmpl_0123456789abcdef0123456789abcdef",
  "object": "chat.completion",
  "created": 1789430400,
  "model": "lume-3.5",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The Tagus runs through Lisbon."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 42,
    "completion_tokens": 8,
    "total_tokens": 50
  }
}
```

The response includes `X-Request-Id`.

`temperature` runs from 0 to 2. You set the length of the reply with `max_tokens` or `max_completion_tokens`. On `lume-3.5` the ceiling is 16,384 tokens. A larger value is brought down to that ceiling. If you leave the cap out, the request reserves the full 16,384. `stop` is one string, or up to four, where the reply should end.

The messages and that cap need to fit in the 400,000 token context window. When they do not, you'll get HTTP 400:

```json theme={null}
{
  "error": {
    "message": "Request exceeds the context window for model 'lume-3.5'",
    "type": "invalid_request_error",
    "param": "messages",
    "code": "context_length_exceeded"
  }
}
```

[Models](/models) has the window and the output cap for every ID.

## Messages

`POST /v1/messages` needs the scope `messages:create` and the header `anthropic-version: 2023-06-01`. The instruction goes in the top-level `system` field. `messages` holds the turns: the person's question, the assistant reply, and the new question.

`max_tokens` is required. The official Python SDK adds `/v1/messages` itself, so its base URL is `https://api.overcontrolgroup.com`.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.overcontrolgroup.com/v1/messages \
    -H "x-api-key: $OVERCONTROL_API_KEY" \
    -H "anthropic-version: 2023-06-01" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "lume-3.5",
      "max_tokens": 256,
      "system": "Answer in one sentence.",
      "messages": [
        {"role": "user", "content": "What is the capital of Portugal?"},
        {"role": "assistant", "content": "The capital of Portugal is Lisbon."},
        {"role": "user", "content": "Which river runs through it?"}
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from anthropic import Anthropic

  client = Anthropic(
      api_key=os.environ["OVERCONTROL_API_KEY"],
      base_url="https://api.overcontrolgroup.com",
  )

  message = client.messages.create(
      model="lume-3.5",
      max_tokens=256,
      system="Answer in one sentence.",
      messages=[
          {"role": "user", "content": "What is the capital of Portugal?"},
          {"role": "assistant", "content": "The capital of Portugal is Lisbon."},
          {"role": "user", "content": "Which river runs through it?"},
      ],
  )

  print(message.content[0].text)
  ```
</CodeGroup>

The new reply is in `content`, as a text block. `stop_reason` is `end_turn` when the model finishes on its own. The counts below are an illustration.

```json theme={null}
{
  "id": "msg_0123456789abcdef0123456789abcdef",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "The Tagus runs through Lisbon."
    }
  ],
  "model": "lume-3.5",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 42,
    "output_tokens": 8
  }
}
```

Messages responses include `X-Request-Id` and `request-id`.

`temperature` and `top_p` each run from 0 to 1. On `lume-3.5`, `max_tokens` has to be from 1 through 16,384. Above that, the response is HTTP 400 and the message is `'max_tokens' exceeds the model limit`. `stop_sequences` takes up to four strings, and each one can be at most 1,024 bytes.

The messages and `max_tokens` need to fit in the 400,000 token context window. When they do not, you'll get HTTP 400. The body has no `error.code`:

```json theme={null}
{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "Request exceeds the model context window"
  },
  "request_id": "req_0123456789abcdef0123456789abcdef"
}
```

## Next

<CardGroup cols={2}>
  <Card title="Streaming" icon="wave-pulse" href="/text/streaming">
    Read the reply as it is written.
  </Card>

  <Card title="Tools" icon="wrench" href="/text/tools">
    Let the model call a function in your application.
  </Card>

  <Card title="Structured outputs" icon="braces" href="/text/structured-outputs">
    Ask Chat Completions for JSON.
  </Card>

  <Card title="Models" icon="brain" href="/models">
    Compare context windows and output caps.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.