> ## Documentation Index
> Fetch the complete documentation index at: https://docs.overcontrolgroup.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Read a Lume reply as it is written, on Chat Completions or Messages.

Streaming sends the reply in pieces, so your application can show text while the model is still writing. Set `stream` to `true`. Every current model can stream. The examples use `lume-3.5`.

A client SDK is the comfortable way to read the stream. The frames below are what that SDK is parsing, so you can also read them yourself.

## Chat Completions

`POST /v1/chat/completions` needs the scope `chat:completions`. The stream is server-sent events. Each frame is a `data:` line, a blank line, and then the next frame. The stream ends with `data: [DONE]`.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.overcontrolgroup.com/v1/chat/completions \
    -H "Authorization: Bearer $OVERCONTROL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "lume-3.5",
      "max_tokens": 256,
      "stream": true,
      "messages": [
        {"role": "user", "content": "Explain what an API is in one sentence."}
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      api_key=os.environ["OVERCONTROL_API_KEY"],
      base_url="https://api.overcontrolgroup.com/v1",
  )

  stream = client.chat.completions.create(
      model="lume-3.5",
      max_tokens=256,
      stream=True,
      messages=[
          {"role": "user", "content": "Explain what an API is in one sentence."}
      ],
  )

  for chunk in stream:
      text = chunk.choices[0].delta.content
      if text:
          print(text, end="", flush=True)
  ```

  ```javascript Node.js theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    apiKey: process.env.OVERCONTROL_API_KEY,
    baseURL: "https://api.overcontrolgroup.com/v1",
  });

  const stream = await client.chat.completions.create({
    model: "lume-3.5",
    max_tokens: 256,
    stream: true,
    messages: [
      { role: "user", content: "Explain what an API is in one sentence." },
    ],
  });

  for await (const chunk of stream) {
    const text = chunk.choices[0]?.delta?.content;
    if (text) process.stdout.write(text);
  }
  ```
</CodeGroup>

The first chunk sets the assistant role. Later chunks carry the new text in `choices[0].delta.content`. The chunk that finishes the reply has an empty `delta` and `finish_reason` set to `stop`.

```text theme={null}
data: {"id":"chatcmpl_0123456789abcdef0123456789abcdef","object":"chat.completion.chunk","created":1789430400,"model":"lume-3.5","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl_0123456789abcdef0123456789abcdef","object":"chat.completion.chunk","created":1789430400,"model":"lume-3.5","choices":[{"index":0,"delta":{"content":"An API"},"finish_reason":null}]}

data: {"id":"chatcmpl_0123456789abcdef0123456789abcdef","object":"chat.completion.chunk","created":1789430400,"model":"lume-3.5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
```

If you also want token counts, set `stream_options.include_usage` to `true`. That field is only valid when `stream` is `true`. Content chunks then include `"usage": null`, and a later chunk has an empty `choices` array and the usage object:

```json theme={null}
{
  "id": "chatcmpl_0123456789abcdef0123456789abcdef",
  "object": "chat.completion.chunk",
  "created": 1789430400,
  "model": "lume-3.5",
  "choices": [],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 16,
    "total_tokens": 34
  }
}
```

Leave `include_usage` out, and the chunks do not contain `usage` at all.

Once the stream has started, a failure is not a new HTTP status. You get an error object, then `data: [DONE]`.

```json theme={null}
{
  "error": {
    "message": "Internal server error",
    "type": "api_error",
    "param": null,
    "code": "internal_error"
  }
}
```

If the model produces no visible text, the same object uses the message `The model returned no visible assistant content` and the code `empty_model_response`.

## Messages

`POST /v1/messages` needs `messages:create` and `anthropic-version: 2023-06-01`. Set `stream` to `true`. Each frame is an `event:` line, a `data:` line, and a blank line. This stream does not end with `data: [DONE]`.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.overcontrolgroup.com/v1/messages \
    -H "x-api-key: $OVERCONTROL_API_KEY" \
    -H "anthropic-version: 2023-06-01" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "lume-3.5",
      "max_tokens": 256,
      "stream": true,
      "messages": [
        {"role": "user", "content": "Explain what an API is in one sentence."}
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from anthropic import Anthropic

  client = Anthropic(
      api_key=os.environ["OVERCONTROL_API_KEY"],
      base_url="https://api.overcontrolgroup.com",
  )

  with client.messages.stream(
      model="lume-3.5",
      max_tokens=256,
      messages=[
          {"role": "user", "content": "Explain what an API is in one sentence."}
      ],
  ) as stream:
      for text in stream.text_stream:
          print(text, end="", flush=True)
  ```
</CodeGroup>

You will see these events, in this order. A `ping` can arrive among them.

| Event | What it carries |
| - | - |
| `message_start` | The message shell. `content` is empty, `stop_reason` is null, and `usage` is the input count. |
| `content_block_start` | The start of a text block, or of a tool call. |
| `content_block_delta` | `text_delta` with the next characters, or `input_json_delta` with the next piece of tool input. |
| `content_block_stop` | The end of that block. |
| `message_delta` | `stop_reason`, `stop_sequence`, and `usage.output_tokens`. |
| `message_stop` | The stream is finished. |

A text delta looks like this:

```text theme={null}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"An API"}}
```

Once bytes have started, a failure is an `event: error` frame with the Messages error body, and then the stream ends. That is not a new HTTP status. The body has `error.type` and `error.message`, and no `error.code`. An empty reply uses `error.type` `api_error` and the message `The model returned no visible assistant content`.

## Next

<CardGroup cols={2}>
  <Card title="Generate text" icon="message" href="/text/generate">
    The same request without a stream.
  </Card>

  <Card title="Tools" icon="wrench" href="/text/tools">
    A streamed tool call arrives as its own deltas.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.