Skip to main content
Streaming sends the reply in pieces, so your application can show text while the model is still writing. Set stream to true. Every current model can stream. The examples use lume-3.5. A client SDK is the comfortable way to read the stream. The frames below are what that SDK is parsing, so you can also read them yourself.

Chat Completions

POST /v1/chat/completions needs the scope chat:completions. The stream is server-sent events. Each frame is a data: line, a blank line, and then the next frame. The stream ends with data: [DONE].
The first chunk sets the assistant role. Later chunks carry the new text in choices[0].delta.content. The chunk that finishes the reply has an empty delta and finish_reason set to stop.
If you also want token counts, set stream_options.include_usage to true. That field is only valid when stream is true. Content chunks then include "usage": null, and a later chunk has an empty choices array and the usage object:
Leave include_usage out, and the chunks do not contain usage at all. Once the stream has started, a failure is not a new HTTP status. You get an error object, then data: [DONE].
If the model produces no visible text, the same object uses the message The model returned no visible assistant content and the code empty_model_response.

Messages

POST /v1/messages needs messages:create and anthropic-version: 2023-06-01. Set stream to true. Each frame is an event: line, a data: line, and a blank line. This stream does not end with data: [DONE].
You will see these events, in this order. A ping can arrive among them. A text delta looks like this:
Once bytes have started, a failure is an event: error frame with the Messages error body, and then the stream ends. That is not a new HTTP status. The body has error.type and error.message, and no error.code. An empty reply uses error.type api_error and the message The model returned no visible assistant content.

Next

Generate text

The same request without a stream.

Tools

A streamed tool call arrives as its own deltas.