stream to true. Every current model can stream. The examples use lume-3.5.
A client SDK is the comfortable way to read the stream. The frames below are what that SDK is parsing, so you can also read them yourself.
Chat Completions
POST /v1/chat/completions needs the scope chat:completions. The stream is server-sent events. Each frame is a data: line, a blank line, and then the next frame. The stream ends with data: [DONE].
choices[0].delta.content. The chunk that finishes the reply has an empty delta and finish_reason set to stop.
stream_options.include_usage to true. That field is only valid when stream is true. Content chunks then include "usage": null, and a later chunk has an empty choices array and the usage object:
include_usage out, and the chunks do not contain usage at all.
Once the stream has started, a failure is not a new HTTP status. You get an error object, then data: [DONE].
The model returned no visible assistant content and the code empty_model_response.
Messages
POST /v1/messages needs messages:create and anthropic-version: 2023-06-01. Set stream to true. Each frame is an event: line, a data: line, and a blank line. This stream does not end with data: [DONE].
ping can arrive among them.
A text delta looks like this:
event: error frame with the Messages error body, and then the stream ends. That is not a new HTTP status. The body has error.type and error.message, and no error.code. An empty reply uses error.type api_error and the message The model returned no visible assistant content.
Next
Generate text
The same request without a stream.
Tools
A streamed tool call arrives as its own deltas.

