Skip to main content
Chat Completions can return JSON as the assistant’s text. You ask for a JSON object, or for an instance of a schema you name. You parse the string yourself. The API does not hand you a decoded object. This works on every current model. The example uses lume-3.5. Messages does not accept response_format. A Messages request that needs JSON should say so in the prompt, and your application should parse the text block.

Endpoint and auth

POST /v1/chat/completions needs the scope chat:completions. response_format.type is text, json_object, or json_schema. For a schema, json_schema.name is required. schema, description, and strict are optional. strict is accepted and sent on with the schema. The whole response_format object can be at most 32,768 bytes.
For a JSON object with no schema, set response_format to {"type": "json_object"} and tell the model, in the message, to return JSON.

What you get back

The JSON is the string in message.content. finish_reason is stop when the model finishes inside max_tokens.
strict: true is forwarded with the schema. That is the contract. The page does not promise a separate guarantee that the text will satisfy the schema. Set max_tokens so the JSON has room to finish. lume-3.5 allows 16,384 output tokens and a 400,000 token context window. A response_format.type other than text, json_object, or json_schema is HTTP 400:
The response includes X-Request-Id.

Next

Tools

Return a function call instead of JSON text.

Generate text

A normal reply, with no response format.