Skip to main content
Image understanding sends a picture and gets text back. The model does not return an image. Only lume-omni accepts image input. On any other model the image is rejected. Chat Completions and Messages describe the image differently. The limits are different too, so each section below is only about that API.

Chat Completions

POST /v1/chat/completions needs the scope chat:completions. Put the image in a user message, next to the question, as an image_url part. The URL can be http or https, and the path ends in .gif, .jpeg, .jpg, .png, or .webp. It can also be a base64 data URL. detail is optional: auto, low, or high. There is no cap on how many images you attach, and no combined-size cap. Each image can be at most 20 MiB. A data-URL image over that size is HTTP 413, with the message Image input is too large.
The reply is ordinary text in choices[0].message.content. lume-omni has a 400,000 token context window and a 16,384 token output cap. The response includes X-Request-Id. If you send the same image to lume-3.5, and the image is the second part of the first user message, you get HTTP 400:

Messages

POST /v1/messages needs messages:create and anthropic-version: 2023-06-01. The image is base64. A URL source is rejected. You can send at most 20 images, at most 20 MiB each, and at most 25 MiB combined. Media types are image/gif, image/jpeg, image/png, and image/webp. Going over those limits is HTTP 400, with Image data exceeds the model limit or Image input exceeds the request limit. It is not HTTP 413.
Replace <base64> with the encoded image. The reply is a text block in content, the same shape as any other Messages reply. Generate text shows that body.

Next

Audio understanding

Send an mp3 or wav clip to Lume Omni.

Models

See which models accept an image.