> ## Documentation Index
> Fetch the complete documentation index at: https://docs.overcontrolgroup.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Image understanding

> Send an image to Lume Omni and get a text reply.

Image understanding sends a picture and gets text back. The model does not return an image. Only `lume-omni` accepts image input. On any other model the image is rejected.

Chat Completions and Messages describe the image differently. The limits are different too, so each section below is only about that API.

## Chat Completions

`POST /v1/chat/completions` needs the scope `chat:completions`. Put the image in a user message, next to the question, as an `image_url` part.

The URL can be `http` or `https`, and the path ends in `.gif`, `.jpeg`, `.jpg`, `.png`, or `.webp`. It can also be a base64 data URL. `detail` is optional: `auto`, `low`, or `high`. There is no cap on how many images you attach, and no combined-size cap. Each image can be at most 20 MiB. A data-URL image over that size is HTTP 413, with the message `Image input is too large`.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.overcontrolgroup.com/v1/chat/completions \
    -H "Authorization: Bearer $OVERCONTROL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "lume-omni",
      "max_tokens": 256,
      "messages": [
        {
          "role": "user",
          "content": [
            {"type": "text", "text": "What is in this picture?"},
            {
              "type": "image_url",
              "image_url": {
                "url": "https://example.com/diagram.png",
                "detail": "auto"
              }
            }
          ]
        }
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      api_key=os.environ["OVERCONTROL_API_KEY"],
      base_url="https://api.overcontrolgroup.com/v1",
  )

  response = client.chat.completions.create(
      model="lume-omni",
      max_tokens=256,
      messages=[
          {
              "role": "user",
              "content": [
                  {"type": "text", "text": "What is in this picture?"},
                  {
                      "type": "image_url",
                      "image_url": {
                          "url": "https://example.com/diagram.png",
                          "detail": "auto",
                      },
                  },
              ],
          }
      ],
  )

  print(response.choices[0].message.content)
  ```
</CodeGroup>

The reply is ordinary text in `choices[0].message.content`. `lume-omni` has a 400,000 token context window and a 16,384 token output cap. The response includes `X-Request-Id`.

If you send the same image to `lume-3.5`, and the image is the second part of the first user message, you get HTTP 400:

```json theme={null}
{
  "error": {
    "message": "Parameter 'messages.0.content.1' is not supported for model 'lume-3.5'",
    "type": "invalid_request_error",
    "param": "messages.0.content.1",
    "code": "unsupported_feature"
  }
}
```

## Messages

`POST /v1/messages` needs `messages:create` and `anthropic-version: 2023-06-01`. The image is base64. A URL source is rejected.

You can send at most 20 images, at most 20 MiB each, and at most 25 MiB combined. Media types are `image/gif`, `image/jpeg`, `image/png`, and `image/webp`. Going over those limits is HTTP 400, with `Image data exceeds the model limit` or `Image input exceeds the request limit`. It is not HTTP 413.

```bash theme={null}
curl https://api.overcontrolgroup.com/v1/messages \
  -H "x-api-key: $OVERCONTROL_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lume-omni",
    "max_tokens": 256,
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "What is in this picture?"},
          {
            "type": "image",
            "source": {
              "type": "base64",
              "media_type": "image/png",
              "data": "<base64>"
            }
          }
        ]
      }
    ]
  }'
```

Replace `<base64>` with the encoded image. The reply is a text block in `content`, the same shape as any other Messages reply. [Generate text](/text/generate) shows that body.

## Next

<CardGroup cols={2}>
  <Card title="Audio understanding" icon="waveform" href="/text/audio-understanding">
    Send an mp3 or wav clip to Lume Omni.
  </Card>

  <Card title="Models" icon="brain" href="/models">
    See which models accept an image.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.