> ## Documentation Index
> Fetch the complete documentation index at: https://docs.overcontrolgroup.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio understanding

> Send an audio clip to Lume Omni and get a text reply.

Audio understanding sends a clip and gets text back. The model does not return audio. Only `lume-omni` accepts it, and only on Chat Completions.

Messages rejects an audio block. You will get HTTP 400, `error.type` `invalid_request_error`, and the message `Unsupported content block type`.

## Chat Completions

`POST /v1/chat/completions` needs the scope `chat:completions`. Put the clip in a user message as an `input_audio` part. `data` is the base64-encoded file. `format` is `mp3` or `wav`. Each clip can be at most 25 MiB. A larger clip is HTTP 413, with the message `Audio input is too large`.

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.overcontrolgroup.com/v1/chat/completions \
    -H "Authorization: Bearer $OVERCONTROL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "lume-omni",
      "max_tokens": 256,
      "messages": [
        {
          "role": "user",
          "content": [
            {"type": "text", "text": "Summarize this clip in one sentence."},
            {
              "type": "input_audio",
              "input_audio": {
                "data": "<base64>",
                "format": "mp3"
              }
            }
          ]
        }
      ]
    }'
  ```

  ```python Python theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      api_key=os.environ["OVERCONTROL_API_KEY"],
      base_url="https://api.overcontrolgroup.com/v1",
  )

  response = client.chat.completions.create(
      model="lume-omni",
      max_tokens=256,
      messages=[
          {
              "role": "user",
              "content": [
                  {"type": "text", "text": "Summarize this clip in one sentence."},
                  {
                      "type": "input_audio",
                      "input_audio": {"data": "<base64>", "format": "mp3"},
                  },
              ],
          }
      ],
  )

  print(response.choices[0].message.content)
  ```
</CodeGroup>

Replace `<base64>` with the encoded mp3 or wav. The reply is ordinary text in `choices[0].message.content`. `lume-omni` has a 400,000 token context window and a 16,384 token output cap. The response includes `X-Request-Id`.

The HTTP body of the whole request is still capped at 40 MiB. A clip under 25 MiB can still make the request too large once the JSON around it is counted. That response is HTTP 413 with the message `Request body is too large`.

Sending audio to `lume-3.5` is HTTP 400, code `unsupported_feature`, with a message that the parameter is not supported for that model. A format other than `mp3` or `wav` is HTTP 400, message `Audio format '<format>' is not supported for model 'lume-omni'`.

## Next

<CardGroup cols={2}>
  <Card title="Image understanding" icon="image" href="/text/image-understanding">
    Send a picture to Lume Omni instead.
  </Card>

  <Card title="Generate text" icon="message" href="/text/generate">
    A text-only reply on any Lume model.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.