Skip to content
LUNAROUTEDocs

Audio transcription

POST /v1/audio/transcriptions speaks the OpenAI transcription shape. Send an audio file as multipart/form-data with your lr_ key and the gateway returns the transcript. This endpoint is operator-gated and ships switched off — see Availability.

Terminal window
curl https://gw.lunaroute.com/v1/audio/transcriptions \
-H "LUNAROUTE-API-KEY: $LUNAROUTE_API_KEY" \
-F file=@recording.mp3 \
-F model=whisper-large-v3 \
-F response_format=verbose_json

The same call through the official OpenAI SDK (set base_url to https://gw.lunaroute.com/v1):

from openai import OpenAI
client = OpenAI(api_key="lr_...", base_url="https://gw.lunaroute.com/v1")
with open("recording.mp3", "rb") as audio:
transcription = client.audio.transcriptions.create(
model="whisper-large-v3",
file=audio,
)
print(transcription.text)

POST /v1/audio/transcriptions, multipart/form-data.

Part Rule
file Required, exactly once. The audio bytes. The part’s filename and Content-Type are advisory only — the gateway does not decode or inspect audio.
model Required, exactly once. Must be transcription-capable; see Models.
language Optional. An ISO-639-1 code with an optional region (en, pt-BR).
prompt Optional. Context to bias transcription; bounded in length.
temperature Optional. Must be within [0, 1] inclusive.
response_format Optional. json (default), text, or verbose_json.
timestamp_granularities[] Optional, repeatable. segment or word. Requires response_format=verbose_json.

Any part name not listed above is rejected with 400 unsupported_parameter — nothing is silently ignored.

Limit Value
Upload size 25 MiB per file
Parts per request 16
Text field size 64 KiB per text part
Duration Per model — audio_config.max_duration_seconds (between 1 and 3600; models ship configured at 3600).

Limits are enforced while the body is read, so an over-cap upload fails mid-stream with 413 audio_too_large rather than after being buffered. The cap is enforced by the gateway, but the transport in front of it may reject a larger request body first with its own opaque 413/400 (the production deployment sits behind a Cloud Run HTTP/1 request ceiling of roughly 32 MiB).

Duration is bounded by size, not only by the model

Section titled “Duration is bounded by size, not only by the model”

The upload cap binds before the duration limit for most real audio, because duration is not something the gateway can read out of a compressed container — it only sees bytes. How long 25 MiB lasts depends entirely on bitrate:

Encoding 25 MiB carries
MP3 128 kbps (common) ~27 minutes
64 kbps ~55 minutes
48 kbps speech-optimised ~73 minutes
32 kbps Opus (speech) ~109 minutes

So the 3600-second model limit is reachable, but it needs roughly 58 kbps or below. A one-hour recording at 128 kbps is about 58 MiB and is refused with 413 audio_too_large before duration is ever considered. For long recordings, encode for speech — mono, 32-48 kbps — rather than at a music-grade bitrate.

response_format Body Content-Type
json (default) {"text": "…"} application/json
text The transcript, raw text/plain; charset=utf-8
verbose_json {"task", "language", "duration", "text", "segments": […]} application/json

verbose_json passes through the model server’s verbose payload, so duration is in seconds and segments carries per-segment timing. Transcribing an hour of audio can produce a large segments (and words) array.

Responses are not streamed, and there is no stream option: the field is not recognised, so sending it is rejected with 400 unsupported_parameter rather than silently ignored. The request stays open until the whole transcript is ready and then returns in a single body — the time to first byte equals the total request time, and the response carries a Content-Length.

Size your client timeout for the audio, not for a typical HTTP call. At the fleet’s measured throughput (roughly 40x real time) an hour of audio takes about 90 seconds, and the gateway’s overall deadline is 600 seconds. A default 30-second client timeout will fail on perfectly normal files.

Every response carries an x-lunaroute-request-id. Quote it if you need to report a problem.

Audio with no detectable speech returns an empty transcript — {"text": ""} with zero segments — and a 200, not an error. Speech-free audio is detected before it reaches the transcription model, so a silent upload, or one carrying only room tone, transcribes to nothing after the job runs rather than to an invented phrase.

This is not a promise about noisy audio: a room with music, breathing or faint background speech is not silence, and the model will attempt to transcribe whatever it can hear. If you need to know whether a file contained speech at all, the transcript itself is the answer — empty means none was detected.

A speech-free upload is not metered: the model server reports duration as 0, so no audio_second usage is recorded for it. That is why duration reads 0 rather than the file’s length on a silent response — a file containing speech reports its real duration as usual.

segment timestamps are always available in verbose_json. Word-level timestamps depend on the model: a model advertises support through its audio_config.supports_word_timestamps, and requesting timestamp_granularities[]=word against a model that does not support it is rejected with 400 unsupported_parameter — the gateway never returns silently null timestamps. When word timestamps are both requested and supported, verbose_json additionally carries words: [{"word", "start", "end"}].

Terminal window
GET /v1/audio/transcriptions/models

Returns the transcription models available to your organization with their declared limits (audio_config), mirroring /v1/embeddings/models. A model appears here only when it is transcription-capable and its configuration is valid; a model marked for transcription without a valid audio_config is never advertised. whisper-large-v3 is the reference model.

Audio duration is metered on every request: the gateway always obtains the model server’s reported duration and records it, in milliseconds, in the audio_duration_ms field of the request’s usage event (modality audio). The price unit for audio is audio_second.

The metered quantity is the model server’s reported duration, and for a speech-free upload that is 0 — a silent file is not counted as billable audio. See Silence.

At launch the audio policy is included: requests are metered for measurement but customers are not charged. Metered pricing is built (the audio_second price unit and the would-charge decision exist), but charging is not yet enabled, and it additionally requires the admin price-rule authoring path for audio_second rules. Nothing in this endpoint debits a wallet today.

Audio, the transcript, the prompt, and the uploaded filename are never logged and never persisted. The only thing recorded for a transcription is the measured duration on the usage event. The transcript exists only in the response body returned to you.

Audio transcription is served by LunaRoute’s managed ASR fleet and is live on every auth surface below.

A platform kill switch can take transcription down without a deploy. While it is engaged, every transcription call — and the models listing — answers 503 audio_transcription_disabled, and the route stays registered rather than disappearing. Treat that as “temporarily unavailable”, not “not found”: it is a reversible operator action, not a missing endpoint, so nothing about your request is wrong and retrying later is the right response.

The endpoint is served on the same auth surfaces as the rest of the API:

Surface Path
Header auth /v1/audio/transcriptions
Path auth /a/{api_key}/v1/audio/transcriptions
Project path /a/{api_key}/p/{project}/v1/audio/transcriptions
Routing plan /r/{routing_plan_id}/v1/audio/transcriptions

Responses on the /a/... path surfaces carry Deprecation: true, like every other path-auth route.

Errors use the standard envelope, {"error": {"code", "message"}}, with lowercase codes. Common ones:

Code Status Meaning
audio_transcription_disabled 503 The operator kill switch is engaged. Retry later; see Availability.
audio_disabled 403 Audio is switched off for your organization.
invalid_request 400 Malformed multipart body, a missing or duplicate field, or a model that cannot serve transcription.
unsupported_parameter 400 An unknown part, a bad response_format, an out-of-range temperature, a bad language, or word timestamps that the model does not support.
audio_too_large 413 The file part exceeded the size cap.
audio_too_long 400 The decoded audio exceeded the model’s duration limit.
body_read_timeout 408 The upload did not finish within the read deadline.
audio_at_capacity 503 The gateway is at its concurrent-audio ceiling. Retry.
audio_lanes_exhausted 429 Your organization is at its audio concurrency allowance. Back off and retry.
audio_backend_rate_limited 429 The audio backend is shedding load. Back off and retry.
audio_backend_unavailable 502 The audio backend was unreachable or failed.
audio_deadline_exceeded 504 The request exceeded the overall deadline.