Audio transcription
POST /v1/audio/transcriptions speaks the OpenAI transcription shape. Send an
audio file as multipart/form-data with your lr_ key and the gateway returns
the transcript. This endpoint is operator-gated and ships switched off — see
Availability.
Quickstart
Section titled “Quickstart”curl https://gw.lunaroute.com/v1/audio/transcriptions \ -H "LUNAROUTE-API-KEY: $LUNAROUTE_API_KEY" \ -F file=@recording.mp3 \ -F model=whisper-large-v3 \ -F response_format=verbose_jsonThe same call through the official OpenAI SDK (set base_url to
https://gw.lunaroute.com/v1):
from openai import OpenAI
client = OpenAI(api_key="lr_...", base_url="https://gw.lunaroute.com/v1")
with open("recording.mp3", "rb") as audio: transcription = client.audio.transcriptions.create( model="whisper-large-v3", file=audio, )
print(transcription.text)Request
Section titled “Request”POST /v1/audio/transcriptions, multipart/form-data.
| Part | Rule |
|---|---|
file |
Required, exactly once. The audio bytes. The part’s filename and Content-Type are advisory only — the gateway does not decode or inspect audio. |
model |
Required, exactly once. Must be transcription-capable; see Models. |
language |
Optional. An ISO-639-1 code with an optional region (en, pt-BR). |
prompt |
Optional. Context to bias transcription; bounded in length. |
temperature |
Optional. Must be within [0, 1] inclusive. |
response_format |
Optional. json (default), text, or verbose_json. |
timestamp_granularities[] |
Optional, repeatable. segment or word. Requires response_format=verbose_json. |
Any part name not listed above is rejected with 400 unsupported_parameter —
nothing is silently ignored.
Limits
Section titled “Limits”| Limit | Value |
|---|---|
| Upload size | 25 MiB per file |
| Parts per request | 16 |
| Text field size | 64 KiB per text part |
| Duration | Per model — audio_config.max_duration_seconds (between 1 and 3600; models ship configured at 3600). |
Limits are enforced while the body is read, so an over-cap upload fails
mid-stream with 413 audio_too_large rather than after being buffered. The cap
is enforced by the gateway, but the transport in front of it may reject a
larger request body first with its own opaque 413/400 (the production
deployment sits behind a Cloud Run HTTP/1 request ceiling of roughly 32 MiB).
Duration is bounded by size, not only by the model
Section titled “Duration is bounded by size, not only by the model”The upload cap binds before the duration limit for most real audio, because duration is not something the gateway can read out of a compressed container — it only sees bytes. How long 25 MiB lasts depends entirely on bitrate:
| Encoding | 25 MiB carries |
|---|---|
| MP3 128 kbps (common) | ~27 minutes |
| 64 kbps | ~55 minutes |
| 48 kbps speech-optimised | ~73 minutes |
| 32 kbps Opus (speech) | ~109 minutes |
So the 3600-second model limit is reachable, but it needs roughly 58 kbps or
below. A one-hour recording at 128 kbps is about 58 MiB and is refused with 413 audio_too_large before duration is ever considered. For long recordings, encode
for speech — mono, 32-48 kbps — rather than at a music-grade bitrate.
Responses
Section titled “Responses”response_format |
Body | Content-Type |
|---|---|---|
json (default) |
{"text": "…"} |
application/json |
text |
The transcript, raw | text/plain; charset=utf-8 |
verbose_json |
{"task", "language", "duration", "text", "segments": […]} |
application/json |
verbose_json passes through the model server’s verbose payload, so duration
is in seconds and segments carries per-segment timing. Transcribing an hour of
audio can produce a large segments (and words) array.
Streaming
Section titled “Streaming”Responses are not streamed, and there is no stream option: the field is not recognised, so
sending it is rejected with 400 unsupported_parameter rather than silently ignored. The request
stays open until the whole transcript is ready and then returns in a single body — the time to
first byte equals the total request time, and the response carries a Content-Length.
Size your client timeout for the audio, not for a typical HTTP call. At the fleet’s measured throughput (roughly 40x real time) an hour of audio takes about 90 seconds, and the gateway’s overall deadline is 600 seconds. A default 30-second client timeout will fail on perfectly normal files.
Every response carries an x-lunaroute-request-id. Quote it if you need to report a problem.
Silence
Section titled “Silence”Audio with no detectable speech returns an empty transcript — {"text": ""} with zero
segments — and a 200, not an error. Speech-free audio is detected before it reaches the
transcription model, so a silent upload, or one carrying only room tone, transcribes to nothing
after the job runs rather than to an invented phrase.
This is not a promise about noisy audio: a room with music, breathing or faint background speech is not silence, and the model will attempt to transcribe whatever it can hear. If you need to know whether a file contained speech at all, the transcript itself is the answer — empty means none was detected.
A speech-free upload is not metered: the model server reports duration as 0, so no
audio_second usage is recorded for it. That is why duration reads 0 rather than the file’s
length on a silent response — a file containing speech reports its real duration as usual.
Timestamps
Section titled “Timestamps”segment timestamps are always available in verbose_json. Word-level
timestamps depend on the model: a model advertises support through its
audio_config.supports_word_timestamps, and requesting
timestamp_granularities[]=word against a model that does not support it is
rejected with 400 unsupported_parameter — the gateway never returns silently
null timestamps. When word timestamps are both requested and supported,
verbose_json additionally carries words: [{"word", "start", "end"}].
Models
Section titled “Models”GET /v1/audio/transcriptions/modelsReturns the transcription models available to your organization with their
declared limits (audio_config), mirroring /v1/embeddings/models. A model
appears here only when it is transcription-capable and its configuration is
valid; a model marked for transcription without a valid audio_config is never
advertised. whisper-large-v3 is the reference model.
Billing
Section titled “Billing”Audio duration is metered on every request: the gateway always obtains the
model server’s reported duration and records it, in milliseconds, in the
audio_duration_ms field of the request’s usage event (modality audio). The
price unit for audio is audio_second.
The metered quantity is the model server’s reported duration, and for a
speech-free upload that is 0 — a silent file is not counted as billable
audio. See Silence.
At launch the audio policy is included: requests are metered for
measurement but customers are not charged. Metered pricing is built (the
audio_second price unit and the would-charge decision exist), but charging is
not yet enabled, and it additionally requires the admin price-rule authoring
path for audio_second rules. Nothing in this endpoint debits a wallet today.
Privacy
Section titled “Privacy”Audio, the transcript, the prompt, and the uploaded filename are never
logged and never persisted. The only thing recorded for a transcription is the
measured duration on the usage event. The transcript exists only in the
response body returned to you.
Availability
Section titled “Availability”Audio transcription is served by LunaRoute’s managed ASR fleet and is live on every auth surface below.
A platform kill switch can take transcription down without a deploy. While it
is engaged, every transcription call — and the models listing — answers
503 audio_transcription_disabled, and the route stays registered rather than
disappearing. Treat that as “temporarily unavailable”, not “not found”: it is a
reversible operator action, not a missing endpoint, so nothing about your
request is wrong and retrying later is the right response.
The endpoint is served on the same auth surfaces as the rest of the API:
| Surface | Path |
|---|---|
| Header auth | /v1/audio/transcriptions |
| Path auth | /a/{api_key}/v1/audio/transcriptions |
| Project path | /a/{api_key}/p/{project}/v1/audio/transcriptions |
| Routing plan | /r/{routing_plan_id}/v1/audio/transcriptions |
Responses on the /a/... path surfaces carry Deprecation: true, like every
other path-auth route.
Errors
Section titled “Errors”Errors use the standard envelope, {"error": {"code", "message"}}, with
lowercase codes. Common ones:
| Code | Status | Meaning |
|---|---|---|
audio_transcription_disabled |
503 |
The operator kill switch is engaged. Retry later; see Availability. |
audio_disabled |
403 |
Audio is switched off for your organization. |
invalid_request |
400 |
Malformed multipart body, a missing or duplicate field, or a model that cannot serve transcription. |
unsupported_parameter |
400 |
An unknown part, a bad response_format, an out-of-range temperature, a bad language, or word timestamps that the model does not support. |
audio_too_large |
413 |
The file part exceeded the size cap. |
audio_too_long |
400 |
The decoded audio exceeded the model’s duration limit. |
body_read_timeout |
408 |
The upload did not finish within the read deadline. |
audio_at_capacity |
503 |
The gateway is at its concurrent-audio ceiling. Retry. |
audio_lanes_exhausted |
429 |
Your organization is at its audio concurrency allowance. Back off and retry. |
audio_backend_rate_limited |
429 |
The audio backend is shedding load. Back off and retry. |
audio_backend_unavailable |
502 |
The audio backend was unreachable or failed. |
audio_deadline_exceeded |
504 |
The request exceeded the overall deadline. |