If your app already routes language models through OpenRouter, it can now speak and listen with the same key. As of October 7, 2026, OpenRouter lists nine ElevenLabs text-to-speech models and two speech-to-text models, callable through POST /api/v1/audio/speech and POST /api/v1/audio/transcriptions. No separate ElevenLabs plan is required, and every ElevenLabs model is 50% off OpenRouter’s list price through October 19, 8am PT.
The practical change is architectural: audio stops being a separate integration. As I argued in my piece on routing as the layer above your model, the value of an aggregator is that swapping providers becomes a one-line change. That now extends to voice — swap openai/gpt-5-mini for another chat model without touching the audio calls around it.
Three models cover most jobs
The announcement suggests starting small:
- Eleven v4 for expressive narration and character work. Audio tags like
[whispering]and[laughing]in the script are read as delivery directions — and they count toward your billed characters. - Eleven v4 Turbo for voice agents, built for real-time replies at half the v4 rate.
- Scribe v2 for transcription, with speaker labels, word timestamps, and 90+ languages. Scribe v2 Medical cuts clinical-term errors by 35% versus Scribe v2, per OpenRouter.
For high volume, Flash v2.5 takes up to 40,000 characters per request; Multilingual v2 adds speed control (0.7–1.2) that v4 models don’t support.
The three-call voice agent
The walkthrough in OpenRouter’s announcement builds a full loop with the same API key: transcribe the question with Scribe v2, answer with any chat model, speak the reply with v4 Turbo.
def speak(text, out_path):
r = requests.post(
f"{API}/audio/speech",
headers=HEADERS,
json={
"model": "elevenlabs/eleven-v4-turbo",
"input": text,
"voice": "george",
"response_format": "mp3",
},
)
r.raise_for_status()
with open(out_path, "wb") as f:
f.write(r.content)
The transcription response includes a usage object with audio seconds and dollar cost, which matters if you’re tracking per-turn agent economics.
Gotchas worth knowing before your first request
A few details will save you debugging time:
- Omit
response_formatand the speech endpoint returns raw PCM at 24 kHz — fine for pipelines, unplayable in most players. Ask formp3to get a 44.1 kHz file. - ElevenLabs-specific settings like
seed,previous_text, andnext_textgo underprovider.options.elevenlabs. Unknown keys are silently ignored; a listed key with a bad value returns a 400. - Splitting long scripts? Pass neighboring text as
previous_textandnext_textso intonation stays consistent across the joins. - Uploads cap at 25 MB (about 27 minutes of 128 kbps MP3). Longer files can go by public URL, but the upstream timeout is 180 seconds, so split very long recordings.
- Cloned and designed voices stay in your own ElevenLabs account — those need BYOK.
What this does and doesn’t change
For builders, the case is consolidation: one billing relationship, one usage stream, one key rotation story across text and audio. Realtime Scribe v2 over WebSocket isn’t part of this launch, and ElevenLabs-only features like voice cloning require bringing your own key — so this doesn’t replace a direct ElevenLabs integration for every use case. But if you’ve been deferring a voice feature because it felt like a second vendor to manage, the cheapest experiment now starts with the key you already have.
