Realtime TTS-2, ranked #1 on Artificial Analysis, is a next generation voice model built for realtime conversation. It hears the full audio of the exchange, picks up the user's tone, pacing and emotional state, and takes voice direction in plain English.
Launch partners & customers say
Capabilities
A natural-language description of how a line should be delivered, passed inline at the start of your text. Not a fixed list of preset emotions. Not a slider. Write the prompt the way you'd write a stage direction.
“I missed you. How was today?”
[speak warm and soothing, welcoming someone home after a long day]Independent benchmark rankings
Inworld models are consistently top-ranked on Artificial Analysis.Blind tests by thousands of real users, not internal evals.
ELO score vs. cost per 1M characters. Higher and further left is better. Source: Artificial Analysis Controlled Voice Arena Leaderboard, August 2026.
A tired, slower delivery for someone winding down at the end of the day.
Direct and purposeful. It reads the urgency and matches it with action.
Three languages inside one generation. Same speaker, same person on the other end.
One prose direction reshapes the read.
Built for natural realtime conversation
A real conversation isn't just words. It's the tone someone uses, the pause before they answer, the energy they carry into a sentence. Most voice agents stitch a pipeline together from four vendors and lose all of that signal at every handoff. We built each layer ourselves and pass the full audio context, the user's state, and the conversation history through one persistent connection — so the system can decide not just what to say, but how to say it.
Realtime STT transcribes and profiles the speaker in one pass. Age, accent, pitch, vocal style, emotional tone, and pacing become structured signals on the same connection. The rest of the pipeline knows who is talking and how they feel, not just what they said.
200+ models
Realtime Router takes the user's state and the conversation context and selects the right model, prompt, and tools for the moment. Reasoning, retrieval, and tool calls all happen on the same persistent connection.
Realtime TTS-2 takes the prior audio, the user's emotional state, the conversation history, and the developer's natural-language direction and decides how to deliver the line. Same words, different read for the moment. Sub-200ms first chunk, identity-preserved across >200 languages.
The capabilities that change what you can build, not feature counts. Quality ranks come from the Artificial Analysis Controlled Voice Arena.
| Capability | ElevenLabs | Cartesia | OpenAI | ||
|---|---|---|---|---|---|
| Voice quality (Artificial Analysis Controlled Voice Arena) | #1 | Not ranked | #5 | #2 | Not ranked |
| Natural conversational delivery | Not supported | Not supported | |||
| Realtime latency | Not supported | Not supported | Not supported | ||
| Multi-turn aware speech synthesis | Not supported | Not supported | Not supported | ||
| Simple voice direction (inline tags) | |||||
| Advanced voice direction (free-form descriptions) | Not supported | Not supported | Not supported | ||
| Voice cloning | Not supported | Not supported | |||
| Voice design | Not supported | Not supported | Not supported | ||
| Crosslingual (single voice, >200 languages) | Not supported | Not supported | Not supported | ||
| Voice profiling (understand user context) | Not supported | Not supported | Not supported | Not supported | |
| Single customizable speech-to-speech API | Not supported | Not supported | Not supported | Not supported | |
| User-aware LLM routing | Not supported | Not supported | Not supported | Not supported | |
| Optimized alphanumeric support | Not supported | Not supported | Not supported |
Verified August 2026 from public docs and the Artificial Analysis Controlled Voice Arena leaderboard. Based on the latest models from each provider.
Realtime TTS-2 ships through the Inworld API and the Inworld Realtime API. Customers on Realtime TTS 1.5 upgrade by changing the model identifier, no other code changes. Code samples at docs.inworld.ai. Pricing at inworld.ai/pricing.