TL;DR: Choose Inworld Realtime TTS-2 when expressive delivery and direct authoring control matter most. Choose Realtime TTS-2 Flash when P99 time to first byte, sustained volume, or cost per character dominates. Both support WebSocket streaming, cloning, and 200-plus languages; teams should test quality and latency on their own traffic.
Inworld Realtime TTS-2 and Realtime TTS-2 Flash serve different priorities. TTS-2 is the flagship for expressive output and direct steering. Flash is the speed and cost tier, with under 25ms P99 server-side time to first byte. Choose from the application's latency budget, session volume, required delivery control, and cost per useful interaction.
Both models use Inworld's TTS APIs and share voice assets, so teams can test or route them by turn type. An application may use TTS-2 for important dialogue and Flash for short, latency-sensitive responses. Log the decision so teams can connect each output to latency, cost, quality, and user outcome.
What separates TTS-2 from TTS-2 Flash?
Realtime TTS-2 prioritizes expressive delivery and natural-language steering. Realtime TTS-2 Flash prioritizes response speed, volume, and a lower cost floor. Both support streaming, instant cloning, Professional Voice Cloning in preview, and more than 200 languages. The decision turns on which constraint causes product failure first: expression, latency, or economics.
The latency figures are measured server-side and exclude network transit, client buffering, and playback startup. Pricing reflects rates published on September 2, 2026 and may change. Enterprise terms depend on volume and deployment requirements, so teams should confirm commitments before using the lowest rate in a business case. Record the exact contract assumptions with the decision.
When should teams choose Realtime TTS-2?
Choose Realtime TTS-2 when the performance of a line carries product meaning. It is the stronger starting point for companions, roleplay, interactive media, tutoring, narration, and voice agents that need direct control over tone, pacing, volume, pitch, pauses, or non-verbal expression. It still streams with under-100ms server-side P99 TTFB.
TTS-2 is appropriate when the same words must land differently across turns. A tutor may soften a correction; a character may move from suspicion to relief; a support agent may slow down before a sensitive instruction. Those cases require control accuracy, not simply a pleasant default voice. Teams should evaluate the model with their actual scripts and score whether listeners hear the requested intent without being shown the instruction.
The application generates emotionally variable dialogue.
Writers or developers need request-level delivery direction.
Non-verbal cues and pauses carry meaning.
Under-100ms server-side P99 TTFB fits the complete latency budget.
The additional unit cost is justified by engagement or conversion.
When should teams choose TTS-2 Flash?
Choose Realtime TTS-2 Flash when the speech layer has an extremely small latency budget or the product generates enough audio that cost per character shapes the architecture. Flash reports under 25ms P99 server-side TTFB, starts at $15 per million characters on demand, and retains streaming, cloning, and multilingual support.
Flash is suited to short conversational turns, rapid acknowledgments, high-volume social features, and other cases where waiting for the first audio chunk is more damaging than losing some expressive range. It can also serve utility speech inside a mixed application: menus, confirmations, navigation, or low-emotion turns. The product should still test pronunciation, long-session consistency, and listener preference because a faster first byte does not guarantee a better complete interaction.
Every millisecond materially affects interruption and turn-taking.
The workload produces millions of repeated short utterances.
Cost per character is a primary margin constraint.
The application can use simpler delivery patterns.
An under-25ms server-side TTFB creates useful network headroom.
How should teams compare quality and cost?
Compare models on cost per accepted minute or useful session, not list price alone. Include failed generations, retries, abandoned turns, and revenue or retention attached to each interaction. A cheaper request becomes more expensive when it requires regeneration or reduces the chance that users continue.
According to the Artificial Analysis Speech Arena, Realtime TTS-2 ranked second on September 2, 2026 with an Elo score of 1,250, while the model listed by Artificial Analysis as TTS-2 Flash research preview ranked sixth at 1,222. The table listed normalized prices of $20.80 and $10.40 per million characters. Teams should reproduce those results with their voices, text, rates, methods, and samples under production conditions before deciding.
Assign traffic randomly, keep voice identity and text constant, and measure preference, altered words, playable-audio latency, completion, retries, and downstream behavior. If TTS-2 improves retention on valuable turns, its price may be efficient. If Flash produces equivalent outcomes, lower price and latency should win. Predeclare decision thresholds before the experiment starts, retain raw results, and rerun the test after material model updates.
Which competing TTS models should teams test?
A useful shortlist should represent different quality, latency, deployment, and pricing choices. Artificial Analysis placed Cartesia Sonic 3.6, SpeechifyAI Simba 3.2, Alibaba Qwen-Audio-3.0-TTS-Plus, ElevenLabs v3 Conversational, and Google Gemini 3.1 Flash TTS near Inworld's models publicly on September 2, 2026.
Microsoft Azure Speech, Amazon Polly, OpenAI TTS-1, Hume AI Octave 2, and Deepgram Flux TTS belong in broader evaluations. Cartesia emphasizes conversational speed; ElevenLabs serves creative production; hyperscalers suit cloud consolidation. Compare named versions on identical text, voices, regions, and dates. A fair test may show Inworld is not the right fit. Use production traffic, not vendor demonstrations, for the final choice.
Creative speech: include ElevenLabs and Hume AI.
Conversational latency: include Cartesia and Deepgram.
Cloud consolidation: include Google, Microsoft, Amazon, and OpenAI.
What latency numbers belong in the decision?
Use P95 and P99 time to first byte, first playable audio, and complete turn latency at expected concurrency. Inworld's published under-25ms and under-100ms figures are server-side P99 TTFB, not end-to-end response times. Network distance, audio encoding, buffering, device playback, and upstream model delays still affect the user.
A model should be chosen inside the full voice pipeline. Speech recognition, endpointing, language-model generation, routing, synthesis, and playback each consume part of the budget. Measure from the user's final audio packet to the first audible response, then break that total into components. Flash creates more headroom for the rest of the system, but that headroom matters only if the network and application preserve it.
Published model TTFB is one component of end-to-end latency. Always record client region, server region, audio format, concurrency, warm or cold state, and percentile.
TTS-2 prioritizes expressive control; TTS-2 Flash prioritizes under 25ms P99 TTFB, sustained volume, and lower unit cost.
Both models support WebSocket streaming, voice cloning, and more than 200 languages behind the Inworld API.
Server-side TTFB excludes network, buffering, playback, and upstream speech or language-model delays.
Mixed-model routing can reserve TTS-2 for high-value turns and Flash for latency-sensitive utility speech.
The winning model is the one with the lowest cost per useful interaction under the application's real traffic.
Frequently Asked Questions
What is the main difference between TTS-2 and TTS-2 Flash?
Realtime TTS-2 is Inworld's flagship for expressive delivery and natural-language steering. Realtime TTS-2 Flash is designed for lower latency, higher volume, and lower cost. TTS-2 reports under-100ms P99 server-side TTFB, while Flash reports under 25ms. Both support streaming, cloning, and more than 200 languages.
Which model should a realtime voice agent use?
Start with TTS-2 when the agent needs expressive, directed delivery and under-100ms server-side TTFB fits the full latency budget. Choose Flash when each millisecond affects turn-taking or when high volume makes unit cost decisive. Many applications should test both and route different turn types to different models.
Is TTS-2 Flash always cheaper than TTS-2?
Its published list prices are lower, but delivered cost depends on the contracted tier, text volume, retries, failed generations, and product outcomes. Compare cost per accepted minute or useful session rather than cost per request alone. Confirm enterprise terms directly because the lowest rates depend on volume and deployment requirements.
Can developers change models without rebuilding the application?
Yes, the models are available through the same Inworld TTS product family, so a team can change the model selection behind a feature flag or routing rule. Safe migration still requires regression testing, traffic controls, version logging, and rollback thresholds because pronunciation, variation, timing, and quality can change between models.
Published by Inworld. Product specifications and pricing reflect Inworld's public TTS page on September 2, 2026 and may change. TTFB figures are server-side P99 measurements that exclude network latency. Artificial Analysis rankings and normalized prices reflect its live Speech Arena on September 2, 2026. Teams should test both models on their own workloads.