TL;DR: Cartesia Sonic 3.6 led the Artificial Analysis Speech Arena on September 2, 2026, while Inworld TTS-2 ranked second and TTS-2 Flash sixth. Choose from production requirements: listener preference, delivery control, P99 latency, languages, streaming, deployment, and cost. Reproduce the comparison on your own voices and text.
How do the models compare?
Cartesia Sonic 3.6 and Inworld TTS-2 are close competitors in realtime speech, while Inworld Flash provides a separate lower-latency and lower-cost tier. Cartesia ranked higher in the current listener-preference snapshot. Inworld combines its TTS models with steering, cloning, multilingual delivery, and a broader realtime API and inference system.
Which model has stronger voice quality?
The two Artificial Analysis leaderboards disagree, and both are worth reading. On the Artificial Analysis Controlled Voice Arena, which holds the voice constant by testing every model on the same eight cloned voices, Inworld Realtime TTS-2 ranks first at Elo 1123 and Cartesia Sonic 3.6 second at 1119; both carry a 1-2 rank range because their confidence intervals overlap. On the Speech Arena on September 2, 2026, Sonic 3.6 ranked first at 1,282 against 1,250 for TTS-2. Neither result settles every production use case.
Test pronunciation, emotional direction, long-session consistency, and target languages separately. A leaderboard uses selected prompts and native voices. It does not measure a cloned brand voice, application-specific jargon, interruption recovery, or cost per useful session.
How should latency be compared?
Inworld reports under-100ms server-side P99 TTFB for TTS-2 and under 25ms for Flash. Cartesia publishes realtime latency claims for Sonic, but a valid comparison requires identical client regions, audio formats, concurrency, connection state, and timing boundaries. Measure first playable audio and complete turn latency from the client.
- P50, P95, and P99 latency.
- Warm and cold connections.
- Peak concurrent requests.
- Errors, retries, and cancelled audio.
- Regional network performance.
How do pricing and workload fit differ?
Artificial Analysis listed normalized API prices of $20.80 per million characters for Inworld TTS-2, $10.40 for the model it listed as Flash research preview, and $49 for Cartesia Sonic 3.6 on September 2, 2026. Contracted pricing may differ. Compare accepted audio, retries, commitments, and user outcomes rather than assuming the listed rate is delivered cost.
Cartesia may justify its price when its output wins a workload-specific preference test. Inworld Flash may fit high-volume utility speech; TTS-2 may fit directed emotional turns. A mixed application can route valuable turns to one model and routine turns to another if the orchestration remains observable.
When is Cartesia the stronger choice?
Cartesia is a strong choice when Sonic 3.6 wins blind tests on the application's voices and text, when its SDK and streaming behavior fit the stack, or when its latency profile performs better from target regions. Inworld is stronger when its steering, two-model choice, pricing, or broader realtime infrastructure provides more product value.
Neither answer should be assumed. Run the same corpus through both providers, hide model identity, and record preference, word accuracy, identity, P99 latency, failures, and cost. State limitations and the date because both systems can change.
Related Guides
Key Takeaways
- Inworld TTS-2 leads the Artificial Analysis Controlled Voice Arena; Cartesia Sonic 3.6 leads its Speech Arena.
- Inworld offers separate expressive and latency-focused TTS-2 models.
- Normalized prices favor Inworld in the September 2 snapshot.
- Client-side P99 tests are required for a valid latency comparison.
- The workload, not one leaderboard, should choose the provider.
Frequently Asked Questions
Is Cartesia Sonic 3.6 better than Inworld TTS-2?
It depends on the leaderboard. On the Artificial Analysis Controlled Voice Arena, which compares models using the same eight cloned voices, Inworld Realtime TTS-2 ranks first at Elo 1123 and Cartesia Sonic 3.6 second at 1119, both sharing a 1-2 rank range because their confidence intervals overlap. On the Artificial Analysis Speech Arena on September 2, 2026, Cartesia ranked higher. Neither result establishes universal superiority, so run a same-experiment comparison under production conditions.
Which is cheaper?
Artificial Analysis listed lower normalized prices for both Inworld models than Cartesia Sonic 3.6. Actual cost depends on plan, commitments, retries, accepted output, and engineering. Confirm live rates and calculate cost per useful interaction before deciding.
Which is faster?
Inworld publishes under 25ms server-side P99 TTFB for Flash and under 100ms for TTS-2. Compare those with Cartesia using identical timing boundaries, concurrency, regions, connections, and audio formats. Client-measured first playable audio is more useful than unmatched vendor figures.
Published by Inworld. Rankings and normalized prices reflect the Artificial Analysis Speech Arena on September 2, 2026 and may change.