Native-quality voices for any character, in 100+ languages
Pick from hundreds of library voices, clone a voice from seconds of audio, or design one from a text prompt. Steering tags like [laughing] and [whispering] bring your experience to life across 100+ languages.
Native-quality voices for any character, in 100+ languages
Pick from hundreds of library voices, clone a voice from seconds of audio, or design one from a text prompt. Steering tags like [laughing] and [whispering] bring your experience to life across 100+ languages.
Know how the player sounds, not just what they said
Use Inworld STT across 30+ languages to capture commands and conversation, detect turn-taking, and return voice signals like emotion, pace, pitch, and vocal style, each with confidence scores. When a player is tense, excited, or whispering, the world can react in kind, without wiring it by hand.
Know how the player sounds, not just what they said
Use Inworld STT across 30+ languages to capture commands and conversation, detect turn-taking, and return voice signals like emotion, pace, pitch, and vocal style, each with confidence scores. When a player is tense, excited, or whispering, the world can react in kind, without wiring it by hand.
Use the Inworld Realtime API for low-latency, speech-to-speech conversations. One realtime session handles listening, semantic turn detection, barge-in interruptions, and spoken responses, fast enough for live, interruptible dialogue.
Use the Inworld Realtime API for low-latency, speech-to-speech conversations. One realtime session handles listening, semantic turn detection, barge-in interruptions, and spoken responses, fast enough for live, interruptible dialogue.
Route each line by metadata: scene, platform, NPC load, locale, or custom fields. Throwaway barks go to fast open models, while story-critical cutscenes get the strongest reasoning. A/B test dialogue styles and measure engagement, not just latency. With automatic fallback across 200+ models, a scene never stalls on a provider outage.
Route each line by metadata: scene, platform, NPC load, locale, or custom fields. Throwaway barks go to fast open models, while story-critical cutscenes get the strongest reasoning. A/B test dialogue styles and measure engagement, not just latency. With automatic fallback across 200+ models, a scene never stalls on a provider outage.
Every model is a standard HTTP API you call over REST or WebSocket, with WebRTC for in-browser realtime. The LLM router is a drop-in for the OpenAI and Anthropic SDKs, so you point your existing code at Inworld with only a base-URL change. Wire voiced characters into Unreal, Unity, a custom engine, or a media pipeline with the tools your team already uses.
Every model is a standard HTTP API you call over REST or WebSocket, with WebRTC for in-browser realtime. The LLM router is a drop-in for the OpenAI and Anthropic SDKs, so you point your existing code at Inworld with only a base-URL change. Wire voiced characters into Unreal, Unity, a custom engine, or a media pipeline with the tools your team already uses.
We understand game/consumer economics and are committed to providing the highest quality models and inference for the lowest possible price. Whether it's TTS, STT, or LLMs, you will never get better price-performance elsewhere. And with our Inference service, you can access consumer-optimized open models for up to 50% lower than public third-party rates.
Provider API rates, June 2026. Inworld on the Growth plan; as low as $5 at enterprise scale. *Gemini 3.1 Flash TTS is billed by audio output tokens; ~$180/1M is the effective rate on a typical query. β Cartesia estimated from published tier pricing.
Consumer-scale pricing
$ per 1M characters
Text-to-speech price vs the market
Other providersInworld
Provider API rates, June 2026. Inworld on the Growth plan; as low as $5 at enterprise scale. *Gemini 3.1 Flash TTS is billed by audio output tokens; ~$180/1M is the effective rate on a typical query. β Cartesia estimated from published tier pricing.
Scale affordably on closed and open models
We understand game/consumer economics and are committed to providing the highest quality models and inference for the lowest possible price. Whether it's TTS, STT, or LLMs, you will never get better price-performance elsewhere. And with our Inference service, you can access consumer-optimized open models for up to 50% lower than public third-party rates.
Create characters that respond to player speech, world state, and story context.
IP-based experiences
Let fans interact with approved characters, hosts, creators, or franchises.
Interactive avatars
Power realtime speaking avatars for apps, media, events, commerce, and entertainment.
Interactive content and ads
Create media that adapts to viewer input, context, or campaign logic.
News, sports, and entertainment
Build conversational hosts, explainers, recaps, fan experiences, or personalized content.
FAQ
Inworld is API-first: call TTS, STT, and the Realtime API over REST, WebSocket, or WebRTC from Unreal, Unity, a custom engine, or a media pipeline. The LLM router is also available via API and is OpenAI- and Anthropic-compatible, so you can point your existing SDK at Inworld with a base-URL change.
Yes. Inworld TTS streams word and character timestamps with the audio, plus visemes for real-time lip-sync, so animated characters and video stay in sync with generated speech.
Realtime TTS generates speech in milliseconds, fast enough for live, interruptible dialogue under high concurrent load. The Realtime API runs speech-to-speech conversation under a second end-to-end.
Yes. Design a unique voice for each character from a text description, or clone one from a short sample.
Yes. Inworld TTS can power narration, dubbing, and audio for all forms of media products. Word-level timestamps drive captions and follow-along highlighting, and cross-lingual voices keep one identity across languages.
Yes. Inworld Inference serves top open models, including Gemma, DeepSeek, Kimi, MiniMax, GLM, and more, tuned for consumer workloads at up to 50% below public third-party rates, and the Router falls back across 200+ models.
Start building
Join millions of developers building the next wave of AI applications.