Get started
Games & Media

Build expressive gaming & media experiences

Narrated media, interactive stories, and characters that can speak, listen, and react in realtime.
AstroBeam
Disney
Latitude
Little Umbrella
NBCUniversal
Niantic
Particle
Streamlabs
Ubisoft
Weekend
Xbox

Native-quality voices for any character, in 100+ languages

Pick from hundreds of library voices, clone a voice from seconds of audio, or design one from a text prompt. Steering tags like [laughing] and [whispering] bring your experience to life across 100+ languages.
  • English
  • Spanish
  • French
  • German
  • Japanese
  • Korean
Explore all voices

Know how the player sounds, not just what they said

Use Inworld STT across 30+ languages to capture commands and conversation, detect turn-taking, and return voice signals like emotion, pace, pitch, and vocal style, each with confidence scores. When a player is tense, excited, or whispering, the world can react in kind, without wiring it by hand.
Voice profile signals
Emotion
Frustrated
84%
Age
Adult
84%
Accent
British
84%
Pitch
High
84%
Vocal Style
Shouting
84%
More signals coming soon

Characters players can actually talk to

Use the Inworld Realtime API for low-latency, speech-to-speech conversations. One realtime session handles listening, semantic turn detection, barge-in interruptions, and spoken responses, fast enough for live, interruptible dialogue.
In-game dialogue

Route based on any dimension

Route each line by metadata: scene, platform, NPC load, locale, or custom fields. Throwaway barks go to fast open models, while story-critical cutscenes get the strongest reasoning. A/B test dialogue styles and measure engagement, not just latency. With automatic fallback across 200+ models, a scene never stalls on a provider outage.
scene
boss_fight
platform
quest_3
npcs_on_screen
12
locale
ja
Inworld Router
claude-sonnet-4.6
Story-critical cutscenes
gemma-4-31b
Combat barks and ambient NPCs Β· lowest latency
deepseek-v3.2
Companion NPCs with memory

APIs that plug into any engine

Every model is a standard HTTP API you call over REST or WebSocket, with WebRTC for in-browser realtime. The LLM router is a drop-in for the OpenAI and Anthropic SDKs, so you point your existing code at Inworld with only a base-URL change. Wire voiced characters into Unreal, Unity, a custom engine, or a media pipeline with the tools your team already uses.
Plug into any engine
REST & WebSocket
TTS, STT, Realtime API, Router
WebRTC
in-browser realtime
Drops into
Unreal EngineUnity
Custom engines
SDKs
Node.jsPython
Any language

Scale affordably on closed and open models

We understand game/consumer economics and are committed to providing the highest quality models and inference for the lowest possible price. Whether it's TTS, STT, or LLMs, you will never get better price-performance elsewhere. And with our Inference service, you can access consumer-optimized open models for up to 50% lower than public third-party rates.
Consumer-scale pricing
$ per 1M characters

Text-to-speech price vs the market

Other providersInworld
Provider API rates, June 2026. Inworld on the Growth plan; as low as $5 at enterprise scale. *Gemini 3.1 Flash TTS is billed by audio output tokens; ~$180/1M is the effective rate on a typical query. †Cartesia estimated from published tier pricing.

Use cases

Game characters and NPCs

Create characters that respond to player speech, world state, and story context.

IP-based experiences

Let fans interact with approved characters, hosts, creators, or franchises.

Interactive avatars

Power realtime speaking avatars for apps, media, events, commerce, and entertainment.

Interactive content and ads

Create media that adapts to viewer input, context, or campaign logic.

News, sports, and entertainment

Build conversational hosts, explainers, recaps, fan experiences, or personalized content.

FAQ

Inworld is API-first: call TTS, STT, and the Realtime API over REST, WebSocket, or WebRTC from Unreal, Unity, a custom engine, or a media pipeline. The LLM router is also available via API and is OpenAI- and Anthropic-compatible, so you can point your existing SDK at Inworld with a base-URL change.
Yes. Inworld TTS streams word and character timestamps with the audio, plus visemes for real-time lip-sync, so animated characters and video stay in sync with generated speech.
Realtime TTS generates speech in milliseconds, fast enough for live, interruptible dialogue under high concurrent load. The Realtime API runs speech-to-speech conversation under a second end-to-end.
Yes. Design a unique voice for each character from a text description, or clone one from a short sample.
Yes. Inworld TTS can power narration, dubbing, and audio for all forms of media products. Word-level timestamps drive captions and follow-along highlighting, and cross-lingual voices keep one identity across languages.
Yes. Inworld Inference serves top open models, including Gemma, DeepSeek, Kimi, MiniMax, GLM, and more, tuned for consumer workloads at up to 50% below public third-party rates, and the Router falls back across 200+ models.

Start building

Join millions of developers building the next wave of AI applications.
Copyright Β© 2021-2026 Inworld AI