Get started
Social & Companions

Build engaging companions & social apps

AI friends, characters, and social worlds that listen, express emotion, stay in character, and respond in realtime.
Animation.inc
Azimov
Bible Chat
Isekai Zero
Janitor
Kaon
Status
Twizl
Weekend

Native-quality voices for any character, in 100+ languages

Pick from hundreds of library voices, clone a voice from seconds of audio, or design one from a text prompt. Steering tags like [laughing] and [whispering] bring your experience to life across 100+ languages.
  • English
  • Spanish
  • French
  • German
  • Japanese
  • Korean
Explore all voices

Know how the user sounds, not just what they said

Use Inworld STT across 30+ languages to capture speech, detect turn-taking, and return voice signals like emotion, pace, pitch, and vocal style, each with confidence scores. When a user is tired, anxious, or excited, the companion's reply can match, softer, faster, leaning in, without wiring it by hand.
Voice profile signals
Emotion
Frustrated
84%
Age
Adult
84%
Accent
British
84%
Pitch
High
84%
Vocal Style
Shouting
84%
More signals coming soon

Companions users can actually talk to

Use the Inworld Realtime API for low-latency, speech-to-speech conversations. One realtime session handles listening, semantic turn detection, barge-in interruptions, and spoken responses. Users can trail off, interrupt, and banter naturally, with instant replies.
Companion conversation

Right-size the model for each user

Route each conversation by metadata: plan, region, mode, or custom fields. Free tiers stay on fast, cheap models, while paying users get premium models. A/B test personalities and measure retention, not just latency. With automatic fallback across 200+ models, a conversation never dies with a provider outage.
plan
free
mode
roleplay
tone
playful
region
jp
Inworld Router
claude-sonnet-4.6
Premium plans only
gemma-4-31b
Free tier Β· in-character Β· lowest cost
deepseek-v3.2
Longer, more involved threads

Your users' confidential chats remain theirs

Your user's conversations remain theirs. Inworld supports zero-data-retention configurations, on-prem deployment on your own GPUs, and enterprise-grade security, so your users can be themselves without worrying about privacy.
Built for confidential conversations
Zero data retention
Run TTS with ZDR, no audio or text persisted. Nothing to delete because nothing is kept.
On-prem deployment
Deploy TTS on your own H100 / B200 / B300 GPUs. Conversations never leave your network.
Enterprise controls
Retention, deployment, and access controls for privacy-sensitive AI voice workflows.
ZDR
SOC 2 Type II
GDPR
H100 / B200 / B300

Scale affordably on closed and open models

We understand consumer app economics and are committed to providing the highest quality models and inference for the lowest possible price. Whether it's TTS, STT, or LLMs, you will never get better price-performance elsewhere. And with our Inference service, you can access consumer-optimized open models for up to 50% lower than public third-party rates.
Consumer-scale pricing
$ per 1M characters

Text-to-speech price vs the market

Other providersInworld
Provider API rates, June 2026. Inworld on the Growth plan; as low as $5 at enterprise scale. *Gemini 3.1 Flash TTS is billed by audio output tokens; ~$180/1M is the effective rate on a typical query. †Cartesia estimated from published tier pricing.

Use cases

Personal companions

Always-available voice companions for conversation, reflection, entertainment, and daily engagement.

Character chatbots

Bring characters to life with distinct voices and speaking styles.

Roleplay and story companions

Support immersive scenarios where timing, tone, and emotional response matter.

Spiritual & lifestyle companions

Create guided conversational experiences for reflection, motivation, or support, with your policies and guardrails.

FAQ

Yes. Cross-lingual cloning preserves a companion's voice identity across supported languages. Voice Design can also produce new companion or roleplay voices from a text description.
Inworld TTS supports 100+ languages and STT supports 30+, so agents can handle multilingual queues and keep one consistent brand voice across every language.
Yes. Realtime STT-1 returns a voice profile with every utterance, covering emotion, pace, pitch, and vocal style, each with confidence scores. Use those signals to adapt tone, pacing, and how the companion responds. Learn more about voice profiling here.
Yes. The Realtime API runs speech-to-speech conversation under a second end-to-end, with semantic turn detection and barge-in interruption handling, fast enough for live conversations under concurrent load.

Start building

Join millions of developers building the next wave of AI applications.
Copyright Β© 2021-2026 Inworld AI