Get started
Realtime STT

The lowest word error rate on real-world audio

Coval's production benchmark measures accuracy on real-world, noisy audio, not clean lab recordings. Inworld STT 1 scores 2.4% WER, the lowest of 30+ models tested (all-model average: 4.5%), starting at $0.10 per hour.

Word error rate by model

Lower is better. The top 12 of 30+ models on Coval's production benchmark, with the all-model average marked.

What is WER?

Word error rate is the percentage of words a model transcribes incorrectly, through substitutions, insertions, or deletions. Lower is better.

Why production audio matters

Most benchmarks use clean, studio-quality recordings. Coval's production dataset uses the noisy audio real applications face, so scores reflect what you'll actually ship.

Top-of-the-benchmark accuracy, bottom-of-the-market price

Published streaming price per hour across the same models. Inworld STT 1 lists at $0.10 per hour, the lowest streaming rate of the providers shown.

Streaming price per hour, lower is better

12 models

STT 1, Inworld AI

$0.10

STT RT v5, Soniox

$0.12

Ink 2, Cartesia

~$0.13

Grok STT, xAI

$0.20

Enhanced, Speechmatics

$0.43

Universal 3.5 Pro, AssemblyAI

$0.45

Nova 3, Deepgram

$0.46

Pulse, Smallest

~$0.54

Chirp 3, Google

~$0.96

Chirp 2, Google

~$0.96

Default, Azure

~$1.00

GPT Realtime Whisper, OpenAI

~$1.02

Published streaming list rates as of July 27, 2026; approximate where marked. Volume, batch, and commitment tiers can be lower.

Source: Coval STT benchmark, production dataset, trailing 30 days. Retrieved July 27, 2026. Independent third-party benchmark; scores are rolling and change over time. Google Chirp 3 shares the 2.4% top score. Live results: benchmarks.coval.ai/stt.

Measure it on your own audio

Free tier included. Send a clip to the Realtime STT playground and check the accuracy yourself.
Copyright © 2021-2026 Inworld AI