← Blog

Introducing mono-1,
built for dictation

· Monologue

Today, we’re introducing mono-1, our in-house dictation model, built to turn natural speech into text you can use.

Dictation needs to handle corrections, formatting, context, and writing preferences. mono-1 is the culmination of six months of development, production testing, and feedback from Monologue users—improving how Monologue recognizes your words, handles corrections, and formats your writing.

Against our previous API-based pipeline, mono-1 required about 55% fewer estimated edits and delivered finished dictation 3× faster by median API response time.

Accurate recognition, fast responses

Accurate speech recognition is the foundation of good dictation. For Monologue, that means maintaining high accuracy and low latency across the kinds of conditions people actually dictate in.

To evaluate mono-1, we used WildSpeech-Bench, which tests speech recognition across conditions including background noise, overlapping speech, changes in loudness, pauses, and stuttering.

We measured transcription accuracy using word error rate (WER). Lower WER means fewer words were substituted, omitted, or added.

Across the full benchmark, mono-1 achieved a 7.22% word error rate, the second-lowest among the systems tested. Its median raw-transcription response time was 109 ms, tied with Cartesia Ink for the fastest measured result. OpenAI’s GPT-4o Transcribe was more accurate overall at 6.68% WER, with higher measured latency.

Raw transcription results: mono-1 has 7.22% word error rate and 109 ms median latency; GPT-4o Transcribe has 6.68% and 570 ms; Cartesia Ink has 7.39% and 109 ms. Lower is better on both axes.
WildSpeech-Bench contains 1,100 English clips: 1,000 synthetic voice-cloned samples and 100 real recordings focused on prosody. Timings were collected across separate runs and environments and should be treated as indicative.

About 55% fewer estimated edits

To measure how much correction each output needed, we ran the same set of dictation cases through mono-1 and our previous API-based pipeline. A blinded AI editor estimated the number of word-level edits needed to turn each result into finished text.

mono-1 required about 55% fewer estimated edits, reducing the number of edits from 13.64 to 6.21 per 100 words.

Estimated edits per 100 words: previous API pipeline, 13.64; mono-1, 6.21.
Editing effort was estimated from model outputs using a blinded AI editor; not from actual user edits.

3× faster

Median API response time for speech recognition and formatting, compared with our previous pipeline on the same workload.

Speak naturally. Get text you can use.

We don’t speak in perfectly formed sentences. We pause, repeat ourselves, correct things midway, and give formatting instructions as we go.

mono-1 is designed to handle those patterns while producing text that reflects what you meant to write. It can also use your vocabulary, writing preferences, and available screen context to help get the details right.

Illustrative examples of intended behavior, not recorded mono-1 outputs.

Change your mind while speaking

Let’s meet at six, actually six thirty. I’ll book a table for four.

Let’s meet at 6:30. I’ll book a table for four.

Dictate a short email

Hi Maya, new paragraph, thanks for sending this over. Um, I’ve had a chance to, to look through it and everything looks good. Could you send me the final version by Friday? New paragraph, thanks, Alex.

Hi Maya,

Thanks for sending this over. I’ve had a chance to look through it and everything looks good. Could you send me the final version by Friday?

Thanks, Alex.

Turn a spoken list into something useful

Shopping list. We need milk, eggs, coffee, bananas, and some washing-up liquid.

Shopping list

  • Milk
  • Eggs
  • Coffee
  • Bananas
  • Washing-up liquid

Powering Stream, Sandbar’s voice ring

mono-1 now powers dictation in Monologue. Through our enterprise API, approved developers can bring the same focus on accurate, usable text to their own voice products.

Sandbar is already using mono-1 through the Monologue API with Stream, its voice ring.

At Sandbar, we're building Stream—a private voice ring. Thanks to mono-1, Stream customers will have access to fast, accurate, and quiet voice dictation anywhere. We're excited to work with Monologue towards a future where voice can unlock new thinking for everyone.
Mina Fahmi, CEO of Sandbar

Whether you’re writing a message or building a voice product, mono-1 helps turn what you say into text you can use.

Enterprise API access is currently limited to a small number of approved users. Tell us what you’re building by submitting the request form to be considered for access.