Today, we’re introducing mono-1, our in-house dictation model, built to turn natural speech into text you can use.
Dictation needs to handle corrections, formatting, context, and writing preferences. mono-1 is the culmination of six months of development, production testing, and feedback from Monologue users—improving how Monologue recognizes your words, handles corrections, and formats your writing.
Against our previous API-based pipeline, mono-1 required about 55% fewer estimated edits and delivered finished dictation 3× faster by median API response time.
Accurate recognition, fast responses
Accurate speech recognition is the foundation of good dictation. For Monologue, that means maintaining high accuracy and low latency across the kinds of conditions people actually dictate in.
To evaluate mono-1, we used WildSpeech-Bench, which tests speech recognition across conditions including background noise, overlapping speech, changes in loudness, pauses, and stuttering.
We measured transcription accuracy using word error rate (WER). Lower WER means fewer words were substituted, omitted, or added.
Across the full benchmark, mono-1 achieved a 7.22% word error rate, the second-lowest among the systems tested. Its median raw-transcription response time was 109 ms, tied with Cartesia Ink for the fastest measured result. OpenAI’s GPT-4o Transcribe was more accurate overall at 6.68% WER, with higher measured latency.
About 55% fewer estimated edits
To measure how much correction each output needed, we ran the same set of dictation cases through mono-1 and our previous API-based pipeline. A blinded AI editor estimated the number of word-level edits needed to turn each result into finished text.
mono-1 required about 55% fewer estimated edits, reducing the number of edits from 13.64 to 6.21 per 100 words.
3× faster
Median API response time for speech recognition and formatting, compared with our previous pipeline on the same workload.
Speak naturally. Get text you can use.
We don’t speak in perfectly formed sentences. We pause, repeat ourselves, correct things midway, and give formatting instructions as we go.
mono-1 is designed to handle those patterns while producing text that reflects what you meant to write. It can also use your vocabulary, writing preferences, and available screen context to help get the details right.
Illustrative examples of intended behavior, not recorded mono-1 outputs.
Change your mind while speaking
“Let’s meet at six, actually six thirty. I’ll book a table for four.”
Let’s meet at 6:30. I’ll book a table for four.
Dictate a short email
“Hi Maya, new paragraph, thanks for sending this over. Um, I’ve had a chance to, to look through it and everything looks good. Could you send me the final version by Friday? New paragraph, thanks, Alex.”
Hi Maya,
Thanks for sending this over. I’ve had a chance to look through it and everything looks good. Could you send me the final version by Friday?
Thanks, Alex.
Turn a spoken list into something useful
“Shopping list. We need milk, eggs, coffee, bananas, and some washing-up liquid.”
Shopping list
- Milk
- Eggs
- Coffee
- Bananas
- Washing-up liquid
Powering Stream, Sandbar’s voice ring
mono-1 now powers dictation in Monologue. Through our enterprise API, approved developers can bring the same focus on accurate, usable text to their own voice products.
Sandbar is already using mono-1 through the Monologue API with Stream, its voice ring.
At Sandbar, we're building Stream—a private voice ring. Thanks to mono-1, Stream customers will have access to fast, accurate, and quiet voice dictation anywhere. We're excited to work with Monologue towards a future where voice can unlock new thinking for everyone.
Whether you’re writing a message or building a voice product, mono-1 helps turn what you say into text you can use.
Enterprise API access is currently limited to a small number of approved users. Tell us what you’re building by submitting the request form to be considered for access.
