Non-native English

Your accent is notthe problem.

Most people who speak English as a second language assume dictation is not for them. Modern speech to text handles accented English far better than the reputation suggests. What genuinely breaks is narrower than that, and most of it is fixable in five minutes.

The short answer

Yes, dictation works when English is not your first language. Today's models are trained on very varied audio and transcribe a Brazilian, Portuguese, Indian, German or Nigerian speaker of English with no special setup. What still fails is predictable: proper nouns, company and product names, technical jargon, and sentences that mix your own language into the English. Fix those four and dictation is faster than your typing.

Download for MacFree while in betamacOS · Windows next
01What breaks

The real failure list

Six things that break, and what to do about each.

Not one of these is about your pronunciation. They are the specific places speech to text loses, and every one of them has a workaround that takes seconds.

The problemWhat you getWhat to do
Names of people and placesYour colleague Thiago comes back as Diego. Your city becomes whichever English word sounds closest.Let it be wrong while you speak, then fix the names once at the end. Stopping mid-sentence to repeat a name costs more than the correction does.
Company and product namesSonira, Figma, Vercel arrive as ordinary words, or split into two.Same fix, at the end. If you say one product name twenty times a day, add a text replacement in the OS so the typo corrects itself.
Technical jargon and acronymsKubernetes survives. Your team's internal acronym does not, and spelled-out letters run together into one word.Say acronyms with a small pause between letters. Accept that internal vocabulary, published nowhere, will not be recognised by anything.
Your own language in the middleYou drop one Portuguese or Spanish word into an English sentence and it returns as an English word that sounds similar.Use a tool that does not force one language per dictation. If yours has a language setting, it will argue with you on every mixed sentence.
Hesitation transcribed verbatimUm, so, I think, sorry, I mean. On screen it reads worse than what you would have typed, which is why people quit on day two.Only a cleanup pass fixes this. Everyone hesitates more in a second language, so raw transcription is unfair to you in a way it is not to a native speaker.
The tool rewriting your EnglishYour words come back smoother, longer, and less like you. Sometimes the meaning has moved.Decide whether you want a transcriber or a ghostwriter. They are different products, and only one of them lets you keep learning.

Notice what is not on that list. Vowels, stress, rhythm, the things people are self-conscious about. That is precisely the part the models got good at.

02Causes

Why it happens

The model is guessing words, not hearing letters.

It learned many accents, and some far more than others

Large speech models are trained on enormous, varied audio, which is the reason accented English works at all now. The distribution is still uneven. An accent that appears often in the training audio is recognised more reliably than one that appears rarely, which is why the same tool can feel excellent to you and mediocre to the colleague sitting next to you.

Recognition is prediction

Speech to text does not decode sounds one at a time. It predicts the most likely sequence of words given the audio and everything said so far. Common phrasing survives noise and unfamiliar pronunciation because context holds it up. A name the model has rarely seen has no such support, so it loses to whatever ordinary word sounds closest. That is why your sentences land and your names do not.

The language is decided per dictation, not per sentence

Most tools settle on one language for the whole dictation, either from a setting or from the first seconds of audio. A sentence carrying two languages sits outside that model of the world, so the minority language gets pushed into the majority one. For anyone who works in English and lives in another language, this is the daily failure. The accent is not.

The cleanup pass has opinions

Anything that tidies a raw transcript is a language model, and language models drift toward fluent, standard, faintly formal prose. Point one at a second-language speaker's sentence and it will quietly upgrade it. The output reads better and is no longer yours, which matters more than it sounds when the text is a message from you to a colleague.

The first two belong to whoever trains the acoustic model, and no small app changes them. The last two are product decisions. Those are the ones worth choosing carefully, and the ones a buyer can actually judge.

03Five minutes

The fix

How to make dictation work in your English.

This works on any dictation tool, including the free one built into your Mac. Nothing here asks you to sound more like a native speaker, because that is not the variable that decides the outcome.

  1. 01

    Speak in phrases, not in words

    The model uses context to recover unclear sounds, so a whole phrase is recognised better than the same words said one at a time. Careful, isolated pronunciation is the instinct of a second-language speaker and it makes the result worse. Say the complete thought at your normal pace.

  2. 02

    Say your punctuation out loud

    Period, comma, question mark, new paragraph. Two minutes of practice and it is automatic. This single habit does more for the quality of the output than anything to do with your accent, because it removes the sentence boundaries the tool would otherwise have to guess from your intonation.

  3. 03

    Let the names be wrong until the end

    Every stop to correct a name costs a restart and kills the sentence you were building. Dictate the whole message, then fix the two proper nouns. Second-language speakers stop and correct far more often than native speakers do, and that habit, not the accuracy, is what makes dictation feel slow.

  4. 04

    Test with a sentence from your actual work

    Not the quick brown fox. Take a real sentence, with your product names, your acronyms, and the word from your own language you always end up using. That test predicts whether you will still be dictating next week. A generic sentence predicts nothing.

  5. 05

    Read it back before you send it

    Almost every second-language speaker reads better than they write, so a quick read catches what you would have missed while composing. If your tool can read the text aloud, listening catches a different set again: mostly the wrong word that is entirely plausible in context, which is exactly the error dictation produces.

Five habits, five minutes. If a tool still fails step four after all of them, the tool is the problem, not the accent.

04Sonira

What we built

No language setting, on purpose.

Sonira never asks which language you are about to speak. There is no dropdown to set before dictating and nothing pinned for the session, so an English sentence with two Portuguese words inside it comes back as an English sentence with two Portuguese words inside it. That is not a mode you switch on. It is the only behaviour there is.

The pass that punctuates and tidies the transcript is forbidden in writing from translating, from changing your wording, and from converting your register. It takes out hesitation, false starts and the words you retracted, and then it stops. What lands in your document is your sentence, punctuated. Not a fluent stranger's version of your sentence.

What this does not do is fix the acoustic model. Recognition of accented English is the industry's problem and we buy it like everyone else buys it. There is no custom vocabulary yet either, so a name the model has never seen is still wrong the first time, and you still fix it by hand.

Honest note: the current public build is the dictation core. Read-aloud ships in the next public build, which is waiting on a signing step, not on the code.

Hold a key. Speak. The words land where the cursor is.

05Your words

Transcriber, not ghostwriter

Your English, not a smoother stranger's.

There is a real temptation, when you write in a second language, to let a tool improve the sentence on the way out. It works, and it costs something. Text that has been rewritten stops being evidence of what you can do, and after a year of it you write no better than the day you started.

A tool that only transcribes gives you the opposite deal. The words are yours, the mistakes are yours to find, and the speed you gain is simply the speed of not typing. If you do want a rewrite, that should be a separate, deliberate action, not something that happens silently to every message you dictate.

The read-back trick

Reading is the strongest skill most second-language speakers have, and listening is a close second. Having a draft read back to you catches the plausible-but-wrong word your eyes skate over, which is the exact error class dictation produces. It is why we built voice out as well as voice in.

06FAQ

Good to know

Questions, answered.

01Does speech to text work if I have a strong accent?
Usually, yes. Current models are trained on very varied audio and handle strong accents in English far better than the tools people remember from a few years ago. Accuracy does vary by accent, and no honest tool will hand you a number for yours. Test it with one real sentence from your own work and you will know inside a minute.
02Will it understand me if I mix two languages in one sentence?
That depends entirely on the tool. Anything that pins one language per dictation will force your other language into it. Sonira pins nothing and leaves each part in the language it was spoken in. If you work in English and live in another language, this is the feature that decides whether dictation fits your actual day.
03Why does it get every word right except the names?
Because recognition is prediction. Common words are held up by context, while a name the model has rarely seen is not, so it loses to whatever ordinary word sounds closest. Fix names at the end of the dictation rather than mid-sentence, and the cost drops to a few seconds.
04Does dictation make my English look worse than typing?
Raw transcription does, because it prints your hesitation. Everyone hesitates more in a second language, so it lands harder on you than on a native speaker. A cleanup pass that removes filler and false starts without rewriting your words closes that gap, and that is what Sonira's second pass is for.
05Should I use a tool that improves my English while it transcribes?
Only if you have decided you want that. It is a different product from dictation and it carries a cost: your text stops reflecting your own level, to you and to everyone reading it. We do not do it by default, and we think the choice should be yours and explicit.
06Does it work in my own language too, not only English?
Yes. Sonira has no language setting, so any language the model handles works, and Portuguese in particular is treated as first-class rather than one line in a list of a hundred. If that is your case, there is a separate guide on European versus Brazilian Portuguese.

Dictate in the English you actually speak.

Free while we're in beta.

Download for MacmacOS · Windows next

Also dictating in Portuguese? Read the language guide.

New to dictation? Start with the Mac guide.

Coming from Wispr Flow? Read the honest comparison.

Back to the Sonira home page

SoniraEnd