Languages

99 languages. 16 we can prove.

Every offline dictation tool built on an open speech model can list a hundred languages, because the model lists a hundred languages. Almost nobody tests them. We tested ours, and 16 passed on recordings of real native speakers; the error rate for every one of them is below. Nine more passed our synthetic tests and then missed on real speakers. They are named on this page instead of quietly dropped.

How this works in the app

The Language section of InkBeepAI Settings: a dictation language picker set to English, checkboxes for automatic detection, using the large model for English, and using the graphics card, and a line reading "English is verified: 7.8% word error, measured on recordings of native speakers."
The app quotes the same figure this page does — for whichever language you pick.
The InkBeepAI language drop-down open, showing verified languages such as Russian, Turkish, Ukrainian, Catalan, Vietnamese, Dutch, French, Indonesian and Finnish above a separator reading "Supported, not verified", with Afrikaans, Albanian, Amharic, Arabic, Armenian and Assamese below it.
Verified languages first, then the line, then everything else.
  • You choose the language. It stays chosen. Pick it once in Settings, or switch it in two clicks from the tray menu. The language you picked is named in the tray menu and on the tray tooltip, so you can never be dictating into a language you did not intend — and while automatic detection is still deciding, both read “Detecting language”.
  • Nothing is mixed. One language is active at a time. InkBeepAI does not try to guess per sentence, because guessing needs several seconds of speech and a dictated phrase does not have them. We measured that: on dictations of 3 to 15 seconds, guessing cost nothing in 13 of 14 languages; on short phrases it made 10 of 14 worse, and Hindi was misread as Urdu often enough to ruin a session.
  • Automatic detection is there if you want it, off unless you turn it on, and it works the safe way round: it keeps using the language you picked until one dictation has at least four seconds of actual speech, identifies the language from that one, switches to it, and then stays there. The tray menu shows what it picked.
  • It never translates. You speak Spanish, you get Spanish text. Dictation, not interpretation.
  • Recognition is still entirely on your PC. Both speech models ship inside the download. Adding 98 languages did not add a single network request — the licence check is the same one it always was.

The 16 we verified

The figure is the error rate: out of every hundred words spoken, how many came back wrong (for Japanese, written without spaces, out of every hundred characters). Lower is better and zero is perfect. It counts a missing word, an extra word and a wrong word alike, and it is measured against the reference transcript that came with the recordings — including their own conventions for numbers and names, which we did not tune for.

Measured on recordings of native speakers

LanguageError rateTested against
Spanish español1.9%Native speakers, Google FLEURS
Italian italiano3.7%Native speakers, Google FLEURS
Portuguese português4.2%Native speakers, Google FLEURS
Japanese 日本語4.8%Native speakers, Google FLEURS · per character, because Japanese is written without spaces
German Deutsch4.9%Native speakers, Google FLEURS
Polish polski7.2%Native speakers, Google FLEURS
English7.8%Native speakers, Google FLEURS
Russian русский7.8%Native speakers, Google FLEURS
Turkish Türkçe8.0%Native speakers, Google FLEURS
Ukrainian українська8.1%Native speakers, Google FLEURS
Catalan català8.6%Native speakers, Google FLEURS
Vietnamese Tiếng Việt8.6%Native speakers, Google FLEURS
Dutch Nederlands8.8%Native speakers, Google FLEURS
French français8.9%Native speakers, Google FLEURS
Indonesian Indonesia9.1%Native speakers, Google FLEURS
Finnish suomi9.3%Native speakers, Google FLEURS

English’s 7.8% is for the small English-only model the app runs by default, because that is what an English user actually gets. The same recordings through the large multilingual model score 5.5%, and you can switch English to it in Settings. The fast one is the default deliberately: the gap in speed is far larger than the gap in accuracy, and dictation lives or dies on the wait.

Nine passed on synthetic speech and failed on real speakers

Every language below cleared our synthetic tests, several of them with almost no errors. On recordings of actual native speakers they missed the bar — some on the average, some because too many individual sentences came back wrong. They are still in the app, because they run and some people will want them; the app tells you they missed when you pick one.

LanguageSynthetic speechNative speakersSentences under 15%
Romanian română2.2%11.4%65%
Swedish svenska9.1%11.6%60%
Slovak slovenčina0.0%13.1%65%
Hindi हिन्दी3.4%14.1%50%
Czech čeština4.6%15.8%50%
Danish dansk7.0%16.5%60%
Hungarian magyar1.7%16.5%50%
Estonian eesti4.1%17.2%45%
Latvian latviešu3.8%20.5%45%

To be verified, a language needs an average under 15% on native speakers and at least 70% of its sentences individually under 15%.

That gap is the entire argument for the human test. Slovak scored 0.0% on synthetic speech and 13.1% on real speakers, with 7 of its 20 sentences outside the bar. It is why we do not simply repeat the model's own language list back to you as a feature.

How it was tested

Four passes, all on one machine, using the same speech engine and the same settings the shipped app uses.

PassAudioWhat it was for
Clean641 clips across 69 languages: up to ten native-written sentences each, read by synthetic voicesA wide first sweep, to find which languages were worth testing properly
Noisy358 of those clips, covering 37 languages, with microphone band-limiting and room noise mixed in at a measured levelDictation happens in real rooms, not studios
Very noisyThe same 37 languages, with the noise raised until the speech was only just dominantThe point at which a language stops being usable rather than merely worse
Human500 clips of native speakers reading, 20 for each of 25 languages, from Google’s FLEURS set, with the transcripts that came with themThe pass that decides. Only a language that clears it is verified, and its result is the figure we publish

A language has to be good on average and good sentence by sentence, so one lucky sentence cannot carry it. On native speakers that means an average under 15% and at least 70% of sentences under 15%. The synthetic sentences were chosen without digits, because “2026” against “twenty twenty-six” scores as an error without anybody having misheard anything. The native-speaker transcripts are used exactly as they came, so there a number the model writes as digits does count against it.

What this costs you

Said plainly, because you will find out anyway:

  • The download is much bigger. Two speech models ship inside it — a small English one and a large multilingual one. It is a one-time download and it is what keeps everything on your own machine.
  • Languages other than English want a graphics processor. The graphics built into the processor counts. We have measured speed on one laptop, with Intel Iris Xe graphics; older or weaker graphics will be slower. Without a usable one, the multilingual model runs on the processor alone and is far too slow to dictate with. InkBeepAI times the larger model on your own PC when it loads it, and if it is far too slow it says so once and offers the way out that applies — turning the graphics processor on, or switching back to English. It does not let you find out by waiting.
  • English does not need a graphics processor. It keeps the same small English model it has always used. On a PC with a graphics processor it now runs on that; without one it runs on the processor, as it always did.
  • 83 languages are not verified. They run, and the app tells you when you pick one. We would rather offer them with that caveat than pretend to a hundred verified languages.

The other 74

Like the nine above, none of these is verified. They run; the app labels them as not verified when you pick one.

Did not clear our synthetic screen (44)

Tested on synthetic speech: most came in outside the bar, some of them far outside, and two cleared the quiet pass and fell over once we added room noise. A few of these results are partly scoring artefacts — Chinese output mixes simplified and traditional characters, and Hebrew spelling varies — so read this as “not proven”, not “broken”.

Afrikaans, Albanian, Amharic, Arabic, Azerbaijani, Bengali, Bosnian, Bulgarian, Burmese, Chinese, Croatian, Galician, Georgian, Greek, Gujarati, Hebrew, Icelandic, Kannada, Kazakh, Khmer, Korean, Lao, Lithuanian, Macedonian, Malay, Malayalam, Maltese, Marathi, Mongolian, Nepali, Pashto, Persian, Serbian, Sinhala, Slovenian, Somali, Sundanese, Swahili, Tamil, Telugu, Thai, Urdu, Uzbek, Welsh.

Not in our tests (30)

The model accepts these, and we have no measurement to give you either way.

Armenian, Assamese, Bashkir, Basque, Belarusian, Breton, Faroese, Haitian Creole, Hausa, Hawaiian, Javanese, Latin, Lingala, Luxembourgish, Malagasy, Māori, Norwegian, Norwegian Nynorsk, Occitan, Punjabi, Sanskrit, Shona, Sindhi, Tagalog, Tajik, Tatar, Tibetan, Turkmen, Yiddish, Yoruba.

Common questions

Do I have to download anything extra for another language?

No. Both speech models are inside the installer. Choose a language in Settings or from the tray menu and it is ready — no account, no download, no network request.

Can it switch languages automatically?

It can, and it is off unless you turn it on. When it is on, InkBeepAI keeps dictating in the language you picked until one dictation has at least four seconds of actual speech, identifies the language from that one, switches to it, and stays on it. It deliberately does not re-decide on every sentence: a two-second phrase is not enough to identify a language reliably, and getting that wrong turns speech into confident nonsense with no error message.

Does it translate what I say?

No. It writes down what you said in the language you said it in.

My language is not verified. Should I buy?

Only after trying it, inside the refund window. It will run, and it may be good enough for your work — but we could not verify it, and this page says whether we measured it falling short or never measured it at all. We would rather say that than invent a number. There is a 30-day refund if it is not good enough.

Why do other tools claim more languages?

Because the open speech model underneath most offline dictation tools, including this one, advertises 99. Repeating that number is free. Our list is the same list — the difference is that 16 of ours went through four rounds of testing on this machine and passed the round that decides — recordings of native speakers — and the results, including the 9 that failed on real speakers, are on this page.

Is my audio still private?

Unchanged. Recognition runs on your own PC in every language. InkBeepAI talks to one server, its licence server: once when you activate your key, and then a licence check every two weeks, each carrying your licence key and up to three anonymous hardware IDs — never audio, never text. How that is built.

Dictate in your own language

One-time purchase, 30-day refund, no subscription, no account.