Independent benchmark · Measured 2026-09-17
Speech Recognition Accuracy: 95 Languages
Pikka Speech transcribes 95 languages with a median character error rate of 3.1%. OpenAI gpt-transcribe scores 5.1%, ElevenLabs Scribe v2 3.8%, Google Chirp 3 12.2% and Gemini 3.5 Transcribe 6.6% on the same audio — lower is better.
Full table
Character error rate by language
Character error rate (CER) per language, lower is better. Engines marked n/a returned no usable transcript for that language. Bold marks the best result in each row. † measured on OpenSLR (see method).
| Language | Pikka Speech | OpenAI gpt-transcribe | ElevenLabs Scribe v2 | Google Chirp 3 | Gemini 3.5 Transcribe |
|---|---|---|---|---|---|
| Afrikaans | 2.8% | 3.0% | 1.5% | 10.7% | 5.2% |
| Amharic | 3.9% | 14.9% | 10.7% | 12.3% | 5.9% |
| Arabic | 4.1% | 5.4% | 5.2% | 11.5% | 4.9% |
| Armenian | 12.3% | 14.2% | 12.5% | 21.1% | 20.4% |
| Assamese | 8.5% | 12.3% | 8.2% | 42.7% | 14.6% |
| Asturian | 7.6% | 9.4% | 4.2% | 10.8% | 16.0% |
| Azerbaijani | 0.9% | 2.5% | 0.9% | 28.0% | 2.5% |
| Basque† | 3.6% | 3.1% | 4.4% | 18.7% | 14.6% |
| Belarusian | 2.2% | 2.8% | 2.2% | n/a | 3.2% |
| Bengali | 0.9% | 3.2% | 3.3% | 10.2% | 1.9% |
| Bulgarian | 0.4% | 0.4% | 1.0% | 8.2% | 1.6% |
| Burmese | 9.9% | 41.6% | 8.7% | 7.9% | 13.5% |
| Cantonese | 9.7% | 4.8% | 22.7% | 12.7% | 14.8% |
| Catalan | 2.2% | 3.7% | 2.7% | 3.6% | 2.4% |
| Cebuano | 9.8% | 8.3% | 4.5% | n/a | 9.1% |
| Chinese | 4.9% | 6.8% | 7.6% | 12.8% | 7.9% |
| Croatian | 0.4% | 27.8% | 2.4% | 2.7% | 2.4% |
| Czech | 6.0% | 4.2% | 3.1% | 9.1% | 2.2% |
| Danish | 1.8% | 3.3% | 1.2% | 2.2% | 6.6% |
| Dutch | 1.8% | 0.4% | 0.6% | 14.2% | 2.4% |
| English | 3.1% | 1.5% | 2.2% | 9.7% | 1.2% |
| Estonian | 0.5% | 1.9% | 1.7% | 9.2% | 7.3% |
| Finnish | 0.3% | 0.6% | 0.2% | 8.8% | 1.0% |
| French | 3.9% | 2.5% | 0.7% | 16.8% | 3.4% |
| Galician | 1.1% | 1.5% | 1.6% | 6.1% | 3.9% |
| Georgian | 4.1% | 6.8% | 2.6% | 4.5% | 9.1% |
| German | 5.1% | 0.4% | 0.4% | 5.7% | 1.1% |
| Greek | 2.4% | 2.1% | 1.7% | 9.6% | 7.1% |
| Gujarati | 3.4% | 8.3% | 6.9% | 15.6% | 6.5% |
| Hausa | 5.9% | 17.4% | 6.6% | 7.4% | 9.4% |
| Hebrew | 6.7% | 6.1% | n/a | 27.0% | 14.2% |
| Hindi | 1.7% | 3.6% | 4.4% | 17.3% | 1.7% |
| Hungarian | 4.0% | 2.3% | 3.2% | 12.1% | 11.3% |
| Icelandic | 10.9% | 12.2% | 9.1% | 12.1% | 19.6% |
| Indonesian | 1.1% | 2.6% | 1.1% | 14.5% | 0.3% |
| Irish | 22.8% | 32.4% | 18.6% | n/a | 60.9% |
| Italian | 0.3% | 0.6% | 0.7% | 0.8% | 0.5% |
| Japanese | 1.8% | 3.6% | 1.1% | 0.9% | 2.7% |
| Javanese | 3.1% | 11.4% | 2.4% | 6.0% | 10.0% |
| Kannada | 1.6% | 4.6% | 7.3% | 11.8% | 11.2% |
| Kazakh | 0.9% | 4.5% | 2.9% | 15.1% | 2.4% |
| Khmer | 9.6% | 12.3% | 7.1% | 14.3% | 11.4% |
| Korean | 2.2% | 2.9% | 1.3% | 8.3% | 3.8% |
| Kurdish | n/a | 10.9% | 7.5% | n/a | 16.1% |
| Kyrgyz | 5.6% | 5.7% | 6.5% | 31.0% | 7.9% |
| Lao | 17.6% | 62.0% | 7.4% | 10.1% | 34.1% |
| Latvian | 2.4% | 4.0% | 2.7% | 13.9% | 6.3% |
| Lithuanian | 2.1% | 6.5% | 5.3% | 17.9% | 9.7% |
| Luxembourgish | 35.5% | 16.9% | 6.1% | 36.1% | 26.3% |
| Macedonian | 1.0% | 1.3% | 0.3% | 9.3% | 2.3% |
| Malay | 2.2% | 1.3% | 2.2% | 22.2% | 2.3% |
| Malayalam | 1.2% | 4.3% | 7.4% | 11.8% | 3.0% |
| Maltese | 12.0% | 6.7% | 2.3% | 12.9% | 16.0% |
| Maori | n/a | 16.3% | 6.5% | 21.2% | 18.0% |
| Marathi | 3.7% | 7.0% | 6.5% | 26.3% | 4.8% |
| Mongolian | 5.2% | 12.5% | 7.3% | 5.9% | 7.0% |
| Nepali | 0.0% | 7.0% | 7.0% | 45.3% | 4.7% |
| Northern Sotho | 13.2% | 34.7% | 6.4% | 29.2% | 23.6% |
| Norwegian | 0.7% | 1.0% | 3.4% | 4.2% | 3.8% |
| Odia | 7.8% | 10.6% | 6.8% | 17.0% | 8.2% |
| Oromo | n/a | 63.1% | 31.4% | n/a | 39.0% |
| Pashto | n/a | 24.1% | 22.9% | n/a | 46.1% |
| Persian | 1.9% | 2.3% | 4.4% | 3.2% | 2.8% |
| Polish | 1.6% | 1.0% | 1.9% | 12.1% | 3.3% |
| Portuguese | 1.0% | 2.8% | 3.4% | 5.4% | 3.7% |
| Punjabi | 4.6% | 23.9% | n/a | 22.1% | 4.5% |
| Romanian | 2.9% | 0.8% | 1.1% | 60.1% | 2.5% |
| Russian | 0.2% | 0.2% | 1.2% | 3.2% | 0.4% |
| Serbian | 15.1% | 88.1% | 30.0% | 81.3% | 40.9% |
| Sesotho† | 11.0% | 19.7% | 1.5% | n/a | 28.7% |
| Setswana† | 6.8% | n/a | 3.8% | n/a | 28.7% |
| Sindhi | 4.2% | 98.5% | 5.3% | n/a | 8.9% |
| Sinhala† | 15.3% | 48.8% | 2.0% | n/a | 43.5% |
| Slovak | 1.2% | 1.8% | 1.2% | 5.0% | 3.5% |
| Slovenian | 2.0% | 3.5% | 1.7% | 23.5% | 8.0% |
| Somali | n/a | 24.9% | 12.4% | n/a | 17.6% |
| Spanish | 1.2% | 0.4% | 0.5% | 2.1% | 0.8% |
| Swahili | 1.9% | 5.3% | 1.1% | 7.8% | 3.0% |
| Swedish | 4.5% | 3.8% | 2.4% | 4.0% | 6.0% |
| Tagalog | 2.3% | 4.9% | 3.1% | 8.0% | 5.3% |
| Tajik | 2.9% | 58.3% | 2.4% | n/a | 10.0% |
| Tamil | 16.8% | 18.6% | 18.3% | 28.8% | 16.5% |
| Telugu | 1.4% | 3.4% | 6.3% | 23.0% | 1.4% |
| Thai | 4.0% | 4.9% | 4.6% | 29.5% | 5.0% |
| Turkish | 0.7% | 1.0% | n/a | 44.9% | 0.9% |
| Ukrainian | 1.2% | 3.3% | 2.6% | 14.7% | 1.1% |
| Urdu | 4.7% | 90.5% | 14.7% | 29.2% | 11.4% |
| Uzbek | 7.7% | 9.2% | 7.0% | 14.1% | 12.7% |
| Vietnamese | 1.1% | 3.6% | 1.6% | 12.2% | 4.2% |
| Welsh | 5.7% | 13.2% | 4.8% | 25.8% | 21.2% |
| Wolof | n/a | 19.6% | 13.2% | 25.7% | 29.5% |
| Xhosa | 49.0% | 33.7% | 6.5% | 44.8% | 12.1% |
| Yoruba | n/a | 34.6% | 18.7% | 27.1% | 25.6% |
| Zulu | n/a | 22.9% | 7.2% | 10.3% | 14.2% |
| pt-BR | 1.0% | 2.8% | n/a | n/a | 3.7% |
Method
How we measure
A short, reproducible method — no vendor-run benchmarks, no cherry-picked clips.
Every engine transcribes the same FLEURS clips — the open multilingual speech corpus from Google used across speech-recognition research — through its own official API.
- 95 languages, 10 utterances each, native speakers reading real sentences.
- Scored as character error rate after normalising case and punctuation, so differences reflect recognition, not formatting.
- Word error rate (WER) is reported in the downloadable data for languages that use word spacing; character-based scripts (Chinese, Japanese, Thai, Khmer, Burmese, Lao) are scored on characters only.
- Every row is published, including the languages where another model leads.
- Sinhala, Basque, Sesotho and Setswana are measured on OpenSLR corpora — SLR52, SLR76 and SLR32, all CC BY-SA 4.0 — because FLEURS has no audio for them. Every other language here is FLEURS.
- Not yet included: Akan, Bosnian, Faroese, Frisian, Quechua, Romansh, Kinyarwanda, Albanian and Turkmen — no openly licensed benchmark corpus covers them yet.
Read the corpus paper for the recording methodology: FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.
FAQ
Questions about this benchmark
What is character error rate (CER)?
CER is the share of characters that differ from the reference transcript, as a percentage. A CER of 2% means the engine transcribed 98 of every 100 characters correctly. Lower is better.
What audio was used for this benchmark?
FLEURS, the open multilingual speech corpus from Google, used across speech-recognition research. We ran 10 utterances per language across 91 FLEURS languages — native speakers reading real sentences, identical clips for every engine.
Did every model get the same audio and settings?
Yes for the audio: every engine transcribed the exact same clips through its own official API. Each engine was measured in its default configuration; results were scored after normalising case and punctuation so differences reflect recognition, not formatting.
Why do some cells show n/a?
n/a means that engine returned no usable transcript for that language on the day of the run — usually because the model does not support it. We publish those gaps instead of hiding them.
Are all Pikka Speech languages in this benchmark?
95 languages are measured here — 91 on FLEURS and, because FLEURS has no audio for them, Sinhala, Basque, Sesotho and Setswana on OpenSLR corpora (CC BY-SA 4.0). 9 further offered languages — Akan, Bosnian, Faroese, Frisian, Quechua, Romansh, Kinyarwanda, Albanian and Turkmen — have no openly licensed benchmark corpus yet; they will be added as corpora become available.
How often is this table updated?
It is re-run whenever a model we use or benchmark changes. This version was measured on 17 September 2026.
Keep exploring
Which languages Pikka Speech covers: language coverage. How interpretation and captions work together: platform capabilities and hybrid simultaneous interpretation. Or start a free test room.
Hear the difference on your own event
Create a free test room, speak a few sentences in your language, and compare the live captions against any other engine.