How Accurate Is Phone Call Translation? 4 Systems Tested (2026)
Published 2026-08-18 · By Ron Villomo, Founder of LiveLingo · LiveLingo is one of the systems measured; full method and raw data
Quick Answer: How accurate is phone call translation?
In our August 2026 test, LiveLingo scored highest at 96.9% on phone-line audio. Across 120 English utterances translated into Spanish, Japanese, Chinese and German, LiveLingo reached 4.85 out of 5 (96.9%), Google Cloud Speech-to-Text v2 with Translate v3 reached 4.70 (94.0%), Azure Speech Translation 4.59 (91.8%) and a Whisper-large plus GPT-4o-mini baseline 4.44 (88.9%). On harder language pairs those fall to 87.0%, 77.7%, 70.0% and 66.7%. The phone line itself costs almost nothing: the same systems on wideband microphone audio scored within 0.04 points of their phone numbers. Scoring was by three independent frontier LLM judges on a fixed studio-grade corpus carried over a G.711 phone channel, downloadable below.
1. How accurate is phone call translation? The measured answer
We scored four systems on identical audio passed through a simulated phone leg. Each translation was rated 0-5 for comprehension fidelity by three independent frontier LLM judges, and the score below is the mean of the three, shown also as a percentage of the maximum. The rubric penalises wrong-word substitution, dropped entities or numbers, and language-script mixing; it does not penalise style or valid synonyms.
| System | Phone-line score | Accuracy | Harder pairs |
|---|---|---|---|
| LiveLingo | 4.85 / 5 | 96.9% | 87.0% |
| Google Cloud STT v2 + Translate v3 | 4.70 / 5 | 94.0% | 77.7% |
| Azure Speech Translation | 4.59 / 5 | 91.8% | 70.0% |
| Whisper-large + GPT-4o-mini | 4.44 / 5 | 88.9% | 66.7% |
n=120 utterances, English into Spanish, Japanese, Chinese and German. “Harder pairs” is a separate 160-utterance set covering Arabic, Turkish, Vietnamese, Polish, Portuguese, Indonesian, Malay and Mandarin sources into European and Asian targets. Measured August 2026 on a fixed studio-grade corpus carried over a G.711 phone channel, downloadable above; see the full method and raw data.
Download the test audio and results
Both arms of the 30-utterance English set, plus every score in machine-readable form. The two archives contain the same utterances; the only difference is the phone channel. Released under CC BY 4.0, so anyone can re-run these systems and check the numbers.
audio-wideband.zip (3.1 MB, 16 kHz) · audio-phoneline.zip (2.6 MB, G.711) · results.json
The spread between language pairs matters more than the spread between the top systems. Even the strongest system loses three points moving from German to Japanese, and every system loses far more on the harder corridor pairs than it loses to the phone line.
| Language pair (best system, phone-line) | Score | Accuracy |
|---|---|---|
| English → German | 4.92 / 5 | 98.4% |
| English → Spanish | 4.86 / 5 | 97.1% |
| English → Chinese | 4.84 / 5 | 96.9% |
| English → Japanese | 4.77 / 5 | 95.3% |
Tap and speak in English
Tap to start
2. Does Apple Live Translation work on phone calls?
Apple Live Translation handles phone calls on recent iPhones and Samsung Galaxy AI offers a comparable in-call feature. Both are good at what they do, and for an iPhone-to-iPhone call in a supported language they are the most convenient option available because there is nothing to install.
Their limit is coverage, not quality. Both require a recent handset from that specific manufacturer, both are restricted to the language sets those makers ship, and critically both put the capability on your device rather than on the line. That is fine for calling a friend with the same phone. It does not help when you need to reach a clinic switchboard, a landlord, a hotel desk, or a supplier whose office phone is a decade old. Those calls are the reason most people go looking for a translated-call product in the first place.
3. Why is phone call audio quality bad, and does it reduce accuracy?
A phone call is encoded with ITU-T G.711, which samples at 8 kHz and carries roughly 300 Hz to 3400 Hz. The discarded band is not empty. Acoustic phonetics work from the University of Oxford's phonetics laboratory puts the energy peaks of /s/ at roughly 4500 Hz and 7500 Hz, entirely above the telephone ceiling, which is also why band-limited /s/ is systematically misheard as /sh/. The physics predicts phone audio should be harder. ITU-T G.711 · University of Oxford Phonetics Laboratory
To test whether that prediction survives to the output, we produced two versions of every clip. One is the original 16 kHz recording. The other is the same file resampled to 8 kHz with anti-alias filtering, quantised with G.711 mu-law, then decoded back into the original 16 kHz container, so both files are identical in format and no system can branch on sample rate. Energy above 3.4 kHz falls from 14-22% of the signal to about 3%, while RMS loudness is unchanged, so the comparison measures bandwidth and not volume.
Every system landed within 0.04 points of itself. LiveLingo scored 97.2% on microphone audio and 96.9% over the phone line; Google 93.4% and 94.0%; Azure 91.7% and 91.8%; the Whisper baseline 88.1% and 88.9%. 88% of individual cells received an identical score in both arms. Telephone bandwidth destroys real information and modern translation absorbs nearly all of it.
4. Recognition drifts about 2%, and robustness varies 5x between systems
Quality holding steady does not mean nothing happened underneath. Comparing the source transcript each system produced from the wideband file against the one it produced from the phone-line file, Google Cloud Speech-to-Text v2 changed 0.4% of recognised words and the Whisper-large baseline changed 2.0%, with 96% and 81% of transcripts respectively coming back completely unchanged.
Raw comparison overstates this badly, and the reason is worth stating because it is how this kind of number goes wrong. Whisper hallucinates stock phrases onto trailing silence, emitting fixed subtitle-style boilerplate unpredictably in either arm. Chinese output flips between Simplified and Traditional script at random. Arabic differs mostly by hamza spelling. Before removing those, Whisper's apparent drift measured 5.7%; after, it is 2.0%.
The useful finding is that narrowband robustness differs by about 5x between recognisers, which is a far larger effect than the phone line itself. Arabic and Indonesian drifted most at around 5%, consistent with the fricative prediction, and even there the fidelity score did not move.
5. Does AI translation get worse on phone calls than in person?
Because the evidence usually cited for it measures something else. In the Whisper paper (Radford et al., 2022), wav2vec 2.0 Large and Whisper Large V2 both score an identical 2.7% word error rate on LibriSpeech clean read speech. On the CallHome telephone corpus they diverge to 34.8% and 17.6%. That gap is real, large, and routinely offered as proof that phone audio is the problem. Radford et al., 2022 · LDC CallHome
But CallHome is not merely narrowband. It is spontaneous, overlapping, unscripted conversation between people who know each other well, dense with contractions, false starts and crosstalk. LibriSpeech is one person reading a book aloud. Changing two variables at once cannot tell you which one mattered. Isolating the bandwidth finds almost nothing, which means the CallHome penalty is mostly the other variable.
That is the more useful conclusion, because those are the factors you can actually control. Taking turns cleanly, moving somewhere quieter, and letting each speaker finish will do more for a translated call than any concern about line quality.
6. What actually separates phone call translation apps
Since bandwidth costs nothing measurable and the top systems sit within three points of each other on easy pairs, the differences that matter are elsewhere: how far quality falls on your specific language pair, whether the product can reach the number you need to call, and how the call handles turn-taking.
Language pair is the biggest quality lever by a wide margin, so check your own corridor rather than a headline average. On reach, LiveLingo places the translated call from the app to any mobile or landline number, so the recipient answers an ordinary phone call and installs nothing. And on conduct, it is worth knowing whether a product announces itself honestly or clones your voice to hide that a translator is involved. livelingo.io/phone-call-translation
7. What this test covers, and what it does not
The numbers above are bounded by four specific design choices. Stating them is not a hedge on the results; it is the scope within which the results hold.
| Dimension | What was tested |
|---|---|
| Source speech | A fixed studio-grade corpus, chosen so both arms receive acoustically identical input with zero speaker variation. Both arms are published above, so every number here can be re-derived from the same audio we used. The clips are more clearly articulated than spontaneous conversation, which makes them the most favourable case for the bandwidth result in section 3; corpus construction is documented in full on the method page. |
| Judging | Three frontier LLM judges, not human raters. 66% of wideband cells scored a perfect 5.0, so the rubric measures whether meaning survived rather than fine-grained quality. |
| Channel | G.711 band-limiting and mu-law quantisation only. Packet loss, jitter, automatic gain control and noise suppression are not modelled, and all occur on real calls. |
| Independence | LiveLingo is our product and placed first on both corpora. One measurement per arm, so run-to-run judge variance is not averaged out. |
The practical reading: the ranking in section 1 is a clean like-for-like comparison, since every system received identical audio and identical judging. The bandwidth result in section 3 is the claim most sensitive to the corpus choice, and replication on spontaneous conversational speech is the next step. Full derivations, statistical power and the measurement artifacts we removed are on the full method and raw data.
Frequently asked questions
How accurate is phone call translation?
In our August 2026 test, the best system scored 4.85 out of 5 (96.9%) on phone-line audio across 120 English utterances translated into Spanish, Japanese, Chinese and German, scored by three independent frontier LLM judges on a comprehension-fidelity rubric. Google Cloud Speech-to-Text v2 with Translate v3 scored 4.70 (94.0%), Azure Speech Translation 4.59 (91.8%), and a Whisper-large plus GPT-4o-mini baseline 4.44 (88.9%). On harder language pairs such as Arabic, Turkish and Vietnamese into European and Asian targets, the same systems dropped to 87.0%, 77.7%, 70.0% and 66.7% respectively.
Does a phone line make translation less accurate than in person?
Almost not at all. We scored the same utterances twice, once as wideband microphone audio and once after passing the identical recording through a simulated G.711 phone leg, and comprehension fidelity moved by less than 0.04 points on a 0-5 scale for every system tested. LiveLingo scored 97.2% on microphone audio and 96.9% over the phone line. Telephone bandwidth is a real loss of information, but modern translation absorbs nearly all of it.
Why is phone call audio quality bad?
A standard phone call uses the ITU-T G.711 codec, which carries only 300 Hz to 3400 Hz and samples at 8 kHz. Human speech carries useful information well above that: the /s/ sound has energy peaks around 4500 Hz and 7500 Hz, entirely above the telephone ceiling. That is why phone audio sounds muffled and why band-limited /s/ is often misheard as /sh/. It is a genuine loss, it is just a much smaller problem for modern speech systems than it used to be.
Does Apple Live Translation work on phone calls?
Yes, on recent iPhones, and Samsung Galaxy AI offers a comparable in-call feature. Both are limited to recent handsets from that manufacturer and to the languages those makers support, and both require the feature on your own device rather than on the line, so neither helps when you need to call an office switchboard, a landline, or someone on an older phone. LiveLingo places the translated call from the app to any mobile or landline number and the recipient installs nothing.
Which translation system handles phone audio best?
Two things separate them. On translation quality over a phone line, LiveLingo scored 96.9%, Google 94.0%, Azure 91.8% and a Whisper-large baseline 88.9% across 120 utterances. On raw recognition robustness the spread is wider: when the same clip went through a phone leg, Google Cloud Speech-to-Text v2 changed only 0.4% of recognised words while the Whisper baseline changed 2.0%, a 5x difference. Published research shows the same pattern at scale: on the CallHome telephone corpus, wav2vec 2.0 Large scores 34.8% word error rate against Whisper Large V2 at 17.6%, even though both score an identical 2.7% on clean read speech.
Does a bad phone connection make translation worse?
Bandwidth alone does not, but packet loss, jitter and background noise are separate problems and they do. Our test isolated the frequency range, holding loudness constant, so it measures the phone band and nothing else. A call with dropouts or heavy street noise degrades translation for the same reasons it degrades a human listener, and no codec quality compensates for speech the microphone never captured cleanly.
Full method, per-language results, statistical power, limitations and the measurement artifacts we removed are documented at livelingo.io/research/phone-line-bandwidth-method.