NativeCall quality report · August 23, 2026 experiment
Measured call quality: Japanese recognition results and customer reviews
On the same public Japanese test clips, candidate configuration D recorded a 2.61% character error rate, compared with 15.29% for historical control C. This measured result under telephone-audio conditions gives us data for evaluating call quality.
Historical control C
15.29% CER
Character error rate; lower is better
Candidate D
2.61% CER
Character error rate; lower is better
Measurement scope: August 23, 2026; 5 unique Japanese clips, each repeated 3 times, with an 8 kHz telephone-codec simulation. This is a historical speech-recognition component benchmark; the full method and measurement boundary are below.
Read all four languages’ results and methodology ↓Historical call playback logs · August 29–September 1, 2026
Time until translation playback starts
Toward the app caller · median
1.36 seconds
Toward the app caller · p95
4.276 seconds
These 33 turns have both timestamps, measured from the recorded end of speech to the server’s translation playback-start event. The measurement covers the translation and playback path after recognition, making it closer to the call flow than recognition timing alone.
Measurement boundary and sample details
The previously selected sample contains 10 calls and 115 retained timing turns: 34 have this measurement and 81 lack one or both timestamps. Thirty-three measured turns are toward the app caller. There is one recipient-direction turn, at 1.315 seconds, so no percentile is calculated for that direction. The event ends at server playback start; it excludes later device, network and acoustic-arrival time and does not measure translation accuracy.
Test configurations use anonymous group labels. Provider and model names are not published; these labels distinguish historical experiment configurations.
Historical experiment: all tested results
Each language has only 5 unique public FLEURS clips, each repeated 3 times: 15 trials per model and language. Repetitions do not create 15 independent clips.
The candidate did not improve every language: it had lower error rates on these Japanese and English clips, but higher error rates on Mandarin and Cantonese. All eight result rows are included below; the best numbers from different models are not combined into a single service score.
| Language / test group | Error rate | Completed trials | ASR p50 | ASR p95 |
|---|---|---|---|---|
| English · Historical control A | 7.50% WER | 15/15 | 279 ms | 411 ms |
| English · Candidate D | 6.67% WER | 15/15 | 484 ms | 664 ms |
| Mandarin · Historical control B | 5.52% CER | 15/15 | 241 ms | 330 ms |
| Mandarin · Candidate D | 6.21% CER | 15/15 | 277 ms | 625 ms |
| Cantonese · Historical control B | 2.34% CER | 15/15 | 225 ms | 323 ms |
| Cantonese · Candidate D | 3.13% CER | 15/15 | 193 ms | 489 ms |
| Japanese · Historical control C | 15.29% CER | 15/15 | Not comparable | Not comparable |
| Japanese · Candidate D | 2.61% CER | 15/15 | 208 ms | 433 ms |
Method and measurement scope
WER is word error rate; CER is character error rate. Both divide the total insertions, deletions and substitutions by the number of reference words or characters; lower is better. English uses words; Japanese, Mandarin and Cantonese use characters. Different languages and units are not directly comparable. Subtracting error rate from 100% does not produce a translation-accuracy score.
Scoring normalizes Unicode, punctuation, Traditional/Simplified Chinese and spoken/Arabic numeral formatting. Some writing differences therefore do not count as errors. A mechanical score does not establish preserved meaning, safety or task completion.
Audio underwent an 8 kHz mono G.711 mu-law encode/decode roundtrip and was streamed in real time in 100 ms frames. It did not traverse a live phone network, and no task scripts, corpus prompts or hotwords were supplied. p50 is the median; p95 is the 95th percentile. Timing starts when audio submission ends and stops when the recognition provider reports completion. Translation, speech synthesis and playback are excluded.
The control C harness adds a fixed waiting period and has no comparable completion event, so its completion delay is withheld. “Control” refers to the historical experiment configuration, not today’s production routing. This small sample has not been independently audited and cannot establish performance across real phone calls.
These figures measure speech recognition and the listed server playback-start event. They are not a translation-meaning accuracy score and do not include device, network or acoustic-arrival time after that server event.
Sources and next measurements
Google FLEURS · Download the sanitized benchmark JSON
The aggregate records the date, anonymous test groups, sample sizes, scoring method and SHA-256 digest of the source score file. Further testing needs more unique public clips, reference text fixed before testing, and measurement from the end of speech to translated audio actually reaching the listener.
Watch app translation demos and read the excerpts →
See task reporting and our completion standard →