Native Call app icon Native Call

NativeCall quality report · August 23, 2026 experiment

Measured call quality: Japanese recognition results and customer reviews

On the same public Japanese test clips, candidate configuration D recorded a 2.61% character error rate, compared with 15.29% for historical control C. This measured result under telephone-audio conditions gives us data for evaluating call quality.

Historical control C

15.29% CER

Character error rate; lower is better

Candidate D

2.61% CER

Character error rate; lower is better

Measurement scope: August 23, 2026; 5 unique Japanese clips, each repeated 3 times, with an 8 kHz telephone-codec simulation. This is a historical speech-recognition component benchmark; the full method and measurement boundary are below.

Read all four languages’ results and methodology ↓

Historical call playback logs · August 29–September 1, 2026

Time until translation playback starts

Toward the app caller · median

1.36 seconds

Toward the app caller · p95

4.276 seconds

These 33 turns have both timestamps, measured from the recorded end of speech to the server’s translation playback-start event. The measurement covers the translation and playback path after recognition, making it closer to the call flow than recognition timing alone.

Measurement boundary and sample details

The previously selected sample contains 10 calls and 115 retained timing turns: 34 have this measurement and 81 lack one or both timestamps. Thirty-three measured turns are toward the app caller. There is one recipient-direction turn, at 1.315 seconds, so no percentile is calculated for that direction. The event ends at server playback start; it excludes later device, network and acoustic-arrival time and does not measure translation accuracy.

Download anonymous aggregates, measurement scope and source digest →

Selected App Store reviews

4.749 ratings
More reviews

“Someone in our group felt unwell first thing in the morning, and we urgently needed to contact a restaurant to cancel our reservation. NativeCall help…”

loveguavarTaiwan App StoreApp Store ↗

“I needed to call to book a room at short notice and was still trying to memorise the Japanese when I found this app. It was so convenient, and the rep…”

shasi shajinTaiwan App StoreApp Store ↗

“I lost my passport while travelling in Japan and was not even sure where I had left it. Thankfully, NativeCall helped cut down on the international ca…”

芙啾Taiwan App StoreApp Store ↗

Ratings across 7 storefronts · Selected excerpts · English translations where applicable

Test configurations use anonymous group labels. Provider and model names are not published; these labels distinguish historical experiment configurations.

Historical experiment: all tested results

Each language has only 5 unique public FLEURS clips, each repeated 3 times: 15 trials per model and language. Repetitions do not create 15 independent clips.

The candidate did not improve every language: it had lower error rates on these Japanese and English clips, but higher error rates on Mandarin and Cantonese. All eight result rows are included below; the best numbers from different models are not combined into a single service score.

Speech recognition error and completion timing, August 23, 2026
Language / test groupError rateCompleted trialsASR p50ASR p95
English · Historical control A 7.50% WER 15/15 279 ms 411 ms
English · Candidate D 6.67% WER 15/15 484 ms 664 ms
Mandarin · Historical control B 5.52% CER 15/15 241 ms 330 ms
Mandarin · Candidate D 6.21% CER 15/15 277 ms 625 ms
Cantonese · Historical control B 2.34% CER 15/15 225 ms 323 ms
Cantonese · Candidate D 3.13% CER 15/15 193 ms 489 ms
Japanese · Historical control C 15.29% CER 15/15 Not comparable Not comparable
Japanese · Candidate D 2.61% CER 15/15 208 ms 433 ms
Method and measurement scope

WER is word error rate; CER is character error rate. Both divide the total insertions, deletions and substitutions by the number of reference words or characters; lower is better. English uses words; Japanese, Mandarin and Cantonese use characters. Different languages and units are not directly comparable. Subtracting error rate from 100% does not produce a translation-accuracy score.

Scoring normalizes Unicode, punctuation, Traditional/Simplified Chinese and spoken/Arabic numeral formatting. Some writing differences therefore do not count as errors. A mechanical score does not establish preserved meaning, safety or task completion.

Audio underwent an 8 kHz mono G.711 mu-law encode/decode roundtrip and was streamed in real time in 100 ms frames. It did not traverse a live phone network, and no task scripts, corpus prompts or hotwords were supplied. p50 is the median; p95 is the 95th percentile. Timing starts when audio submission ends and stops when the recognition provider reports completion. Translation, speech synthesis and playback are excluded.

The control C harness adds a fixed waiting period and has no comparable completion event, so its completion delay is withheld. “Control” refers to the historical experiment configuration, not today’s production routing. This small sample has not been independently audited and cannot establish performance across real phone calls.

These figures measure speech recognition and the listed server playback-start event. They are not a translation-meaning accuracy score and do not include device, network or acoustic-arrival time after that server event.

Sources and next measurements

Google FLEURS · Download the sanitized benchmark JSON

The aggregate records the date, anonymous test groups, sample sizes, scoring method and SHA-256 digest of the source score file. Further testing needs more unique public clips, reference text fixed before testing, and measurement from the end of speech to translated audio actually reaching the listener.

Watch app translation demos and read the excerpts →

See task reporting and our completion standard →

Read customer stories and their outcome evidence →

See capabilities and use cases →