DeskCaller
Craft

What Happens When the Caller Switches Language?

Aneeq Iftikhar
Aneeq Iftikhar · Senior Software Engineer, DeskCaller
· 8 min read
Voice agents with languages and noise, cover reading 'What happens when the caller switches language?' above 'the two callers your defaults handle worst', beside a glowing blue line-art telephone handset with jagged red noise waves on one side and two overlapping speech bubbles on the other

Two callers give most agents trouble, and a UK trade line gets both.

The first is standing next to a cut-off saw. The second starts in English, gets to the part they care about, and finishes the sentence in Urdu.

Neither is unusual, and neither is handled by the defaults. Worse, the documentation for one of the most popular setups will tell you it is handled when it is not.

The multilingual claim that does not survive reading

Vapi's own provider comparison table lists Deepgram as "✅ Full auto-detection" with "100+" languages, and recommends it. That number is Deepgram's single-language catalogue, not what auto-detection covers.

Deepgram's own documentation is clear on the actual figure: language=multi on Nova-3 covers English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch. Ten. On Nova-2 it is Spanish and English. Two.

Urdu, Punjabi, Polish and Welsh are all in that 100+ catalogue as single-language options. None of them is in the auto-detect set. So a caller who switches into any of them mid-sentence is transcribed by a model that was not told to expect it.

There is a second Deepgram feature that looks like the answer and is not. detect_language covers roughly 35 languages, and the documentation states plainly: "Language Detection is not currently supported for streaming." Batch only. It never runs on a live call. Any architecture that puts it in a phone agent is drawing a box that will not exist at runtime.

The defaults are stricter than you think

Two configuration traps in the same platform, both of which fail silently.

Vapi's Gladia transcriber has languageBehaviour, and its documented default is "automatic single language". The agent detects one language at the start and locks to it for the call. Vapi's multilingual guide lists Gladia as supporting code-switching without mentioning that the default prevents it.

Vapi's Soniox transcriber sets languageHintsStrict to true by default, which is Vapi's default rather than Soniox's. If you pass ["en"] as a hint and the caller answers in Polish, strict mode restricts transcription to English. You get English-shaped nonsense rather than an error.

What actually covers UK callers

One configuration, and it is not the recommended one. Soniox on Vapi accepts 185 language codes including Welsh, Punjabi, Urdu, Polish, Bengali and Gujarati, and Vapi documents that setting transcriber.languages to an empty array lets it "automatically detect and transcribe any supported language, including code-switching within a conversation".

Empty array, not a list. A list plus the default strict flag is the trap above.

For comparison, AssemblyAI's multilingual model on Vapi covers 18 languages: English, Spanish, French, German, Italian, Portuguese, Turkish, Dutch, Swedish, Norwegian, Danish, Finnish, Hindi, Vietnamese, Arabic, Hebrew, Japanese and Chinese. Vapi's guide recommends it when "language detection is inaccurate". For a UK caller base that recommendation is inert, because the languages that turn up are not on the list.

On Retell, Welsh exists as cy-GB, with Azure and Soniox on the recognition side and ElevenLabs and OpenAI for the voice.

One cost note: ElevenLabs Agents charges a flat per-minute rate with no multilingual surcharge, identical across every tier from Free to Business. Multilingual is a configuration decision, not a pricing one.

Now the building site

Noise is the other half, and here the research is older and better than anything the vendors publish.

The Aurora-2 benchmark measured recognition accuracy against signal-to-noise ratio on 8 kHz telephone-filtered audio, which is the right channel for our purposes. Under multi-condition training: 98.52% word accuracy on clean audio, 97.69% at 20 dB SNR, 96.94% at 15 dB, 94.89% at 10 dB, 87.82% at 5 dB, 61.71% at 0 dB, and 24.55% at -5 dB.

Read the shape rather than the numbers. Accuracy holds up remarkably well down to about 10 dB and then falls off a cliff. Between 5 dB and 0 dB it loses twenty-six points.

Two caveats that matter. This is a 2000-vintage recognition system on a small digit vocabulary, so a modern model will sit higher across the board. And the second finding is the more useful one: training matters more than the noise. The same benchmark scored 61.34% when trained on clean audio and tested in noise, against 87.81% for multi-condition training. Babble was the worst case at 49.88%.

Babble being worst is why a busy restaurant is harder than a busy road. Speech-shaped noise defeats a recogniser in a way that engine noise does not.

Why a caller cannot simply speak up

People do raise their voices in noise, and it is a measurable reflex rather than a courtesy. ITU-T P.1100 quantifies it: speech level rises about 3 dB for every 10 dB of ambient noise above 50 dB(A), capped at roughly +8 dB, and it stops rising above about 77 dB(A). Note the edition, since 03/2017 has been superseded twice, most recently in 10/2025.

So the caller's compensation runs out. Past a certain ambient level they cannot get louder, and the signal-to-noise ratio degrades from there whatever they do.

For scale, the Control of Noise at Work Regulations 2005 set a lower exposure action value of 80 dB(A), an upper action value of 85 dB(A) at which hearing protection must be provided, and an exposure limit value of 87 dB(A), with construction specifically named in HSE guidance. Those are daily or weekly average exposures rather than instantaneous readings, so they are not directly comparable to the P.1100 ceiling. They do tell you that a working site routinely sits in the region where the Lombard reflex has nothing left to give.

The practical consequence for a trade line: an agent for electricians or roofers will meet callers it genuinely cannot transcribe, and the design question is what it does then rather than how to squeeze another two points of accuracy out of it. Detect the condition, say plainly that the line is difficult, and offer a callback or a text. That is a better answer than three rounds of "sorry, could you repeat that", which is covered in the names guide.

What the law actually requires, which is less than you have read

Three things, and two of them are commonly overstated online.

Welsh. If you are an ordinary private business, the Welsh Language Standards do not apply to you, and the Welsh Language Commissioner says so in its own published material. Section 25 of the Welsh Language (Wales) Measure 2011 sets six cumulative conditions before any duty exists, and an ordinary private company fails at the first. The Measure does reach some private bodies: water and sewerage undertakers were brought in under the No. 9 Regulations in 2023, and because those bodies sit in Schedule 8, only service delivery and record keeping standards can be applied to them.

Where the Standards do apply, they are demanding. Standard 22 requires the complete automated telephone service to be available in Welsh, and speech-driven option selection counts as an automated system. So a covered body cannot bolt Welsh onto the front of an English agent.

Language on its own. Not a protected characteristic. Section 9 of the Equality Act 2010 defines race as colour, nationality, and ethnic or national origins. An English-only phone line is not unlawful because a caller speaks Polish.

Disability, which is different. An English-only voice line is a provision, criterion or practice under section 20(3), and the reasonable-adjustments duty bites where a disabled person is put at a substantial disadvantage. Section 20(11) treats a relay call or an interpreter as an auxiliary aid. The EHRC Code, whose 2026 edition came into force on 5 August 2026, gives an example directly on point: software that cannot store a caller's access preference. If your agent cannot remember that this caller uses relay, that is the failure the Code describes.

If you deliver NHS services, and dental practices, pharmacies and opticians often do, you must also have regard to the Accessible Information Standard under section 250 of the Health and Social Care Act 2012. Have regard to, rather than a flat binding requirement, but it is the standard a commissioner will ask about.

What we would set

  • Pick the transcriber for the languages your callers actually use, not for the biggest number in a comparison table. On Vapi that currently means Soniox with an empty languages array for a genuinely mixed caller base.
  • Check the strictness defaults. languageHintsStrict and Gladia's automatic single language both silently do the opposite of what a multilingual deployment wants.
  • Assume narrowband and design the readback accordingly.
  • Detect the unworkable call and exit it honestly. A caller on a site does not need four attempts; they need a text message and a callback.
  • Persist a caller's communication preference. It is good practice, and for a disabled caller the EHRC Code treats failing to do so as a live example of an inadequate adjustment.

The thing not to do is add languages you cannot support properly. An agent that greets a caller in Polish and then cannot follow the answer is worse than one that never offered.

🎙️ Talk to Our AI Agent

Try it now - it's live!