Building a customised voice bot with your own NLU data | Connect

Building a customised voice bot with your own NLU data.

Custom NLU data, domain-trained speech recognition and integrated voice AI can improve accuracy across accents, code-switching and under-represented languages, enabling more natural and effective customer conversations.

By Alfredo Gemma, AI Solutions Director at Connect


When attempting to deploy conversational artificial intelligence (AI) across diverse global markets, enterprise operators frequently discover that off-the-shelf voice solutions degrade rapidly when confronted with under-represented languages, regional dialects, and complex code-switching patterns.

The fundamental flaw lies in the standard voice bot architectural paradigm – a rigid, half-duplex pipeline where Voice Activity Detection (VAD), Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), and Text-to-Speech (TTS) execute sequentially without feedback.

While this process works for English, this linear dependency guarantees failure in real-world scenarios when faced with under-represented languages where ASR errors compound with limited NLU training data, resulting in poor intent classification and broken customer experience delivery (CX).

voice bot blog image

Knitting together an integrated conversational fabric.

Solving this challenge requires replacing static architecture with a full-duplex, agent-orchestrated framework built on a streaming-first service.

Rather than stitching disconnected models together, the underlying software is engineered into an integrated conversational fabric where each layer continuously informs the next.

This design allows the voice bot to process incoming audio while simultaneously generating speech, dynamically handling natural human interruptions while preserving contextual model selection.


Make every layer of AI work together.

Discover how an integrated voice AI architecture creates more natural, context-aware conversations.


Bridging the ASR–NLU gap.

Generic ASR systems misrecognise domain terminology at rates that render them operationally unreliable for enterprise applications.

In high-volume banking environment, for example, specialised terms like "EMI" are regularly transcribed as "army", while in healthcare applications serving isiZulu speakers, critical medical terms are routinely lost to acoustic noise and variations in dialect.

To eliminate this operational friction, acoustic models must undergo domain-specific training on telephony-grade audio infused with natural code-switching patterns, rather than relying on clean, read-speech datasets.

This acoustic foundation is subsequently paired with an in-domain language model designed to resolve lexical ambiguities directly from context. ASR hypotheses are automatically rescored using deep domain knowledge, biasing system outputs toward verified technical nomenclature before downstream processing occurs.

This process of converting under-utilised linguistic resources into training infrastructure is a type of strategic "tax avoidance" – a mechanism to bypass the prohibitive "Large Language Models (LLM) Context Tax" associated with throwing massive, context-heavy prompts at generic models.


Fine-tuning models with curated data.

For under-represented languages, enterprise development teams rarely possess millions of annotated conversational transcripts.

However, structured lexical databases, including expert-curated dictionaries, WordNets, and formal grammatical descriptions, frequently exist. These structured resources can be programmatically converted into millions of high-quality instruction-response pairs.

When applied to the Hindi WordNet, for instance, this conversion methodology produced 1.25 million targeted training pairs.

When combined with parameter-efficient fine-tuning methodologies, a targeted 12B-parameter model can comfortably outperform frontier-scale LLMs on domain-specific execution tasks.

Where lexical resources are severely constrained, synthetic-hybrid data generation paired with human-in-the-loop validation provides an effective path forward.

For severely under-resourced languages, a tightly curated dataset of 10,000 conversational pairs is sufficient to establish new operational benchmarks.

This is not uncalibrated synthetic data; it is curated data engineered to mirror authentic conversational cadence, localised slang, and dialectal nuances.


Rapid deployment across niche domains.

By pairing custom training data with knowledge-aware, audio-grounded generative slot-filling frameworks, systems achieve robust zero-shot and few-shot adaptation.

This architectural approach enables enterprise operators to launch a domain-tuned voice bot using as few as 20 annotated examples per intent.

This rapid deployment model rests on three practical accelerators that convert limited intent data into production-ready conversational capability:

  • Synthetic conversation synthesis: Multi-speaker dialogue pipelines generate scenario-level interactions, mapping speaker attributes to TTS profiles to assemble speaker-aware simulated conversations.
  • Acoustic model initialisation: Synthesised multi-speaker speech acts as a highly effective initialisation baseline for low-resource ASR training, accelerating acoustic convergence.
  • Prompt optimisation: Advanced prompt engineering reduces overall fine-tuning requirements, accelerating market delivery while maintaining high atomic precision.

As a technology-agnostic systems integrator, Connect builds solutions tailored to how people actually communicate, incorporating local accents, code-switching, and cultural context.

The goal isn't to replicate what works for English; it's to build systems that work for the people who actually speak these languages – with their accents, code-switching, and cultural context.

By deploying a hybrid intelligence approach that integrates models, we ensure your AI infrastructure delivers high containment rates, reduces Cost to Serve, and transforms CX delivery across every interaction.

Frequently asked questions.


How much training data do you need to build a customised voice bot?

You don’t necessarily need millions of annotated conversations. By combining existing linguistic resources, domain-specific data, synthetic conversations and targeted fine-tuning, effective voice AI can be built with much smaller datasets. In some cases, domain-tuned voice bots can be developed with as few as 20 annotated examples per intent.

Why customise voice AI instead of using an off-the-shelf model?

Generic voice AI models can struggle with regional accents, code-switching, local terminology and industry-specific language. Customising the ASR and NLU layers around how your customers actually communicate can improve recognition and intent accuracy, increasing containment, reducing cost to serve and creating more natural conversations.

faq image

Can voice AI work effectively with under-represented languages and regional dialects?

Yes. Domain-specific acoustic training, curated linguistic data and synthetic conversations can help voice AI recognise under-represented languages, regional dialects, accents and code-switching more accurately. This allows businesses to build voice experiences around how customers actually speak rather than relying on models predominantly optimised for English.

About'Connect.

Connect is a global AI-enabled CX specialist and digital transformation partner. Founded in 1990, we help organisations modernise customer journeys and optimise service operations across every touchpoint, applying AI where it delivers measurable operational value.

Our differentiation lies in the experience we’ve gained from operating CX in the real world. We deliver end to end; from the network that carries customer contact, through interactions in the contact centre, to the integrated back-end systems that support them. This end-to-end accountability creates a unified view of the customer and operations, enabling consistent, reliable outcomes at scale.

Connect with us Connect United Kingdom, Connect South Africa, Connect India, Connect USA.

Find out how we can help your business communicate better.

To discuss your communications challenges and requirements, get in touch with us today.

Connect with us now.

New web: Contact Us

"*" indicates required fields

This field is for validation purposes and should be left unchanged.
Consent*