The hidden cost and risk of scaling GenAI in contact centres.
While artificial intelligence (AI) promises a revolution in contact centre productivity and operational efficiency, current deployment strategies are hitting a wall when attempting to bridge the scaling gap.
By Greg Jarvis, Global Customer Success Director at Connect
The biggest contributor to this roadblock is the pervasive belief that the most powerful and largest AI model is the best tool for every job.
The scaling gap.
While Generative AI (GenAI) pilots typically demonstrate immense potential in controlled conditions, the transition to live, high-volume operational environments often fails to replicate this success.
For starters, large-scale models are often slow. In a live customer interaction, a five-second delay in an agent-assist prompt leads to dead air and plummeting CSAT scores.
Generalist models are also prone to hallucinations because they are trained on the entire internet, not your specific product catalogue or compliance scripts. Error-ridden, irrelevant or outright comical responses can erode the customer experience and impact trust.

Additional challenges often stem from cloud-based GenAI tools, which can introduce risks around data security, control, consumer protection, and regulatory oversight.
The hidden cost.
Furthermore, when deploying AI models in contact centres, bigger isn’t always better. Running large language models (LLMs) privately is not usually a viable alternative due to high GPU and specialist infrastructure requirements, and rising energy consumption.
When the cost-per-query of an AI interaction exceeds the cost of the human time it was meant to save, the business case collapses. At scale, the operating model of many private LLM setups is simply prohibitively expensive.
This reality leaves operators contemplating a false choice: Accept the systemic risks of public cloud models, including data sovereignty concerns and regulatory opacity, or shoulder the crushing operational costs of private infrastructure, specialist talent, and unsustainable energy consumption.
A production-grade alternative.
A more effective approach is to use smaller, purpose-built domain or task-specific models, especially when operating in regulated production environments.
What contact centres need is a production-grade AI engine designed to avoid these trade-offs. Constructed with a combination of optimised large, small (SLM), and micro language models (MLM), these production-grade AI engines are purpose-built for specific tasks, rather than trained on the entire internet.
As a result, the models are smaller, more accurate for their intended use, and can be securely hosted within the organisation’s environment.
This approach gives organisations greater control over data and behaviour, supports stronger governance and consumer protection, and significantly reduces the compute and energy required to operate AI at scale.
The outcome is a much more sustainable cost profile and a materially stronger return on investment as automation volumes increase.
Ultimately, sustainable AI isn’t about running the biggest models available. It’s about deploying controllable right-sized AI models that work accurately, securely, and at a cost that holds up in production. That’s how regulated contact centres can successfully scale AI.
Frequently asked questions.
How can contact centres reduce the cost and risk of scaling GenAI?
LLMs, while powerful, are often cost-prohibitive at scale due to high compute, infrastructure, and energy demands. A more sustainable approach is to deploy a combination of large, small, and micro language models, each optimised for specific tasks. This architecture reduces dependency on public cloud risks while avoiding the heavy overhead of private LLM infrastructure.
Can smaller AI models outperform large models in contact centres?
Yes. Smaller, purpose-built models are designed for specific tasks, not general knowledge. Unlike large models trained on broad internet data, smaller models are trained (or tuned) on domain-specific data, such as product information, workflows, and compliance rules. This makes them more accurate, consistent, and reliable in real customer interactions. They also deliver faster response times, are far more cost-efficient, requiring less compute, and they offer greater control and governance.

About'Connect.
Connect is a global customer experience specialist, systems integrator and digital transformation partner with industry-leading, technology-enabled capabilities. Founded in 1990, we’ve evolved alongside every major industry shift; from on-premise to cloud, voice to omni-channel, and now AI-enabled experience. Built on this extensive market experience, our approach is focused on delivering outcomes-based solutions that accelerate value, informed by what it takes to operate, scale, and continuously improve CX in live environments. We deliver end to end, from the network that carries customer contact, through interactions in the contact centre, to the integrated back-end systems that support them. This end-to-end accountability creates a unified view of the customer and operations, enabling consistent, reliable outcomes at scale.
Connect with us Connect UK, Connect South Africa, Connect India, Connect USA.
Find out how we can help your business communicate better.
To discuss your communications challenges and requirements, get in touch with us today.
Connect with us now.
New web: Contact Us
"*" indicates required fields