Modelling your contact centre AI investment: Avoid the hidden tax.
By Sally Hodgin, Principal AI Consultant at Connect and Tim North, Group Vice President of Strategy at Connect
The UK contact centre industry spends around £30 billion* per year handling inbound calls. As organisations race to implement AI-first voice automation, a new cost layer is quietly emerging alongside that human spend: LLM compute, orchestration infrastructure, and the operational inefficiencies that can accompany poorly designed Voice AI architectures.
At national scale, we estimate the combination of LLM overuse, prompt bloat and inefficient context management could introduce £1 billion or more in additional annual cost across UK contact centres. This risk doesn’t come from AI itself. It comes from how AI is implemented.
We are seeing a growing shift towards routing interactions through Voice AI, but in practice this does not remove cost overnight. In the early stages it typically adds a layer of AI orchestration on top of the existing human workforce.

Every inbound call must first pass through some form of Voice AI triage before the system decides whether the interaction can be automated or routed to a human agent.
At UK scale this means billions of interactions passing through AI systems every year, even if only a portion of those calls are ultimately automated. While exact figures depend heavily on architecture and model choice, reasonable approximations suggest that the annual LLM compute layer supporting UK Voice AI deployments could easily surpass £1 billion.
But compute cost is only part of the picture. Even a few seconds of unnecessary latency in Voice AI interactions, caused by large prompts, inefficient context management or repeated LLM calls, can materially increase average handling time.
At national scale, single-digit seconds of additional interaction time translates into hundreds of millions of pounds in operational cost. In other words, marginal inefficiencies compound extremely quickly.
We estimate that misapplied frontier LLM usage combined with context bloat could plausibly introduce £200–£500 million per year in additional cost through increased interaction time alone.
The future of AI in contact centres therefore isn’t about replacing agents with LLMs. It’s about ensuring organisations only pay for frontier-scale intelligence when they genuinely need it.
That is why AI model architecture is rapidly becoming a commercial strategy, not just a technical design decision.
What follows is how Connect is uniquely positioned to safeguard organisations against these potential risks.
Enter the taxman
We never expected to find ourselves discussing tax strategy in the context of AI. But when Nicolas Bustamante coined the term “LLM Context Tax”, it really resonated. LLMs are undeniably powerful. The question is not whether they work. It’s whether they are being applied proportionately.
In high-volume, structured environments like the contact centre, AI model architecture is no longer a technical nuance, it is a commercial strategy.
Understanding language models in business terms.
A language model is an AI model trained to understand and generate human language by learning patterns across large volumes of text data.
Models differ primarily by scale and specialisation.
- Large Language Models (LLMs) typically exceed 20 billion parameters.
- Small Language Models (SLMs) operate in the low billions
- Micro Language Models (MLMs) are under 100 million parameters and are trained for highly specific tasks such as intent detection or entity extraction
A parameter is simply a learned weight inside the model. More parameters generally mean greater expressive capacity and deeper reasoning ability, but also greater computational demand.
Cost is influenced by two factors: model size and usage. Every interaction is broken into “token” fragments of words and numbers. The more context you provide, the more tokens the model must process. Larger models make each of those tokens more computationally expensive. At scale, that combination translates directly into infrastructure cost and energy consumption.
The operational impact on contact centre performance.
Contact centre leaders are accountable for a consistent set of outcomes:
- Cost to Serve
- Average Handling Time (AHT)
- First Contact Resolution (FCR)
- Containment rate
- Customer advocacy (CSAT / NPS
- Compliance and operational risk
- ESG and sustainability performance
AI model architecture has a direct and measurable influence on each of these outcomes. The size of the model selected, the way it is deployed, the amount of context it processes and the degree to which it is specialised for the task all shape operational performance in different ways.
The discussion should not centre on whether AI is capable, that has largely been proven. The more relevant question is whether the chosen architecture improves core contact centre metrics in proportion to the cost and complexity it introduces. When automation scales, marginal inefficiencies in precision, latency, energy consumption or governance discipline compound quickly.
The impact typically emerges across a small number of primary drivers, each of which maps directly to measurable business performance.
1. Precision: Accuracy drives containment and FCR.
Most contact centre interactions are structured: identity verification, balance enquiries, appointment changes, policy updates. These are deterministic workflows that require speed, precision and reliability rather than open-ended reasoning.
Consider Identification & Verification (ID&V): Large conversational models can perform ID&V, but they are not optimised for atomic precision. If one digit is misinterpreted in a 10-digit account number, the system re-prompts. Repetition increases AHT. Escalation reduces containment. FCR drops.
Task-specific micromodels are built for this workload. In benchmark testing, specialist intent and entity models routinely achieve F1 scores in the mid-90s.
For non-technical readers, F1 is a combined measure of precision and recall. Scores above 90% indicate strong production reliability. Connect’s intent models, powered by Elerian AI, operate at approximately 96% F1 with inference times measured in milliseconds.
Translated commercially: fewer re-asks, fewer escalations, higher containment and more predictable FCR. At scale, marginal improvements in precision compound significantly.
2. Latency: Speed influences AHT and customer experience.
Latency is often underestimated, even small delays in voice automation disrupt natural turn-taking. Additional seconds during authentication or workflow confirmation increase Average Handling Time.
Higher AHT directly increases Cost to Serve, smaller models process requests faster because they require less compute per token. In structured journeys, that speed advantage translates directly into shorter interactions and more efficient automation.
In commercial terms, shaving seconds from high-volume journeys is equivalent to adding headcount capacity, without increasing headcount.
3. Energy intensity: Compute drives cost and ESG performance.
Energy benchmarks illustrate the scale of the difference:
- MLMs: ~0.1–0.4 watt-hours per million tokens
- SLMs: ~1–4 watt-hours
- LLMs: often 10–100+ watt-hours
In commercial terms, using a frontier LLM for a high-volume, low-complexity journey can consume 10 to 50 times more energy than a task-specific alternative.
In contact centre environments, processing high volumes of interactions per month, that difference impacts:
- Cloud infrastructure spend
- Cost to Serve
- Margin performance
- ESG and sustainability commitments
- Scalability and cost predictability
At scale, one element of the hidden tax on LLM usage is the cumulative compute overhead of applying frontier-scale models to workloads that do not require frontier-scale intelligence.
4. Governance and deployment: Risk influences compliance and brand.
Frontier LLMs are typically accessed via shared cloud APIs. While often secure, this model introduces external dependency and limits infrastructure control.
Smaller SLMs and micromodels can often be deployed within a client’s own environment on-premise, private cloud or tightly controlled VPC. Because they require significantly less computational resource, they can operate inside regulated infrastructure.
For financial services, utilities and public sector organisations, this supports:
- Data sovereignty
- Reduced external data exposure
- Stronger auditability
- Alignment with regulatory frameworks
Governance directly influences compliance risk and brand protection, so should not be an afterthought when designing your AI solutions.
5. Orchestration: Capital efficiency drives ROI.
The most advanced organisations are carefully orchestrating their use of AI models. Rather than relying on a single model to perform every task, mature AI strategies distribute workloads across model tiers according to task complexity:
- MLMs: handle atomic, precision-based tasks such as intent classification, entity capture, validation and ID&V. These models are optimised for speed, determinism and cost efficiency.
- SLMs: manage workflow orchestration, routing logic, summarisation and structured journey management. They coordinate the interaction, determine next best action and ensure procedural consistency
- LLMs: invoked selectively when deeper reasoning, contextual interpretation or emotional nuance is required.
From the customer’s perspective, the experience remains seamless. From the organisation’s perspective, computational intensity aligns with task complexity. Premium AI cost is incurred only where premium reasoning materially improves outcomes.
This is where capital efficiency emerges. Rather than paying frontier-scale compute for every interaction, the organisation pays for depth of reasoning only when depth of reasoning is required.
Connect’s AI models, powered by Elerian AI, are designed with this principle in mind. Delivered via secure APIs into existing loud Contact Centre Centre Solutions (CCaaS) and CRM platforms, they enable organisations to optimise intent accuracy, workflow performance and containment without replacing core CX investments. The objective is disciplined optimisation of the AI layer within the existing ecosystem, not wholesale platform disruption.
At scale, orchestration is not a technical preference. It is the mechanism through which AI maturity translates into sustainable commercial return.
The bottom line.
In the contact centre, intelligence should be measured by business outcomes, not parameter count.
- Precision improves FCR and containment
- Latency influences AHT and Cost to Serve
- Energy intensity affects margin and ESG performance
- Governance protects compliance and brand.
- Orchestration ensures capital efficiency
The hidden tax of LLMs in your contact centre captures the commercial risk of ignoring these dynamics. Strategic “tax avoidance” in this context is not about limiting ambition. It is about applying the appropriate level of intelligence to each workload so that customer experience, cost discipline and operational resilience improve together.
Footnote (Latency / AHT calculation)
ContactBabel’s UK Decision-Makers’ Guide 2024 estimates 5.38 billion inbound calls handled by agents annually, with an average cost per call of £5.58 and an average handling time of ~421 seconds. This implies a cost of roughly 1.3p per second of interaction time.
At this scale, even small latency increases compound rapidly. For example:
- +3 seconds per interaction ≈ ~£213m/year
- +5 seconds per interaction ≈ ~£355m/year
- +7 seconds per interaction ≈ ~£497m/year
Footnote (Illustrative LLM compute estimate)
ContactBabel reports ~5.38 billion inbound calls annually. If Voice AI ultimately automates 30% of interactions (~1.6 billion calls), most architectures still require all calls to pass through an AI triage layer before routing decisions are made.
Typical LLM usage patterns (roughly 8k–12k input tokens and 2k–4k output tokens per interaction) applied at this scale produce a national compute range of approximately £100m to £1bn+ annually, depending on model choice and context management. These figures are illustrative but highlight how architectural decisions around model usage and prompt design can significantly influence cost at scale.
Frequently asked questions.
What is the “LLM context tax” in a contact centre environment?
The LLM Context Tax refers to the hidden operational and financial cost of pushing every interaction through large, frontier-scale language models. While powerful, LLMs increase compute consumption, latency, energy usage and governance complexity.
When should a contact centre use an LLM instead of a smaller model?
LLMs are most valuable when deeper reasoning, contextual interpretation or emotional nuance materially improve the outcome. For high-volume, structured tasks such as ID&V, balance enquiries or routing, specialised micro or small language models typically deliver faster, more precise and more cost-efficient performance.

About'Connect.
Connect is a global customer experience specialist, systems integrator and digital transformation partner with industry-leading, technology-enabled capabilities. Founded in 1990, we’ve evolved alongside every major industry shift; from on-premise to cloud, voice to omni-channel, and now AI-enabled experience. Built on this extensive market experience, our approach is focused on delivering outcomes-based solutions that accelerate value, informed by what it takes to operate, scale, and continuously improve CX in live environments. We deliver end to end, from the network that carries customer contact, through interactions in the contact centre, to the integrated back-end systems that support them. This end-to-end accountability creates a unified view of the customer and operations, enabling consistent, reliable outcomes at scale.
Connect with us Connect UK, Connect South Africa, Connect India, Connect USA.
Find out how we can help your business communicate better.
To discuss your communications challenges and requirements, get in touch with us today.
Connect with us now.
New web: Contact Us
"*" indicates required fields