• Technology

Why AI Support Pilots Stall on Trust, Not Capability

  • Felix Rose-Collins
  • 4 min read

Intro

Most companies testing AI for customer support run into the same wall. The system works fine for weeks. Then it gives one customer a confident and completely wrong answer. Maybe about a refund window. Maybe about a policy that changed last quarter. Someone senior notices and the rollout that was meant to cover most of the queue quietly freezes at a fraction of it.

Industry data backs up how common this pattern is. The 2025 World Quality Report from OpenText Capgemini and Sogeti found reliability and hallucination concerns were a top barrier to AI adoption for 60% of organizations surveyed. Separate research from Gong found 58% of companies had stalled AI projects and nearly half of planned AI investment was held back by trust concerns rather than budget. The tools are capable. Confidence in them is what lags behind.

The Math Behind One Bad Answer

Support work is asymmetric in a way that punishes average accuracy as a metric. A correct answer saves a few minutes. A wrong one can approve a refund that should never have gone out, invent a policy that never existed, or make a commitment nobody authorized. When the downside of one mistake outweighs the upside of a hundred correct replies, optimizing for the average case is the wrong goal entirely.

There's a quieter cost too. A support lead who catches the AI being wrong once starts double checking everything after that, which erases most of the time it was supposed to save. Trust doesn't get scored per conversation. It builds up or it collapses, and once it collapses teams stop trusting the system even when it's right most of the time.

Hallucination Doesn't Go Away, It Gets Managed

Even the strongest language models fabricate under the right conditions, and the honest numbers are higher than most people assume. On grounded summarization, essentially the same task as answering from your own help docs, Vectara's hallucination leaderboard puts the best models around 3% and shows well known flagship models clustering between 6% and 15%. Some reasoning heavy models climb past 20% on that same task because deeper reasoning gives them more room to introduce claims the source text never made. Push a model into open ended questions with nothing to ground it against and the numbers get much worse. Stanford researchers found leading models hallucinating on a large majority of specific legal questions when no source document was provided.

None of this means any particular model is bad. It means no single model, used alone, is reliable enough to sit in front of paying customers without something watching over it.

Intelligence and Control Pull in Opposite Directions

There's a structural reason a single model can't solve this on its own. As a system gets better at handling ambiguous or unfamiliar cases, it also becomes harder to fully predict, because the same reasoning that lets it handle a case nobody scripted for is the reasoning that occasionally leads it somewhere it shouldn't go. A more capable model isn't automatically a safer one. That tradeoff is why reliability has to be engineered as a layer around the model rather than expected from the model by itself.

Aissist approaches this with four layered techniques instead of one. Prompt engineering sets the baseline rules every task follows, which matters more in an agentic system where a single customer request can expand into more than a dozen sub tasks that all need to respect the same guardrails. A booster step runs uncertain decisions more than once and keeps the answer most agents agree on, at the cost of extra compute. A self inspection pass has the system review its own output, or hand it to a second model in a different role, before anything reaches the customer. And a stacked governance layer sits above all of it, checking whether an output or action matches policy before it goes out, functioning less like a filter and more like a supervisor for the whole system.

That combination is how the platform holds its AI error rate under 1%, a number worth noting mainly because so few vendors in this space publish one at all.

Judging When Not to Answer

The most valuable behavior in a support agent isn't answering more questions correctly. It's recognizing which ones it shouldn't try to answer alone. A system that escalates a tricky billing dispute early, with full context attached, does far less damage than one that pushes ahead and guesses. This is also what separates an agent that describes a fix from one that actually applies it, pulling up the order, making the change, and confirming it back to the customer.

Reliability Has to Be Maintained, Not Just Built

A system that's accurate on launch day won't stay that way without something watching it. Products change, policies get updated, and the questions customers ask drift along with them. Continuous measurement catches that drift as data instead of as a wave of complaints, and a disciplined evaluate, test, and ship cycle keeps closing the gaps it finds, with a human still signing off before anything changes for real.

What to Actually Check Before Trusting a Vendor

Any vendor who claims their AI never gets things wrong should be treated with suspicion, because it usually means nobody is measuring closely enough to know otherwise. The ones worth taking seriously publish an error rate, explain exactly how they catch mistakes before customers see them, and are upfront about when the system hands off to a person instead of guessing. That's a very different pitch from "trust the AI," and it's the one that actually holds up once real ticket volume hits it.

Felix Rose-Collins

Felix Rose-Collins

Ranktracker's CEO/CMO & Co-founder

Felix Rose-Collins is the Co-founder and CEO/CMO of Ranktracker. With over 15 years of SEO experience, he has single-handedly scaled the Ranktracker site to over 500,000 monthly visits, with 390,000 of these stemming from organic searches each month.

Start using Ranktracker… For free!

Find out what’s holding your website back from ranking.

Create a free account

Or Sign in using your credentials

Different views of Ranktracker app