The AI Examiner
Artificial Intelligence. Examined. Est. 2026
← Home

The AI Examiner — AI support vendors, independent-testing gap, 2026-08-27

Headline

Seven vendors publish a headline resolution or containment figure for autonomous AI customer support. Five have no independently conducted test of that figure on record. Two have numbers independent data, or the vendor's own later reversal, contradicts.

Summary

The AI Examiner's claims register tracks seven vendors selling autonomous AI agents for tier-1 customer support, each publishing a headline resolution, automated-resolution, or containment figure. Five of the seven have no independently conducted test of that figure on record. The other two have figures that independent data, or the vendor's own subsequent reversal, contradicts. This snapshot draws entirely from the register compiled 2026-08-26. It does not establish that the five unverified figures are false — only that no independent check of them currently exists.

The evidence

The two contradicted claims

Fin by Intercom publishes a 76% average resolution rate across more than 12,000 customers, improving roughly one percentage point a month, alongside a separately-claimed 93% accuracy figure self-published; see The AI Examiner, tier-1 customer support, 2026-08-07. An independently disclosed production test — four small-business clients, roughly 500 tickets a month combined, tracked over 60 days — measured a 38% average resolution rate, ranging from 28% to 52% depending on documentation quality see The AI Examiner, tier-1 customer support, 2026-08-07. The 93% accuracy figure measures a different thing than resolution rate and was not part of that test.

Klarna stated in 2024 that its AI agent had replaced approximately 700 support roles, was handling two-thirds of chats, and averaged under two minutes to resolve them self-published, Klarna 2024 statements; see The AI Examiner, tier-1 customer support, 2026-08-07. Multiple independent outlets reported Klarna moving to a hybrid AI-and-human support model after customer satisfaction dropped, and Klarna's chief executive officer stated publicly that the company "focused too much on efficiency and cost — the result was lower quality, and that's not sustainable" see The AI Examiner, tier-1 customer support, 2026-08-07. The reversal itself, not a separate resolution-rate test, is the finding.

The five unverified claims

Zendesk's President of Product, Engineering and AI, Shashi Upadhyay, told a reporter that the company's new autonomous support agent resolves 80% of customer issues without human help TechCrunch, 2025-10-08. No published methodology accompanies the figure — it is a spoken claim to a journalist, not a disclosed test — and no independent audit, academic study, or journalistic test of it was found.

Salesforce's "Customer Zero" case study claims Agentforce resolves the large majority of self-service visitor issues on Salesforce's own Help site Salesforce customer story. The percentage on that same, persistent URL has read 75%, 83%, and 85% at different observed dates, with no visible revision date distinguishing which figure is current. No independent test of the underlying Help-site claim was found.

Ada publishes industry-average figures — roughly 72% containment and 52% automated resolution, with best-in-class deployments reaching 84% or higher, across more than 550 deployments Ada resource guide. No independently conducted test of those figures was found. The only third-party data point located was a survey Ada co-commissioned, which found a large share of consumers do not believe their issue was actually resolved — vendor-commissioned self-critical data, not independent verification.

Decagon's case study for Chime claims roughly 70% chat-and-voice resolution and a 60% reduction in support costs, with the cost figure attributed by Decagon to Chime's own 2025 regulatory filing Decagon case study. The only third-party comparison located was a claim made by a competitor, Fin by Intercom, describing an outside-run test that placed Decagon below Fin's own figure — but Fin's separate case-study page for that same test does not name Decagon as the vendor tested, an inconsistency in the one source available. No independent test of Decagon's own published figures was found.

Sierra's case study for Tubi claims 80% containment and a seven-percentage-point increase in customer satisfaction, with resolution time dropping from hours or days to minutes Sierra customer story. Containment, the metric behind that figure, counts a conversation kept away from a human agent — a different measure than the resolution-rate figures the other vendors here publish. No independent test of Sierra's containment or customer-satisfaction numbers was found.

What a buyer should ask for

A buyer evaluating any of these seven vendors on a headline resolution or containment number can ask for the same three things regardless of vendor: the underlying definition of the metric, including whether a customer going silent counts as resolved; the sample the published figure is drawn from, including whether it is the vendor's full customer base or a single named account; and whether any party outside the vendor has measured the figure directly, and on what data. Two vendors carry a specific reason to ask for the current figure in writing rather than take a marketing page at face value: Fin by Intercom's disclosed rate has an independently measured production result well below the number the vendor publishes, and Klarna's own high-profile deployment already prompted a reversal by the company that ran it.

Methodology & sample

This piece draws entirely from The AI Examiner's claims register (examiner/market.db), seven rows covering one vertical — tier-1 customer-support AI agents — compiled 2026-08-26. Each row records a vendor's claim, its source, and the result of a dedicated search for independent verification. No new sourcing, fetching, or searching was performed for this piece; every sentence above traces to a claim_source_url or finding_source_url already stored in the register.

Gaps

  • Five of seven "unverified" ratings mean no independent test was found in the register's search — not that the underlying figures are false. Absence of a test is reported here as a finding in its own right, not as evidence against the vendor.
  • The register covers one vertical and one snapshot in time (2026-08-26). None of these seven vendors has been contacted for comment as part of this register entry; right-of-reply has not been exercised here.
  • Coverage is limited to the seven vendors already in the register. It does not represent the full tier-1 customer-support AI market.

Subscribe & Contribute

Follow The AI Examiner, or send us something you think we should look into.

Subscribe

Older issues are free to read in the archive here.

Charges appear on your statement as LINK.COM* theaiexamine.

Got something we should look into?

Send a question, a correction, or something you think belongs in a future issue.

How we protect what you send

Every submission is reviewed before anything is published. No private information from your submission appears in an issue without your consent.