Invest1 publisher3 min readPublished
Median AI support agent resolves 48% of conversations on Gorgias's live-store test
Gorgias tested 13 AI support vendors on 212 live ecommerce stores and found the top five resolve about 70% of conversations, the median 48%. Resolution is the pricing unit, so the rate to contract on is one measured on the buyer's own traffic, repeatedly.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Zendesk, Intercom and Klaviyo, the helpdesk and email tools most brands already pay for, resolve no more than 42% of conversations in the benchmark.
- Ada resolves 68% of conversations but scores 39 of 100 on blind answer quality, while Sierra ties for the best quality score at 72 and resolves 48%.
- Nine of the ten vendors with four weeks of history lost automation, led by Siena at minus 23 points; Envive alone improved, by 8.
- Gorgias's own dashboards count a conversation as automated once 72 hours pass without a human agent, against the benchmark's rule of zero human touch.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- cost With resolution as the billing unit, the definition written into the contract decides whether a customer who gave up is billed as a success.
- decision Buyers have to budget for repeat measurement after go-live, because a rate taken in a one-month pilot had already slipped for most of the field.
- exposure Brands running a high-automation, low-quality agent end up paying to handle repeat contacts their own agent created.
- constraint Gaps of a few points among the top three cannot settle a purchase when one of the leaders is scored on 34 conversations.
The top of the table rests on few conversations. Rep AI, Decagon and Yuma contributed 272 of the 4,119 scored conversations, about 6.6% [19][20], and Decagon's 73% comes from 34 of them [4]. If each conversation is treated as independent, one standard error at that size is about 7.6 points, which puts a 95% band at roughly 58% to 88% [21]. Gorgias's 64% and Rep AI's 75% both fall inside it [4][5].
Gorgias built the benchmark and is ranked in it. SaaStr, whose fund led Gorgias's seed, describes it as a $100M ARR company [18]. Gorgias places fifth on automation at 64%, with a quality score of 66 [5][8]. SaaStr's write-up names Yuma, Decagon and Gorgias as the only vendors above 64% automation and 65 quality. Those cutoffs sit exactly on Gorgias's automation score and one point under its quality score [9][22]. The method is careful all the same. A judge blind to the vendor scores each response against 26 binary checks, and claims about price, policy and SKU are checked by program against the live store [2].
For a buyer, the four-week trend matters more than the ranking. Quality fell for nine of the ten vendors with history: Ada by 17 points, Zendesk by 15, Sierra by 14 and Gorgias by 12 [12]. The benchmark does not explain the declines. SaaStr lists model updates, store configuration drift or a harder question set as possible contributors [13]. If the question set got harder, the test moved and the agents did not, and a buyer's own traffic would show no fall. Configuration drift would reach the buyer's stores too. Model updates are the vendor's doing, and a contract can require re-measurement after each one.
Price follows the definition. In the benchmark, "email us", a contact form or "call us" counts against the agent, and an agent that pushes customers out of the channel in more than half its replies is scored unresolved [15]. Under Gorgias's dashboard rule, a customer who gave up and never came back counts as a resolution [16]. Gorgias AI Agent charges $0.90 per resolved conversation [17]. SaaStr's account does not say which definition the invoice uses. If it is the looser one, every 100 abandoned conversations counted as resolved add $90 to the bill [23].
The incumbents answer well and hand off often. Intercom scores 66 on quality and Klaviyo 65, while Ada scores 39 [8]. A brand that keeps Klaviyo's agent avoids paying a second vendor. It also leaves 79% of engaged conversations unresolved by the AI, against 58% at Intercom [6][24]. An agent with Ada's profile closes more conversations. The benchmark argues that a wrong answer marked resolved creates a second contact that costs more than the first [10].
I think the case for measuring on the buyer's own traffic holds, with a limit the benchmark itself exposes. A pilot of a few dozen conversations has an error band as wide as Decagon's, and a one-month pilot faces the same decay as every vendor in the table [21][11]. SaaStr's own practice is repetition. "We run 20+ agents in production at SaaStr, and we re-check every one of them on a schedule, including the ones that seem to be working," it wrote [14]. This view is wrong if the next four weeks show the declines reversing and the rankings holding on larger samples. In that case a public benchmark like this one would be a cheaper substitute for a buyer's own test.
What to watch
- Gorgias's next four-week update: whether the nine automation declines continue or reverse, and whether Decagon's sample grows past 34 conversations.
- Whether Gorgias's invoices follow the benchmark's zero-touch definition or its 72-hour dashboard rule.
- A replication run by a party with no vendor in the table, on the same 26 checks.