Skip to content

Product2 publishers3 min readPublished

Salesforce's Koa model reaches pilots with two accounts of what it was trained on

Salesforce and Nvidia post-trained Koa on simulated CRM workflows and claim three times fewer errors than leading general-purpose models. The published record describes the training data two different ways.

The Product Desk · Product desk

Photograph accompanying Salesforce's Koa model reaches pilots with two accounts of what it was trained on
Photo: siliconangle.com

What happened

  • Salesforce said at Dreamforce that it worked with Nvidia to build Koa, its first reasoning model, post-trained on Nvidia's open-weight Nemotron for sales, marketing and customer-support tasks.
  • Kumar said the training set was synthetic, simulating enterprise workflows across 14 industries including healthcare, finance and manufacturing, so no customer data was used at all.
  • Salesforce plans to roll out Nvidia's Nemotron models inside Missionforce Operations in October for customers running air-gapped networks and private clouds.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • contradiction The two published accounts of the training data do not line up, so a data protection review cannot sign off on "no customer data" until Salesforce says in writing what the synthetic generator was conditioned on.
  • decision Agentforce owners now decide per task whether reasoning stays inside Salesforce or leaves for Claude, and the ClaudeForce deal keeps the frontier path funded whichever way they route it.
  • constraint Anyone outside the pilot has to take the error claim on the vendor's terms until winter, because the published figure is a ratio against an unnamed comparison set.
  • capability Government and other regulated buyers can put Nemotron reasoning next to their most sensitive data inside air-gapped networks from October, before Koa itself is generally available.

A privacy reviewer who reads both write-ups of Koa's training comes away with two answers. SiliconANGLE reports that Salesforce "used its customer data to build a massive synthetic dataset crafted out of almost 30 years' worth of internal CRM deployment experiences," in the same paragraph where it quotes Chief Platform and Engineering Officer Rohan Kumar saying the company avoided using customer data at all [5][4]. TechCrunch reports that Salesforce and Nvidia did not use any actual customer data and instead crafted synthetic data that mimicked customers' patterns [12]. Those are different statements, and they lead a data protection officer to a narrow written question: what was the generator conditioned on?

Kumar framed it as institutional knowledge. "There is so much knowledge that we've sort of gathered in building out the CRM over the last 27 years," he said [6]. Jayesh Govindarajan, EVP of Salesforce AI, described the environment they built instead of using records: "We actually simulated a customer service environment with a persona customer service professional, including irate customers that call into the customer service center, all the way to a sales professional who's trying to close a deal" [15].

Provenance is also why Salesforce picked the base model. Govindarajan said Nemotron was the first available pre-trained American model that was state of the art and had clear data provenance, adding: "We have no idea what Qwen trains on" [16].

The reliability number comes from Salesforce itself. In its benchmark tests, the company says Koa matched or exceeded multiple leading general-purpose models on CRM-based tasks with three times fewer errors [8]. Three times fewer puts the error rate at a third of the baseline, a reduction of about 67 percent [9]. Salesforce described the comparison set only as leading general-purpose models and published the ratio without absolute error rates [22]. A pilot team therefore cannot tell whether the baseline missed three steps in a hundred or thirty.

For a team already on Agentforce, this changes a routing rule. Multistep reasoning used to leave the platform for a frontier model such as Claude or ChatGPT through the Agentforce gateway [14]. "But reasoning has always been something that we've relied on the frontier model providers for. Until now," Govindarajan told TechCrunch [13]. Salesforce also announced ClaudeForce with Anthropic, which lets companies use Claude as their interface while data stays in Salesforce's system of records [19].

Then there is cost. TechCrunch reports Koa uses fewer tokens to do the same work [17], and Nvidia's Kari Ann Briski said, "we have a unique architecture for inference to be token efficient" [18]. The Friday review asks a rollout owner about the cases that closed without a human going back in, not tokens per prompt.

Only one pilot voice is on the record, and it is qualitative. Ryan Teeples, chief strategy officer at 1-800Accountant, said "Koa gives our agents the reasoning to work through that complexity step by step and use the right tools along the way" [21]. Koa stays in an expanded pilot with select customers until general availability in winter [10], so that window is where the two claims can be tested. The comparison that settles it is the same hundred cases sent down both paths, scored on the share that closed without a reopen and the tokens each close consumed.

What to watch

  • Whether Salesforce names the general-purpose models and publishes absolute error rates behind the three-times-fewer-errors claim before winter general availability.
  • Whether Koa is metered or priced separately from other Agentforce models when it reaches general availability.
  • Whether the Nemotron rollout for Missionforce Operations actually ships in October for air-gapped and private cloud deployments.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories