Skip to content

Invest1 publisher3 min readPublished

Docusign cut per-document AI cost 50-fold by routing most work to models it trained itself

A Microsoft case study credits Docusign's swap to small task-specific models with 90% lower cost and eight times the throughput. Docusign's own accounting of the whole pipeline claims 50 times cheaper per document.

The Investor · Invest desk

Illustration accompanying Docusign cut per-document AI cost 50-fold by routing most work to models it trained itself

What happened

  • Docusign was running large general-purpose models on every contract it processed, and at more than a million agreements a day that bill had become its biggest infrastructure line.
  • Microsoft said on Sept. 15 that the swap to smaller models trained for specific jobs cut Docusign's AI processing costs 90% and raised throughput as much as eightfold.
  • Docusign also stopped sending whole contracts to the model, and said in a June blog post that per-document cost across the whole pipeline fell 50-fold against the earlier system.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint The extraction layer sets the floor under Docusign's agent line, because every Iris agent acts on data the small models produce first, so the routing choice caps what those products cost to run at volume.
  • contradiction The case study says accuracy held within 2 points of the larger models while Docusign says the small models beat the model they replaced by 7 points, and which figure holds decides whether right-sizing is a discount or an upgrade.
  • decision Any buyer with a repetitive extraction workload now faces a build-or-buy choice between per-call frontier pricing and paying engineers to train, host and maintain in-house models plus a passage filter.
  • exposure The headline savings were published by Microsoft, whose Foundry hosted the training, so anyone repricing their own stack on those numbers is leaning on a supplier's account of its own platform.

The two cost numbers do not stack, and the gap between them is the interesting part. Ninety percent off is a ten-fold cut. Fifty-fold is 98% off [1]. Put the old system at 100 units a document: the swap to small models takes it to 10, and trimming the input plus deleting a cleanup step takes 10 down to 2 [2]. So the model swap removed 90 of the 98 units, and everything after it removed 8, which was still a further five-fold on what was left [2].

Ramachandra Kota, Docusign's senior director of applied science, said in the Microsoft case study: "At a million documents a day, the difference between sending a full 100-page contract and sending the 4,000 tokens that actually matter is the difference between a viable business and an economics problem." [7] If every one of those million-plus documents a day went through at 4,000 tokens, the pipeline still ships 4 billion tokens a day [4].

What Docusign bought with the change was engineering, not capacity. It trained its own extraction models on Microsoft Foundry and held the frontier model back for reasoning work, such as untangling a complex clause or comparing terms across several documents [8]. A filter picks the passages likely to hold the answer and sends only those, and on several extraction types the trimmed input produced better answers than a full document sent to a bigger model [9]. One step went away entirely: the cleanup pass that had been patching the old model's mistakes, retired because the small models came in 7 percentage points more accurate than the frontier model they replaced [10].

The backdrop argues for looking at volume before price. Per-token prices are down about 98% since 2022 while enterprise AI bills are up an estimated 320%, PYMNTS reported in June, because usage grew faster than unit prices fell [12]. Run those two figures together and implied token volume is up roughly 210-fold [3].

Two things limit how far this generalises. The first is the workload: Iris pulls more than 50 facts out of each agreement, including contract value, governing law and renewal dates, the same fields the same way, for customers who arrive with tens of thousands of legacy contracts spread across 20 or more systems [5][6]. The second is provenance, since the 90% and eightfold figures were published by Microsoft, whose Foundry hosted the training [3][8]. The accuracy claims also differ. The case study says accuracy stayed within 2 percentage points of the larger models [4]; Docusign says its small models beat the retired frontier model by 7 points in production [10].

The saving holds only if the agent layer does not spend it. Docusign unveiled an Iris assistant, AI agents and an Agent Studio in May to automate contract work across sales, HR, procurement and legal, and every one of those agents acts on data the extraction models produce first [13]. SEI pulls obligations from its master service agreements and connects them to Workday and HubSpot, SEI CEO Bill Gallagher said in the Microsoft case study [14].

In my view the result is real and narrow: it prices one repetitive high-volume task, and says little about what the reasoning calls will cost. That distinction matters because payback elsewhere is not arriving. PYMNTS Intelligence puts the share of large US enterprises that have broadly deployed or embedded new AI in data and technology functions at 81% to 95%, and the share saying the investment has fully paid back at 5% to 10% [15]. At least half in every industry surveyed put full payback five to six years out, and at least 8 in 10 plan to raise AI budgets next year, with none planning a cut [16].

What to watch

  • Whether Docusign publishes per-agent inference cost for the Iris agents, which would show if reasoning calls are consuming the extraction saving.
  • Whether the 7-point accuracy gain holds on contract types outside the small models' training distribution, or the retired cleanup step comes back.
  • Whether the next Enterprise AI Benchmark moves the 5% to 10% share of enterprises reporting full AI payback.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories