Skip to content

BuildReports disagree4 publishers3 min readPublished

Thomson Reuters priced the middle path at $40M, and still pays Anthropic

The $450,000 training run everyone is quoting is about one percent of the programme behind it. The in-house model took one CoCounsel feature; the agent layer under it stays licensed.

The Engineer · Build desk

How we use AISend a correction

What happened

  • Thomson Reuters spent about $40 million on staff and compute over more than two years to continue training Alibaba's open Qwen, most recently Qwen3.5-397B, into a model it calls Thomson.
  • The $450,000 figure circulating for the model covers only the final training run of the current version.
  • Training drew on Westlaw, Practical Law, Checkpoint and Reuters content with hundreds of subject-matter experts evaluating it, and under 10 percent of the available corpus has been used.
  • CoCounsel Legal itself is built on Anthropic's Claude Agent SDK.

Why it matters

  • cost Anyone costing a domain model from the published run price underfunds it by roughly ninety times, and the gap lands mostly on internal expert time that never appears on an invoice.
  • constraint The entry ticket is the archive, not the GPUs: without a corpus of Westlaw's depth to train and practise inside, this middle option collapses back into renting someone else's model.
  • contradiction The New Stack reads the benchmarks as parity with OpenAI, Anthropic and Google; The Decoder reads the same tables as a method-skewed comparison whose only win is a one-point margin.
  • decision Ownership becomes a per-feature call about whether volume is predictable enough to beat per-token pricing, rather than a platform-wide exit from a vendor.

One part in eighty-nine. That is the ratio between the $450,000 run figure in circulation and the money actually spent getting there [23], and it is the number to fix in your head before costing anything similar. The disclosed spend covers staff and compute over more than two years [3], which works out to under $1.7 million a month, and because the period ran longer than twenty-four months, that is a ceiling rather than an estimate [20]. Standing team, not capital project.

The $40 million is not the whole price either. Thomson Reuters' own accounting leaves out decades of Westlaw, Practical Law, Checkpoint and Reuters material and the working hours of hundreds of domain experts, according to The Decoder [4], and those experts were the people hunting for the places the model failed [5].

What the money bought looks less like weights than like a pipeline. CTO Joel Hron says the open source starting point has been changed "close to a half dozen times already" [25], and research chief Jonathan Schwartz says the bigger finding is "less the individual model and more the model factory we built" [14]. The current base was retrained with Imperial College for safety, ethics and political neutrality into an intermediate version called Snowdon, then pre-trained on in-house content and drilled with agentic reinforcement learning inside the company's own tools [28][29].

The evaluation numbers say where the value actually sits. On the company's in-house Deep Research benchmark with web access only, Thomson scores 0.53 on factual accuracy against GPT-5.4's 0.65; give both access to Thomson Reuters content and Thomson finishes 0.83 to 0.82 [21][16]. Content access is therefore worth 0.30 to the in-house model and 0.17 to the outside one, and the margin at the end is a single point [22]. The Decoder's reading is that data access does nearly as much work as the specialised training, and that newer rival models were not tested at all [18]. The company's evaluation lead, Andrew Bean, grants that on web access alone Thomson is "within the scope of the other models, but certainly not the leader yet" [17].

That is what undercuts the metaphor. Hron describes third-party models as renting a house and the in-house one as buying, with every expert review compounding into equity the company owns [11]. The equity he is describing accrues to the corpus and the review loop, not to the inference bill. Thomson's first production job is a single feature, Tabular Analysis, which can range over as many as 10,000 documents and 100 questions [7], and that is precisely the shape Schwartz says a smaller in-house model pays off on, given that standard fine-tuning tends to degrade general capability while leaving the provider in control of inference pricing and roadmap [10]. The plumbing under that feature is still bought in [2], so the second line on the bill does not go to zero; it just stops growing with document count.

Two hedges survive all of it. The model still retrieves from Westlaw and Practical Law so answers can be checked against material a professional can open [13], and the company says plainly that there is no guarantee any model, Thomson included, is error-free [12].

What to watch

  • Whether moving to a stronger base such as Qwen3.8 and pushing past 10 percent of the corpus widens the 0.01 margin, and whether newer rival models get tested this time.
  • Any move to sell Thomson outside Thomson Reuters' own products, which the company says it is considering, and the license terms that come with it.
  • Whether the agent layer under CoCounsel Legal ever moves in-house, or Anthropic stays the scaffolding while Thomson takes more feature-level inference.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence54
Adoption27
Hype gap+34
Incentives78
Confidence63

Perspective Coverage

4 publishers
Builder
Builder 35%
Operator
Operator 40%
Investor
Investor 25%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Thomson Reuters has developed its own AI model for legal, tax and compliance work, trained on the company's proprietary professional content and designed to power features inside products such as CoCounsel.

  2. [2]

    Thomson Reuters still uses frontier models elsewhere in its products: CoCounsel Legal is built on Anthropic's Claude Agent SDK.

  3. [3]

    The foundation is Alibaba's open Qwen, most recently Qwen3.5-397B, and Thomson Reuters spent about $40 million on staff and computing power over more than two years, according to the company.

Sources

4 independent publishers whose own reporting we read for this story.

  1. letsdatascience.com

    1 article · August 24, 2026

    Thomson Reuters Launches Its Thomson AI Model for Professional Work
  2. runtimewire.com

    1 article · August 24, 2026

    Thomson Reuters launches $40M legal LLM inside CoCounsel
  3. the-decoder.com

    1 article · August 24, 2026

    Thomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic
  4. thenewstack.io

    1 article · August 24, 2026

    Thomson Reuters trained its own AI model. Then it kept using Anthropic’s anyway.

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Legal AI and Professional WorkflowsFollow
  • AI Training Cost EconomicsFollow
  • Proprietary Data MoatsFollow
  • Multi-Model Routing ArchitectureFollow
  • Open-Weight Model AdaptationFollow
  • In-House LLMs vs Frontier LicensingFollow
  • Benchmark Methodology and TransparencyFollow
Loading related stories