BuildReports disagree4 publishers3 min readPublished
Thomson Reuters priced the middle path at $40M, and still pays Anthropic
The $450,000 training run everyone is quoting is about one percent of the programme behind it. The in-house model took one CoCounsel feature; the agent layer under it stays licensed.
The Engineer · Build desk
What happened
- Thomson Reuters spent about $40 million on staff and compute over more than two years to continue training Alibaba's open Qwen, most recently Qwen3.5-397B, into a model it calls Thomson.
- The $450,000 figure circulating for the model covers only the final training run of the current version.
- Training drew on Westlaw, Practical Law, Checkpoint and Reuters content with hundreds of subject-matter experts evaluating it, and under 10 percent of the available corpus has been used.
- CoCounsel Legal itself is built on Anthropic's Claude Agent SDK.
Why it matters
- cost Anyone costing a domain model from the published run price underfunds it by roughly ninety times, and the gap lands mostly on internal expert time that never appears on an invoice.
- constraint The entry ticket is the archive, not the GPUs: without a corpus of Westlaw's depth to train and practise inside, this middle option collapses back into renting someone else's model.
- contradiction The New Stack reads the benchmarks as parity with OpenAI, Anthropic and Google; The Decoder reads the same tables as a method-skewed comparison whose only win is a one-point margin.
- decision Ownership becomes a per-feature call about whether volume is predictable enough to beat per-token pricing, rather than a platform-wide exit from a vendor.
One part in eighty-nine. That is the ratio between the $450,000 run figure in circulation and the money actually spent getting there [23], and it is the number to fix in your head before costing anything similar. The disclosed spend covers staff and compute over more than two years [3], which works out to under $1.7 million a month, and because the period ran longer than twenty-four months, that is a ceiling rather than an estimate [20]. Standing team, not capital project.
The $40 million is not the whole price either. Thomson Reuters' own accounting leaves out decades of Westlaw, Practical Law, Checkpoint and Reuters material and the working hours of hundreds of domain experts, according to The Decoder [4], and those experts were the people hunting for the places the model failed [5].
What the money bought looks less like weights than like a pipeline. CTO Joel Hron says the open source starting point has been changed "close to a half dozen times already" [25], and research chief Jonathan Schwartz says the bigger finding is "less the individual model and more the model factory we built" [14]. The current base was retrained with Imperial College for safety, ethics and political neutrality into an intermediate version called Snowdon, then pre-trained on in-house content and drilled with agentic reinforcement learning inside the company's own tools [28][29].
The evaluation numbers say where the value actually sits. On the company's in-house Deep Research benchmark with web access only, Thomson scores 0.53 on factual accuracy against GPT-5.4's 0.65; give both access to Thomson Reuters content and Thomson finishes 0.83 to 0.82 [21][16]. Content access is therefore worth 0.30 to the in-house model and 0.17 to the outside one, and the margin at the end is a single point [22]. The Decoder's reading is that data access does nearly as much work as the specialised training, and that newer rival models were not tested at all [18]. The company's evaluation lead, Andrew Bean, grants that on web access alone Thomson is "within the scope of the other models, but certainly not the leader yet" [17].
That is what undercuts the metaphor. Hron describes third-party models as renting a house and the in-house one as buying, with every expert review compounding into equity the company owns [11]. The equity he is describing accrues to the corpus and the review loop, not to the inference bill. Thomson's first production job is a single feature, Tabular Analysis, which can range over as many as 10,000 documents and 100 questions [7], and that is precisely the shape Schwartz says a smaller in-house model pays off on, given that standard fine-tuning tends to degrade general capability while leaving the provider in control of inference pricing and roadmap [10]. The plumbing under that feature is still bought in [2], so the second line on the bill does not go to zero; it just stops growing with document count.
Two hedges survive all of it. The model still retrieves from Westlaw and Practical Law so answers can be checked against material a professional can open [13], and the company says plainly that there is no guarantee any model, Thomson included, is error-free [12].
What to watch
- Whether moving to a stronger base such as Qwen3.8 and pushing past 10 percent of the corpus widens the 0.01 margin, and whether newer rival models get tested this time.
- Any move to sell Thomson outside Thomson Reuters' own products, which the company says it is considering, and the license terms that come with it.
- Whether the agent layer under CoCounsel Legal ever moves in-house, or Anthropic stays the scaffolding while Thomson takes more feature-level inference.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence54
- Adoption27
- Hype gap+34
- Incentives78
- Confidence63
Perspective Coverage
4 publishers- Builder
- Builder 35%
- Operator
- Operator 40%
- Investor
- Investor 25%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Thomson Reuters has developed its own AI model for legal, tax and compliance work, trained on the company's proprietary professional content and designed to power features inside products such as CoCounsel.
- [2]
Thomson Reuters still uses frontier models elsewhere in its products: CoCounsel Legal is built on Anthropic's Claude Agent SDK.
- [3]
The foundation is Alibaba's open Qwen, most recently Qwen3.5-397B, and Thomson Reuters spent about $40 million on staff and computing power over more than two years, according to the company.
- [4]
Even the full $40 million leaves out decades of content from Westlaw, Practical Law, Checkpoint and Reuters, plus the working hours of hundreds of domain experts.
- [5]
The model was trained using Thomson Reuters content including Westlaw, Practical Law, Checkpoint and Reuters; hundreds of subject-matter experts evaluated outputs and found places where the model failed; the company says less than 10% of the content available to it has been used for training so far.
- [6]
"Most of our investment went into further training on decades of proprietary content and expert-driven evaluation, not pre-training a foundation model from scratch," Thomson Reuters said.
ReportedSupportedSource: Thomson Reuters, to The New Stack3 sources— create a free account to open themView cited source - [7]
The first use is Tabular Analysis in CoCounsel Legal, which can work across as many as 10,000 documents and answer up to 100 questions about them; Thomson is becoming the default model for that feature.
- [8]
For now customers will not buy access to Thomson directly; it powers specific capabilities inside Thomson Reuters products, a smaller open-weight version is available to researchers on Hugging Face, and the company said it is looking at ways it could commercialise Thomson in future.
- [9]
A smaller version of the model is being released under a non-commercial license.
- [10]
Schwartz says standard fine-tuning techniques "tend to have a strong tendency to degrade general capability" and leave you locked into the provider for inference costs and roadmap, while a smaller in-house model pays off precisely on high-volume work like document review.
ReportedSupportedSource: Jonathan Schwartz, Thomson Reuters3 sources— create a free account to open themView cited source - [11]
Hron describes the choice as "renting a house versus buying a house": every expert review during a product update becomes training data, and with an in-house model "you are building equity in something that you own for the long-term," whereas with third-party models that value evaporates at the provider.
ReportedSupportedSource: Joel Hron, CTO, Thomson Reuters3 sources— create a free account to open themView cited source - [12]
"There's no guarantee any AI model, including Thomson, is error-free," Thomson Reuters said.
ReportedSupportedSource: Thomson Reuters, to The New Stack3 sources— create a free account to open themView cited source - [13]
Thomson is trained to flag uncertainty rather than force a confident-sounding answer, and the company has kept retrieval: the model can pull from sources such as Westlaw and Practical Law so responses are grounded in material a legal professional can check.
- [14]
Research chief Jonathan Schwartz says the bigger finding is "less the individual model and more the model factory we built."
ReportedSupportedSource: Jonathan Schwartz, Thomson Reuters2 sources— create a free account to open themView cited source - [15]
The comparison is skewed by method: Thomson competes with test-time scaling while GPT-5.5 runs without a reasoning mode.
- [16]
Only with access to the company's content does Thomson edge past GPT-5.4, 0.83 to 0.82.
- [17]
Evaluation lead Andrew Bean says there is "a big uplift that comes from being able to train on and practice with your own tools," something outside providers cannot do, and that with web access Thomson is "within the scope of the other models, but certainly not the leader yet."
ReportedSupportedSource: Andrew Bean, evaluation lead, Thomson Reuters2 sources— create a free account to open themView cited source - [18]
The Decoder notes that GPT-5.4 improves just as sharply when given the company's content, so data access does almost as much work as the specialised training, and that the company did not test newer models.
- [19]
The more widely touted figure of $450,000 covers only the final training run of the current version.
- [20]
Spread over twenty-four months, about $40 million is roughly $1.67 million a month, and since the spend ran longer than two years that monthly figure is an upper bound.
- [21]
In Thomson Reuters' in-house Deep Research benchmark with web access alone, Thomson scores 0.53 on factual accuracy while GPT-5.4 hits 0.65.
- [22]
Access to Thomson Reuters content lifts Thomson's factual accuracy by 0.30 (0.53 to 0.83) and GPT-5.4's by 0.17 (0.65 to 0.82), leaving Thomson ahead by 0.01.
- [23]
The $450,000 final training run is about 1.1 percent of the roughly $40 million programme, which is about 89 times the run cost.
- [24]
Thomson was not built from the ground up: the company started with an existing open-source foundation and spent approximately $40 million training the model, including compute and talent.
- [25]
CTO Joel Hron says the company has "changed the open source starting point like probably close to a half dozen times already."
ReportedContestedSource: Joel Hron, CTO, Thomson Reuters3 sources— create a free account to open themView cited source - [26]
On Stanford LegalBench, Thomson scores 0.823 and trails Gemini 3.1 Pro and GPT-5.5; on the Harvey Legal Agent Benchmark it sits just behind Opus 4.8; it leads on instruction following and on PrBench Legal; on reasoning and especially coding it falls off sharply.
- [27]
The New Stack reports that early benchmarks show Thomson competing with models from OpenAI, Anthropic and Google across several professional and general-purpose evaluations.
ReportedContestedSource: The New Stack3 sources— create a free account to open themView cited source - [28]
Working with Imperial College, Thomson Reuters first retrained the Qwen base for safety, ethics and political neutrality; this intermediate version is called Snowdon, after the mountain in Wales.
- [29]
After Snowdon came pre-training on the company's own content, post-training with domain experts, and agentic reinforcement learning inside the company's own tool environments.
Sources
4 independent publishers whose own reporting we read for this story.
- letsdatascience.comThomson Reuters Launches Its Thomson AI Model for Professional Work
1 article · August 24, 2026
- runtimewire.comThomson Reuters launches $40M legal LLM inside CoCounsel
1 article · August 24, 2026
- the-decoder.comThomson Reuters bets $40M on owning its AI instead of renting from OpenAI or Anthropic
1 article · August 24, 2026
- thenewstack.ioThomson Reuters trained its own AI model. Then it kept using Anthropic’s anyway.
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- OpenAIFollow
- Joel HronFollow
- Jonathan H. ChoiFollow
- Andrew BeanFollow
- CheckpointFollow
- GPT-5.5Follow
- AlibabaFollow
- Claude Agent SDKFollow
- Samuel DahanFollow
- WestlawFollow
- CoCounsel LegalFollow
- Jonathan SchwartzFollow
- Claude Opus 4.8Follow
- ThomsonFollow
- GPT-5.4Follow
- Thomson ReutersFollow
- QwenFollow
- Alexander Kardos-NyheimFollow
- Hugging FaceFollow
- Harvey Legal Agent BenchmarkFollow
- Practical LawFollow
- PrBench LegalFollow
- Imperial College LondonFollow
- GeminiFollow
- SnowdonFollow
- Safe Sign TechnologiesFollow
- AnthropicFollow
- Stanford LegalBenchFollow