ProductReports disagree2 publishers3 min readPublished
Thomson Reuters spent $40M to make a $450K training run worth doing
The final run on an open-weight base cost about 1.1% of the two-year program. The expensive part was the corpus and the expert hours, and the live product decision is which task goes to which model.
The Product Desk

What happened
- Thomson Reuters has launched Thomson, its first proprietary large language model, built for legal work.
- It arrives as the default model for Tabular Analysis, the high-volume document review feature inside the CoCounsel assistant, with an admin override.
- Internal tests put the model broadly competitive with leading models on web-only access, and roughly equal or slightly better with company content attached.
Why it matters
- cost The entry ticket for an incumbent with a curated archive is expert payroll and editorial depth, not a compute budget: the run is a rounding error against the program.
- decision Procurement stops being a choice of frontier vendor and becomes a routing policy written task by task, since the model's own owner ships the override switch.
- exposure Every quality claim on the table is self-scored, and the academic weights plus the developer portal are the mechanism by which outsiders get to disagree in public.
- constraint Improvement is now rate-limited by expert-generated training signal rather than by archive volume, which caps progress at the speed of the expensive half of the budget.
About $450,000 buys a training run [5]. It does not buy the reason to do one. Thomson Reuters' own accounting splits those costs: roughly $40 million over two years on people and computing [4], with the final run landing at about 1.1 percent of the total [17]. The residual averages something like $20 million a year [18], and the source material says where it went. Hundreds of subject-matter experts set training objectives, wrote example legal questions and judged responses in blind comparisons [15]. Reinforcement learning taught the model to operate Westlaw and Practical Law [8]. Westlaw itself runs to more than 40,000 databases and over 150 years of editorial curation [9]. The compute was the cheap line.
So the transferable lesson for a publisher sitting on an archive is not that models got cheap to train. It is that the run has stopped being the constraint, and the curated corpus plus the expert payroll are now the whole cost. Thomson Reuters skipped foundation-model work entirely, starting from an open-weight base and layering on its content, its training methods and professional judgment, which it says cut inference cost as well as training cost [6].
The part worth studying is the routing. CoCounsel stays multimodel: Thomson is the default in Tabular Analysis with administrators free to select something else [2], while third-party frontier models keep the tasks where the domain model has no edge [3]. "Thomson needs to set the frontier of intelligence for legal," Joel Hron, global head of AI and TR Labs, told SiliconAngle, calling that a different job from the one the frontier labs are doing [7]. A company that just spent two years building a legal model still sends much of the legal work to somebody else's model by default [3]. Read that as a scope statement rather than a hedge: the domain model wins where the proprietary corpus and the in-house tools matter.
The quality evidence is in-house. Thomson Reuters says internal tests had the model broadly competitive with leading models when everything was restricted to the web, and roughly equal or slightly better once connected to company content, scored on answer completeness and on whether citations supported their claims [10]. Those results have not had extensive independent validation, and the company is now circulating the model to legal experts and academic institutions, planning a smaller open-weight version on Hugging Face under a noncommercial academic licence, and building a portal for outside API keys [20].
Two dependencies sit under the roadmap. Hron concedes the question about keeping pace with faster labs and answers it by arguing that improving open models give Thomson Reuters better foundations for later versions [14], which makes another company's release schedule a planning input. And only about 10 percent of the information base has been used [11], leaving 90 percent [19] that will not simply be poured in: the stated next step is converting the most useful content and product activity into better training signals [16]. That is more expert time, the line that costs $20 million a year [18], not the one that costs $450,000 [5].
Meanwhile the sales argument is ownership. Customer data is not used to train the model, control gives the company more authority over deployment and governance, and it is discussing direct model access with large law firms and corporations that might adapt Thomson to their own knowledge and workflows [12].
What to watch
- The promised technical report, and whether its benchmark results match the internal comparisons already briefed out.
- Whether anyone holding the academic open-weight release or an API key reproduces the citation-support findings independently.
- Whether the direct model access talks with large law firms turn into deals where customers adapt Thomson on their own matter files.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence52
- Adoption28
- Hype gap+24
- Incentives74
- Confidence62
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Thomson Reuters launched Thomson, its first proprietary large language model, combining its legal knowledge with LLMs from outside providers to provide legal advice.
- [2]
Thomson will first be deployed in Tabular Analysis, a high-volume document review capability in the CoCounsel Legal AI assistant, where it will be the default model although administrators can select other models.
- [3]
CoCounsel will remain a multimodel product, using Thomson for work where the domain-specific model has an advantage and third-party frontier models for other tasks.
- [4]
Thomson Reuters spent about $40 million over two years on people and computing for the project.
- [5]
Thomson Reuters said economies reduced the cost of the final training run to about $450,000.
- [6]
Instead of building a foundation model from scratch, the company began with an open-weight model and added proprietary content, training methods and professional expertise; it said the approach reduced both training and inference costs compared with general-purpose frontier models.
- [7]
Joel Hron, global head of artificial intelligence and TR Labs at Thomson Reuters, said the company is not seeking to compete with the largest AI labs across every field: "Thomson needs to set the frontier of intelligence for legal. That's a different job than what I think a lot of the frontier labs are doing."
ReportedSupportedSource: Joel Hron, Thomson Reuters, via SiliconAngle2 sources— create a free account to open themView cited source - [8]
Training included realigning the base model with company values, pretraining on company content, targeted post-training guided by professionals, and reinforcement learning that taught the model to work with company tools such as Westlaw and Practical Law.
- [9]
Thomson Reuters' Westlaw platform encompasses over 40,000 individual databases and more than 150 years of legal publishing and editorial curation.
- [10]
Thomson Reuters said internal tests showed Thomson was broadly competitive with leading models when all had access only to the web, and moved to roughly equal or slightly better performance when connected to Thomson Reuters content, per senior research scientist Andrew Bean; tests assessed answer completeness and whether citations supported claims.
ReportedSupportedSource: Andrew Bean, Thomson Reuters2 sources— create a free account to open themView cited source - [11]
Andrew Bean said only about 10% of the company's total information base has been used so far.
ReportedSupportedSource: Andrew Bean, Thomson Reuters2 sources— create a free account to open themView cited source - [12]
Thomson Reuters said customer data is not used to train the model, that controlling the model gives it more authority over deployment, governance and future development, and that it is discussing direct model access with large law firms and corporations, open to customers adapting Thomson to their own knowledge and workflows.
- [13]
Jonathan Schwartz, head of foundational research, said specialization can damage a model's broader abilities if handled poorly, so the team focused on continual learning, adding domain skills without erasing existing capabilities.
ReportedSupportedSource: Jonathan Schwartz, Thomson Reuters2 sources— create a free account to open themView cited source - [14]
Hron acknowledged that maintaining a proprietary model raises questions about whether Thomson Reuters can keep pace with faster-moving AI laboratories, and argued that improvements in open models will give the company stronger foundations for later versions.
ReportedSupportedSource: Joel Hron, Thomson Reuters2 sources— create a free account to open themView cited source - [15]
Hron said hundreds of subject-matter experts helped define training objectives, create examples of legal questions and judge responses in blind comparisons.
- [16]
The next step is not simply adding more material but turning the most useful content and product activity into better training signals.
- [17]
The roughly $450,000 final training run is about 1.1% of the roughly $40 million spent over two years.
- [18]
The $40 million spent over two years averages about $20 million a year.
- [19]
About 90% of Thomson Reuters' information base has not been used in training so far.
- [20]
The results have not yet received extensive independent validation; Thomson Reuters has begun sharing the model with legal experts and academic institutions, plans a smaller open-weight version on Hugging Face under a noncommercial academic license, and is developing a portal for outside developers to request API keys.
Sources
2 independent publishers whose own reporting we read for this story.
- siliconangle.comThomson Reuters launches proprietary AI model for legal work
1 article · August 24, 2026
- thenextweb.comThomson Reuters built its own model on a Chinese open base
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- Joel HronFollow
- Jonathan H. ChoiFollow
- Andrew BeanFollow
- HarveyFollow
- Model Context ProtocolFollow
- Moonshot AIFollow
- ClaudeFollow
- AlibabaFollow
- Samuel DahanFollow
- WestlawFollow
- Steve HaskerFollow
- Tom CottonFollow
- iManageFollow
- CoCounsel LegalFollow
- Jonathan SchwartzFollow
- ThomsonFollow
- Kimi-K3Follow
- Thomson ReutersFollow
- QwenFollow
- River AIFollow
- Hugging FaceFollow
- Practical LawFollow
- Imperial College LondonFollow
- SnowdonFollow
- AnthropicFollow
- TenetFollow