Invest1 distinct publisher3 min readPublished
Thinking Machines' own figures show that better prompts carried the models it tested from a coin flip to the mid-70s and then quit, leaving the last five points to an 80% trust bar to be bought with annotation time from the analysts being freed.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Reading was never the constraint the lab set out to remove. The post is explicit that the work sits in the small repeated judgments carried over the reading, and that those are the judgments investors get stuck articulating even though the tasks themselves are trivial to them [14][19]. The examples carry the weight: an article touching Greenland reads as noise given its context while one on China tariffs reads as signal, and both are geopolitics bolted to finance [18]. Interestingness is defined against a particular seat, the macroeconomic investor who does not care about a small IPO however financially relevant it plainly is [17]. A prompt can only carry the part of that an expert can spell out, which is the stated reason for going to fine-tuning instead [13].
Five points sounds like polish until you count it in documents: mid-70s accuracy leaves roughly one document in four labelled wrong, 80% leaves one in five, so the stretch that no amount of prompt work could cross is a fifth of the surviving errors [3]. The price of that fifth is annotation by high-quality human hands [15], and the hands belong to the same investors whose attention the system exists to free. The post does not say how many labelled documents it took, nor how many analyst hours, and it reports on a subset of data cleared for public release [12]. Hours in against hours out is the ratio a buyer would want before signing anything.
The cost line is where I would push hardest. GPT 5.4 priced 43% above 5.2 for a marginal accuracy gain [9] means that if a single point moves, 75 to 76, the cost per correctly labelled document rises about 41% [2], and the lab's own reading is that newer models are not improving quickly at this task, per dollar least of all [10].
This is probably wrong, but the more interesting version of the mid-70s plateau is that it describes the labels rather than the models. Automatic prompt optimisation added nothing on top of the expert-written instructions [7], which is what you would expect if the residual disagreement is inside the label set. No inter-annotator agreement figure appears, and no absolute accuracy number or price for the proprietary model either, only that it beat everything tested on information accuracy and recall at a fraction of the cost [4][11]. If two macro investors split on "relevant but uninteresting" one time in six, then 80% is a description of a committee and a fine-tune that clears it has learned one desk's habits.
Which may be the commercial point. Differentiated intelligence, models tuned to specific organisational needs [16], puts the asset in a client's labelled corpus rather than in the weights, and it means research time is going into a narrow relevance classifier fitted to one firm's taste instead of a horizontal product sold off a menu. Two readings would break that. A frontier release that clears 80 on these six tasks [2] from a plain prompt makes the corpus worth close to nothing, and a per-firm labelling cost small enough to disclose turns it into a cheap onboarding step rather than a moat. Until one of those numbers is published, the demonstrated finding is that judgment is still cheaper to rent than to specify.
Ranked by verification strength, evidence, and original report placement.
Thinking Machines Lab published a post titled "Learning to Replicate Expert Judgment in Financial Tasks", describing an attempt to automate the information triage task of identifying what is relevant and interesting for investors to read.
The lab evaluated models on six information filtering tasks drawn from investors' daily workflows, and says it has many other internal tasks showing similar patterns.
The lab measured accuracy as the percentage of documents correctly labeled according to its investors, and also calculated F1 scores for classification tasks.
Variants of Gemini, Claude and GPT averaged roughly 50% accuracy when given a prompt that simply states each of the six tasks to perform.
Instructions written by the lab's experts based on real task descriptions, plus reframing of certain tasks, boosted the tested models' accuracy from a coin flip to the mid-70s.
Performance on the article classification task improved when models were asked to sort news stories into three labels: relevant and interesting, relevant but uninteresting, and irrelevant.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
security
Washington names industrial-scale distillation, then hands the detection bill to abuse teams1 distinct publisher
product
Model choice is becoming a line item, and the differentiator moved up the stack1 distinct publisher
invest
Scalable Capital puts ChatGPT, Claude and Grok inside the European order ticket2 distinct publishers
build
Count invalid JSON as a failed classification, and model choice becomes a reliability problem1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-published account, partly withheld
The frontier-model half of this is specific and internally coherent: a stated baseline, a stated ceiling, a named price gap, a defined accuracy metric. All of it appears in a single post on Thinking Machines' own site, computed on a subset of data cleared for release, with no outside party having run the six tasks. The moment the subject switches to the firm's own model, the numbers stop and the adjectives start — no accuracy figure, no price, and no measure of how consistently the expert labellers agreed with each other, which is the ceiling every one of these scores sits under.
One in-house pipeline, no users named
Something is genuinely running: Qwen3-235B fine-tuned on Tinker with GRPO, fed a training set scrubbed by sending contested labels to staff experts. That is the whole of the adoption record. No customer, no availability, no document volume, no indication the system has taken over any analyst's queue — and by the post's own accounting the frontier baseline never reaches the accuracy its investors said they would need before trusting it.
Adjectives where the score should be
'Expert-level taste and judgement' and 'a fraction of their cost' are the phrases that will travel, and neither has a figure behind it. The quantified part of the post ends below 80% accuracy for other people's models; a general per-dollar plateau is declared on the strength of one version-to-version comparison. The odd thing is that the plain finding — expert prompting bought 25 points and then stopped dead, and automatic optimisers could not find a 26th — is more useful than the framing wrapped around it.
The platform hosts the case for the platform
Read the recipe section and the arrangement becomes plain: the training ran on Tinker from Thinking Machines Lab, and the write-up lives on thinkingmachines.ai. Both sides of the comparison are interested — frontier APIs plateau in the mid-70s and cost 43% more per version, while a bespoke fine-tune on the house platform is said to beat them cheaply. None of that makes the frontier numbers wrong. It does explain which numbers were cleared for release and which were not.
Firm on the diagnosis, thin on the cure
The failure mode described is easy to credit: an analyst makes this call in a second and cannot say why, so the prompt captures the sayable part and misses the rest. The arithmetic around the ceiling holds up too. Confidence falls off at the hinge of the argument, where the story stops measuring and starts claiming, and it stays low because nobody outside this one desk has tried the six tasks.