Build1 distinct publisher3 min readUpdated
A theoretical paper assumes error-free, near-free language models and still finds that in two of three deployment scenarios, researchers respond by doing more work with less care.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The load-bearing part of the model is which stage of a project has someone enforcing it. Pushing a project forward splits into mandatory work, such as making figures, formatting and submitting, and voluntary work, such as extra experiments, deeper analysis or polishing the prose [5]. A venue enforces the first. Nothing enforces the second except the researcher's own judgement, and that is precisely what gives way once the next project starts to look cheap [6].
That asymmetry is why the direction the tooling points matters more than how much time it saves. Accelerate early triage and researchers get choosier, because abandoning an idea costs less, but the survivors still get thinner treatment, since the freed hours pay better in a new project [7]. Accelerate the write-up and the bar for "worth submitting" drops, so marginal projects go out and the average paper is shallower [8]. Only when the tool speeds the voluntary stage does the saving land where the corner-cutting already happens [9].
Hold that taxonomy against the field evidence in the same write-up. OpenAI's report on eight case studies claims up to 60x speedups when rewriting research software, with the bottleneck moving to validation and long-term maintenance [11]. Validation is voluntary-side work in the paper's sense: no reviewer form demands it, and no deadline is missed by skipping it. So the measured win sits in the mandatory column and pushes load into the column the model says gets sacrificed first. That is scenario two with a compiler attached.
The METR result is the sharper test, because it removes the need for the time saving to exist. Experienced open-source developers using AI tools took 19 percent longer while reporting they felt 24 percent faster [12], a 43 point spread between measured and felt speed [13]. The write-up notes the behavioural point directly: perceived savings do not have to be real to change how people allocate effort [14]. An incentive that keys off belief is not fixable by making the model better, and it is not detectable by asking users whether the tool helped.
None of this depends on hallucination, cost or model error. The authors deliberately treat language models as error-free and financially negligible in order to isolate the time effect [3], and the pessimistic result survives that generosity [4]. Their own summary is that as a labor-augmenting technology, LLMs "increase the opportunity cost of our time, impelling us to do more, less well" rather than "the same amount, better" [10]. Read as a procurement note rather than a lament, it says something narrower and more usable: a tool that shortens the parts of the job someone else checks buys throughput, and only a tool that shortens the parts nobody checks can buy care.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
A theoretical economics paper by researchers from Princeton, the University of Washington and other institutions argues that because AI saves time, researchers will spend less effort on each project rather than more.
The model is built on optimal foraging theory from behavioral ecology, a framework describing how organisms allocate effort across competing opportunities, adapted to simulate how researchers distribute labor across projects when LLMs shorten different phases.
The authors deliberately idealize LLMs as tools that cut time costs without introducing errors and at negligible financial cost, to isolate the pure effect of time savings.
In two of the paper's three scenarios thoroughness drops; only one scenario improves quality.
In the model a project has two phases: the researcher checks whether an idea is viable, then either abandons it or pushes forward; pushing forward has a mandatory part (creating figures, formatting text, submitting) and a voluntary part (running extra experiments, deeper analysis, polishing prose).
The authors argue the voluntary part is what gets sacrificed when time becomes scarce, because polishing an already publishable paper stops making sense once time is more valuable.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Stylized model, thin sourcing
The core finding is an analytical result from a deliberately idealized model (error-free, near-costless LLMs) reported by a single publisher, with no paper title, venue or peer-review status supplied. Its direction is corroborated only indirectly by two secondhand empirical items — OpenAI's own field report and the METR slowdown study — neither of which tests the paper's mechanism. One asserted trend, rising submissions straining peer review, carries no data at all.
Behavior signals, no uptake of the model
There is no evidence that the paper's framework has been adopted by institutions, funders or venues; the article only notes that responses 'need to be discipline-specific'. What is observable is adoption of the underlying practice the model reasons about: an OpenAI field report of eight AI-assisted research-software cases, a METR study of developers using AI tools, an AI-generated paper clearing an ICLR workshop, and arXiv tightening submission rules.
Slightly overstated
The framing that better AI makes each paper worse is stronger than the underlying support: the result is a scenario property of one stylized model, and one of its three scenarios points the other way. The write-up does label the work theoretical and does supply a counter-signal (METR's measured slowdown against perceived speedup), which keeps the overstatement modest rather than severe. The gap is mostly in generalizing model scenarios into a claim about science as a whole.
Mixed: vendor self-report plus subscription outlet
Two identifiable incentive structures appear in the supplied material. The 60x speedup figure comes from OpenAI's own field report on AI-assisted research software, a vendor reporting on its own category of tools. The reporting outlet closes with a paid-subscription pitch positioned as 'AI News Without the Hype', which rewards contrarian framing of AI benefits. The academic authors' own funding or affiliation incentives are not disclosed in the source, so they are not scored.
Low-moderate
One publisher, one unnamed theoretical paper, and no independent corroboration of the model or of the peer-review strain claim. Confidence is high that the article accurately describes the paper's structure and scenarios, and moderate on the two cited empirical results, but low that the world-level conclusion holds as stated.
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
product
Nvidia's $6bn Poolside licence is the third run of the same play2 distinct publishers
build
The AI-training bans live on the big infrastructure blogs, not the small publications1 distinct publisher
build
MathCode's bet is that proofs should accumulate, not evaporate1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 23, 2026