Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

AI built stronger judgment only in senior patent attorneys over a 90-day trial

David Autor and Tanya Rodchenko's trial of 133 patent attorneys found AI lifted drafting quality for everyone but built stronger judgment only in seniors. Juniors' judgment scores split toward both ends with no average gain, so better drafts no longer show a firm which juniors are learning.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying AI built stronger judgment only in senior patent attorneys over a 90-day trial
Photo: research.google

What happened

  • AI access raised drafting scores by 0.34 standard deviations after 10 days and 0.38 after 90 days, about a 10- to 11-point gain in percentile rank against the control group.
  • The gains came from fewer poor drafts and more good ones, while the number of excellent drafts did not change.
  • Two-thirds of the lawyers at each firm got early access to the AI tool, and the rest got some AI training but no tool until the three months ended.
  • Independent legal experts graded each draft on Likert scales covering five separate aspects of quality.
  • The authors published the full results through the National Bureau of Economic Research as a paper on expertise in patent drafting.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Draft quality stops working as a progress signal for juniors, since the tool improved their drafts while their review judgment moved up for some and down for others.
  • decision A firm can justify rolling the tool out to senior lawyers on this evidence, but it needs separate evidence before treating the tool as a way to train juniors.
  • cost Firms buying the tool to get more top-tier drafts are paying for a gain the trial did not find, because the count of excellent drafts stayed flat.
  • contradiction Earlier studies disagree on whether AI-assisted work builds lasting skill, and this trial's juniors landed on both sides of that disagreement while using the same tool.

"Judgment is what sets senior professionals apart from their junior peers," David Autor and Tanya Rodchenko wrote [10]. In their account, it takes thousands of hours on the job to build, usually through tedious, routine work done under supervision [11]. The routine work in this trial was patent drafting. Each lawyer drafted a patent from simulated inventor materials after 10 days and again after 90 [8].

The drafting pattern was clearest for juniors, who also saved time [16]. They finished the 10-day draft in about 14.5% less time [19], 18 minutes under the control group's 124-minute average [16]. A grader reading an assisted draft sees a better document. The draft alone cannot show whether the junior could have produced it without help.

Patent lawyers also review each other's work, so the authors added a redline task at 90 days. Subjects had to mark up and correct a hypothetical patent seeded with substantive and stylistic errors [6]. The critical mistakes, sometimes called "patent profanity," include overclaiming an invention's novelty or scope, and they can make a patent unenforceable in court [7].

On that task, senior lawyers who had used the tool for 90 days showed stronger judgment than seniors who had not [9]. Junior judgment did not improve on average, and more juniors landed at each end of the scale [2]. The authors' headline finding does not put a size on either tail.

These are results from one workload. The tool was a then-unreleased Google Labs patent assistant, now part of Gemini Notebook, and all eleven firms do regular, non-exclusive business with Google [14]. Two-thirds of 133 is about 89 lawyers with the tool and 44 without [20]. The junior and senior findings come from subsets of those groups. For the numbers to carry to another team, its juniors' routine work would have to resemble patent drafting, and its review test would have to look like the redline.

The authors argue that the question needs months of sustained use in ordinary work, a credible test of unassisted judgment, and blinded grading by domain experts, and that the three are rarely found together [18]. Three months of routine use [1] is the expensive part. I think it is the main reason this result deserves more weight than a short lab session would get.

The authors place the finding against two sets of earlier studies. In radiology, business problem-solving, job-seeker writing and legal education, AI built into training improved less-experienced workers' independent performance [12]. With software engineers, management consultants, high school students and clinicians, gains often faded once the tool was removed. There, the authors wrote, AI "serves as a temporary exoskeleton that boosts immediate output" [13]. The juniors in this trial show both outcomes inside one cohort, using one tool.

This trial randomized access to the tool [14]. Positive training results come from the other studies, where the tools were "thoughtfully embedded into training workflows" [12]. So the case for redesigning how juniors use AI rests on those studies, plus this trial's evidence that access alone left average junior judgment flat [2]. In my view, a firm that trains juniors through routine drafting needs an unassisted review test, something like the redline, to find out which end of the scale each junior is on.

What to watch

  • Whether the full NBER paper sizes the junior split on the redline task, and where it draws the line between junior and senior lawyers.
  • A follow-up that gives juniors the tool inside a structured training workflow, testing the embedded-training result the authors cite from radiology and legal education.
  • Replication at firms that do not do regular business with Google, using a drafting tool other than the Gemini Notebook assistant.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories