Skip to content

Product1 publisher3 min readPublished

OpenAI's legal index lifted GPT-6 Astra from 38.7% to 54% on 200 bench questions

Astra for Law is GPT-6 Astra with a legal search index and analysis instructions wrapped around it. OpenAI will also sell that index through an API to Harvey and Legora, at a price it has not published.

The Product Desk · Product desk

Illustration accompanying OpenAI's legal index lifted GPT-6 Astra from 38.7% to 54% on 200 bench questions

What happened

  • OpenAI launched Astra for Law: GPT-6 Astra wrapped in a legal search index and instructions for legal analysis, sold to firms and to the software companies serving them.
  • The index reaches U.S. case law, statutes, regulations, court rules and administrative decisions across more than 230 million URLs, with case law supplied by the Free Law Project's CourtListener.
  • It passed the overall correctness check on 54% of 200 questions from the private validation set of Vals AI's Legal Research Bench, against 38.7% for the same model using web search.
  • Only selected firms have access, through a Trusted Access program inside ChatGPT and Codex, and the gpt-6-astra-law API is due later with no date or pricing attached.
  • Thirty-five plugins shipped alongside the model, Thomson Reuters is bringing HighQ matter context into ChatGPT, and ChatGPT for Word reached general availability the same day.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost OpenAI has published no API price, so a firm cannot compare the cost of renting OpenAI's index against what it already pays a vendor to maintain one. The build-or-buy question stays open on the desk.
  • exposure Vendors selling contract review, diligence and IPO preparation are now pitching firms that have those same tools running inside their own ChatGPT Enterprise tenant, built by OpenAI staff.
  • constraint Access runs only to the picked firms for now, so the operator's realistic Monday option is a pilot request.
  • contradiction OpenAI's own benchmark credits an index and a prompt with a 15.3-point gain while Thomson Reuters' CTO argues connectivity alone is not the value, and the two readings point a buyer at different products.

An associate at one of the selected firms opens the model picker in ChatGPT and finds a new entry, GPT-6 Astra Law [7]. The weights are the same GPT-6 Astra that began rolling out on Sept. 3 [20]. The benchmark OpenAI chose rewards finding the right authority and the right passage inside it [4], and on that test the indexed configuration beat the plain one by 15.3 points [1].

That gap is the product. Nothing was retrained; the model got a search index and a set of instructions [1]. OpenAI says the configuration surfaced 24% more reference cases on case-law questions, and on a separate audited set retrieved up to 54% more of the target passages from the correct opinions [6]. It still got 92 of the 200 questions wrong [2]. Wrong answers fell from 61.3% to 46%, about a quarter of the failures [3].

The baseline figure does not divide. 38.7% of 200 questions is 77.4 questions; 77 would be 38.5% and 78 would be 39% [4]. OpenAI does not say what the denominator was.

The buying decision happens a level up from the associate. Harvey and Legora, two of the companies that sell legal AI into firms, are named as API customers for gpt-6-astra-law [9]. So the retrieval advantage a vendor could present as its own engineering becomes a line item those vendors rent, at a rate OpenAI has not disclosed [8]. OpenAI called the launch the start of a long-term investment in law [21].

Thomson Reuters is bringing HighQ matter context into ChatGPT and previewing a connector for its CoCounsel Legal product [13]. "As AI becomes more open and interoperable, the value is not in connectivity alone," said Joel Hron, the company's chief technology officer. "Legal professionals need more than access to information. They need trusted intelligence, relevant enterprise and matter context, purpose built legal capabilities, and the governance required for high stakes work." [16]

Some of that governance is still being drawn. Latham & Watkins is working with OpenAI on governance design for information permissions, ethical walls and client instructions [11]. Eligible firms get zero data retention on the API, and ChatGPT Enterprise usage is excluded from human review by default [10]. A firm holding both sides of a matter will want the wall specified before the pilot starts.

OpenAI engineers embedded at individual firms have been building on ChatGPT Enterprise. Sullivan & Cromwell has an agreement analyzer that pulls its negotiating playbooks and selected precedents into the review of a new deal and turns what it finds into proposed redlines [17]. At Ropes & Gray the work went into deal diligence, and Cooley's tool, GO Public, handles initial public offering preparation, drafting the filing included [18]. John Savva, a Sullivan & Cromwell partner who saw early builds, said the models showed "impressive research depth and sensitivity to authority" [19].

The index is one thing on sale in this launch, and OpenAI will rent it to anyone once the API has a price. Matter context is another: a firm's own documents, permissions and ethical walls. The plugin set is reaching for that [12][14]. The third is the workflow, where a redline or a draft filing comes out. Every line on a vendor invoice belongs in one of the three. Anything in the first column is resold gpt-6-astra-law, at a markup only OpenAI can see today.

What to watch

  • A date and a price for the gpt-6-astra-law API, which is what lets a firm compare renting the index against paying a vendor.
  • What Latham & Watkins and OpenAI publish on ethical walls and information permissions before firms move client matters in.
  • Whether Vals AI releases the full Legal Research Bench run, including the question count behind the 38.7% web-search baseline.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories