Skip to content

Build2 publishers3 min readPublished

OpenAI bundles a 230 million-URL legal index behind the model ID gpt-6-astra-law

The legal toolkit ships as its own entry in the ChatGPT model picker and its own API string. Firms weighing it against retrieval stacks they already built still have no price to compare.

The Engineer · Build desk

Illustration accompanying OpenAI bundles a 230 million-URL legal index behind the model ID gpt-6-astra-law

What happened

  • OpenAI launched Astra for Law on September 17, packaging its GPT-6 Astra flagship with legal instructions, longer-work settings and a Legal Search Index for law firms and legal software vendors.
  • Access begins with selected firms through a Trusted Access program inside ChatGPT and Codex, and OpenAI says API access will follow.
  • OpenAI shipped 26 partner-built plugins and 47 community plugins alongside firm-built systems from Sullivan & Cromwell, Ropes & Gray and Cooley.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A team that already assembled legal retrieval now picks between a corpus it controls and a configuration OpenAI maintains and can change underneath it.
  • exposure Harvey, Legora, Thomson Reuters and iManage gain ChatGPT distribution while OpenAI supplies the model and the maintained legal configuration beneath their products; runtimewire argues standardization pushes their differentiation into content, access controls and audit trails.
  • contradiction Dev.to reports the index as making authoritative material more accessible inside an AI workflow, with lawyer review still required; runtimewire says the URL count does not establish how duplicates, subsequent history or conflicting authorities are handled. Those mechanics decide whether a cite can be relied on.

The detail that touches an engineering plan is the model string. OpenAI lists the legal version as GPT-6 Astra Law in the picker and as gpt-6-astra-law in the API [3]. A distinct ID means another model to qualify. Every eval, regression set and citation check that currently guards your pipeline gets rerun.

What the bundle covers is the cheap half of a legal RAG stack. In its launch thread on X, OpenAI described instructions for legal analysis and writing, settings intended for longer and more thorough work, and a new Legal Search Index [16]. Instructions and settings are the parts a competent team already keeps in version control. The index is the part that costs real money to build and more to keep current. It spans US case law, statutes, regulations, court rules and administrative decisions across more than 230 million URLs [6]. It draws on CourtListener data that OpenAI says covers a substantial share of published US precedential law [7].

The 230 million figure counts addresses in a crawl. Runtimewire's reading is that it describes the size of the search surface and does not establish how the index handles duplicate documents, subsequent legal history, conflicting authorities, or the difference between controlling and persuasive precedent [15]. For the number to transfer to your matters, the corpus would have to cover your jurisdictions and be current to the last filing date you care about. It would also have to collapse the same opinion appearing at forty addresses into one holding.

GPT-6 Astra itself shipped on September 3rd as OpenAI's model for demanding professional and computer-based work [13], fourteen days before the legal packaging [24]. Its API documentation lists a 1.05 million-token context window, tool support including web and file search, and reasoning settings from low to max [14]. Given the description of settings for longer and more thorough work [16], I would expect the legal preset to sit near the top of that range, though that is inference from the general docs rather than a published value.

The buy side cannot be priced yet. OpenAI has not published pricing, a broad API launch date or the eligibility requirements for Trusted Access [9]. The Zero Data Retention commitment is stated for API access by eligible law firms [8]. The first deployments run in ChatGPT and Codex [4]. If your index is CourtListener-derived and your product differentiates on workflow, I think the migration is worth planning for once API access and retention terms are contractual. If what you sell is subsequent-history logic and an audit trail you can defend to a partner, this bundle asks you to swap an inspectable component for an opaque one.

For vendors the position is more awkward. OpenAI shipped 26 partner-built plugins and 47 community plugins for legal work in ChatGPT [10], 73 in total [23]. Thomson Reuters, Harvey, Legora and iManage are among the named partners [11], and OpenAI says API customers such as Harvey and Legora will be able to build on Astra for Law [21]. Runtimewire argues that if the model layer becomes standardized, legal technology vendors will have to distinguish themselves through proprietary content, access controls, audit trails and workflow design instead of basic model access [20]. The three firm-built systems OpenAI named point at the same class of work. They are Sullivan & Cromwell's agreement analyzer, built with OpenAI engineers, Ropes & Gray's M&A diligence system, and Cooley's GO Public workflow for preparing companies to go public [12].

What to watch

  • Published API pricing for gpt-6-astra-law, and whether Zero Data Retention becomes a contractual term or stays a stated commitment.
  • Documentation of the index itself: update cadence, deduplication, subsequent-history treatment and jurisdiction coverage.
  • Whether Harvey and Legora build on Astra for Law's retrieval or keep their own index underneath their products.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories