YFarmX

AI News

OpenAI's Astra for Law puts a case-law index behind GPT-6 Astra

OpenAI introduced Astra for Law on 17 September 2026: GPT-6 Astra with a legal search index, legal instructions and firm tooling, for selected law firms. Two separate figures in the announcement both read 54%, and they measure different things.

Editorial hero: a bound legal reporter volume lying open on clear newsprint, headed ASTRA FOR LAW with the subtitle OPENAI · SELECTED LAW FIRMS

OpenAI introduced Astra for Law on Thursday 17 September 2026, describing it as “a new foundation for law firms and legal technology companies to build AI products and workflows around their expertise.”

The product is a configuration of a model OpenAI already ships. Its own description: it “combines GPT‑6 Astra, our latest and most powerful model, with settings, tools, and context tailored for professional legal work.” GPT-6 Astra shipped on 3 September 2026 at $10 and $50 per million tokens. Astra for Law appears in the ChatGPT model picker as “GPT‑6 Astra Law” and, when it arrives, in the API as gpt-6-astra-law.

What is new is underneath: a legal search index, a set of instructions for legal analysis and writing, and a route into the tools firms already run.

The case law comes from a nonprofit

The index is the substance of the announcement. OpenAI says Astra for Law “can search U.S. case law, statutes, regulations, court rules, and administrative decisions across a corpus of more than 230 million URLs, with sources added daily.”

The case-law portion has a named source. OpenAI credits “our work with Free Law Project, the nonprofit behind CourtListener”, which it says “brings its case-law collection covering more than 99.9% of published U.S. precedential case law” into the research experience.

That 99.9% figure is worth attributing correctly, because it is CourtListener’s claim about its own database rather than a measurement OpenAI performed. Free Law Project’s coverage page puts it in almost identical words: “Our database encompasses more than 99.9% of all precedential legal case law published in the United States and we are working to close the last 0.1% by scanning decisions directly from books.” The page dates that assessment to 2023.

The opening of CourtListener's coverage page. Under the line that CourtListener has one of the most comprehensive collections of American case law on the Internet, the lead paragraph states that as of 2023 the database encompasses more than 99.9% of all precedential legal case law published in the United States, and that the project is working to close the last 0.1% by scanning decisions directly from books. The Diverse Sources section below names Public.Resource.org, the Supreme Court Database, Columbia, Harvard, the Law Library of Congress, the Caselaw Access Project and a partnership with vLex.
The 99.9% coverage claim at its origin: Free Law Project's own page, describing its own collection. Source: CourtListener.

The category named is published precedential opinions, and that boundary does real work in practice: unpublished and non-precedential decisions sit outside it.

Free Law Project has been here before, with a competitor. On 12 May 2026 it announced that “CourtListener is now available inside Claude as a new MCP Connector”, putting the same underlying case-law data behind Anthropic’s model four months before this launch. The data partner is shared. The index OpenAI has built on top of it is its own.

There are two 54% figures and they measure different things

The announcement carries two numbers that both read 54%, close together, describing different quantities. A reader skimming will merge them.

The first is a pass rate. OpenAI says it tested “Astra for Law’s complete setup on 200 U.S. legal research questions from the private validation set of Vals AI’s Legal Research Bench”, and reports that “at the highest reasoning effort for both systems, Astra for Law passed the evaluation’s overall correctness check on 54.0% of questions, compared with 38.7% for GPT‑6 Astra using web search alone – a 40% relative improvement.”

The second is a retrieval measure, and it is relative rather than absolute. On case-law questions, OpenAI says Astra for Law “found 24% more reference cases than GPT‑6 Astra using web search alone at the highest reasoning effort”, and that “on the audited set of target passages, it retrieved up to 54% more relevant passages from the correct court opinions, when comparing the systems at the same reasoning effort.”

So: 54.0% of questions answered correctly, and up to 54% more passages found than the same model working from the open web. Both are measures of retrieval and correctness on a research task. A reader wanting to know how often a citation is accurate is looking at a different quantity, and the announcement leaves it to the courts below.

The public leaderboard is a different test

Vals AI’s Legal Research Bench has a public leaderboard, and OpenAI measured against something else. The 200 questions above come from the benchmark’s private validation set, which Vals holds back.

The public board, updated on 15 September 2026, two days before this launch, shows a three-way tie at the top on the bench’s primary metric: Muse Spark 1.3 Max, Claude Opus 5 and Claude Fable 5.1 all reach 55.29% all-pass accuracy. Astra for Law does not appear on it. GPT-6 Astra, the base model, is among the models Vals has evaluated.

Placing OpenAI’s 54.0% alongside that 55.29% would be a mistake, because the two numbers come from different question sets. The comparison the announcement actually supports is the one it makes: Astra for Law against GPT-6 Astra, on the same private set, at the same reasoning effort.

Vals explains why it grades the way it does, and the reasoning is worth repeating: “We report all-pass as our primary metric because, in legal work, a partially correct answer can be more dangerous than a wrong one: it may read as sound while omitting a critical point.” On the public board the gap between grading styles is wide. Claude Opus 5 reaches 90.58% under weighted partial-credit scoring and 55.29% when every rubric check has to pass.

The Vals AI Legal Research Bench page, updated 15 September 2026, showing the key takeaways: a three-way tie at the top between Muse Spark 1.3 Max, Claude Opus 5 and Claude Fable 5.1 at 55.29% all-pass accuracy, a note that Claude Opus 5 reaches 90.58% under partial-credit scoring against 55.29% under strict all-pass grading, and observations on practice-area variation and tool use.
The public Legal Research Bench two days before the launch. The questions here are a different set from the private validation set OpenAI tested on. Source: Vals AI.

The comparison is against web search alone

Both of OpenAI’s headline gains are measured against “GPT‑6 Astra using web search alone”. That is a fair way to isolate what the index contributes. It is a comparison of the model against itself, and its scope stops there: no competing legal research product is in the measurement.

The distinction has teeth here, because Vals’ own public benchmark does not run its models on web search alone. Its agents work with “case-law search, web search, document retrieval, and HTML parsing”, and the page records that “web_search is the most-used tool across models, averaging 26 calls per session versus 11 for courtlistener_search”. Every model on that public board already has a case-law search tool in hand.

An animation separating the two 54% figures in OpenAI's announcement. The first panel shows the pass rate: on 200 questions from the private validation set of Vals AI's Legal Research Bench, at the highest reasoning effort, Astra for Law reaches 54.0% overall correctness against 38.7% for GPT-6 Astra with web search alone, a 40% relative improvement. The second panel shows the retrieval measure: on case-law questions Astra for Law finds 24% more reference cases, and on the audited set of target passages retrieves up to 54% more relevant passages from the correct court opinions, at equal reasoning effort. A closing panel notes that the public Legal Research Bench leaderboard is a different question set, topped on 15 September 2026 by a three-way tie at 55.29% all-pass.
Two numbers, both 54%, measuring different things. The baseline in each case is the same model without the index.

Only selected firms get it

Astra for Law “will be initially offered to selected law firms through Trusted Access in ChatGPT and Codex, and will be coming soon to the API.” Access runs through what OpenAI calls “a special Trusted Access Program for eligible law firms”, and the page keeps the eligibility criteria in house.

For firms that qualify, the terms include “Zero Data Retention (ZDR) on our API”, and ChatGPT Enterprise usage “excluded from human review by default”.

On price, the page offers one thing: the benchmark charts plot model cost per answer on an axis running from $0.0 to $6.0. The base GPT-6 Astra rate of $10 and $50 per million tokens remains the published figure.

What the firms built

Four firms appear in delivery roles, working with what OpenAI calls its forward-deployed engineers on ChatGPT Enterprise.

Sullivan & Cromwell built an agreement analyser which, in OpenAI’s description, “brings the firm’s negotiating playbooks and selected precedents into the review of a new deal”, producing proposed redlines and draft client advice. Ropes & Gray built a diligence system shaped around how its lawyers work through a data room, designed to trace findings back to source. Cooley built a tool called GO Public for capital markets work, from drafting the IPO filing onwards, which carries a change through the document when the deal moves.

Dave Peinsipp, partner and co-chair of Cooley’s global capital markets group, is quoted on what that changes: “Our collaboration with OpenAI has allowed us to rethink how this work gets done – moving lawyers and management teams more quickly through intensive preparation and into questions that require judgment, market experience and strategic thinking.”

Latham & Watkins appears in a different role, working with OpenAI “to design for information permissions, ethical walls, client instructions, and firm oversight”. Wachtell, Lipton, Rosen & Katz is named on a research collaboration.

Harvey and Legora are API build partners. Niko Grupen, head of applied research at Harvey, is quoted saying that in early testing Astra for Law “showed strength across key aspects of legal research: grounding answers in on-point authorities, citing with precision, and offering practical, advisory guidance.”

Twenty-six partner-built plugins launch with it, connecting ChatGPT to tools firms already run. OpenAI names Relativity, Clio, iManage, Intapp, DeepJudge and Thomson Reuters among them. A further nine community plugins come from LegalQuants, LECG and Skills.law, carrying 47 custom skills.

Thomson Reuters occupies both sides of the board: OpenAI positions the index as something that “complements the licensed content and specialist products firms rely on from providers such as Thomson Reuters”, while Thomson Reuters is bringing HighQ into ChatGPT and previewing a CoCounsel Legal connector.

OpenAI names a rival model in its own demonstration

The announcement includes a worked comparison on a misrepresentation claim, and it names the model it beat. OpenAI’s summary: “Given the same prompt, Astra for Law returned two closely matching precedents; in the litigation example, Claude Fable 5.1 returned a holding that had been reversed on appeal, while in the transactional example it reported finding no such case.”

That is OpenAI’s account of a single prompt, run by OpenAI, published by OpenAI, on a comparison it chose. It is presented on the page as an illustration rather than a measured result, and it sits apart from the benchmark figures above.

The courts already ask lawyers to swear their citations are real

The sharpest test of a product built around grounded citations is waiting in the federal courts, and it predates this launch.

Judge Nina Y. Wang of the US District Court for the District of Colorado has required, effective 1 December 2025, that “every filing shall contain an AI Certification regarding the use, or non-use, of generative AI”. Where AI was used, each individual must certify that AI-drafted language “was personally reviewed by the filer or another human for accuracy and that all legal citations reference actual non-fictitious cases or cited authority.” Filings that skip the certification “may be stricken without substantive consideration”.

The order names its examples of generative AI as “ChatGPT, Harvey.AI, or Google Gemini”. Harvey is one of the two API partners OpenAI names in this announcement.

At the US Court of International Trade, Judge Stephen Alexander Vaden’s order takes a different angle, aimed at confidentiality rather than fabrication. It notes that generative programs “create novel risks to the security of confidential information”, because users “may include confidential information in their prompts, which in turn may result in the corporate owner of the program retaining access to the confidential information”. The Zero Data Retention terms in the Trusted Access programme are aimed squarely at that objection.

What a firm is being offered

The shape of the offer is clear enough. A frontier model, an index over more than 230 million URLs with published precedential case law supplied by a nonprofit, custom legal instructions, plugins into the document and time systems firms already run, retention terms written for privileged work, and a 15-point gain on a 200-question private set against the same model searching the open web.

The question a firm will ask next is how often the citations hold up. Judge Wang’s order has already settled who answers it: the lawyer whose name is on the filing, in a certification signed for every document that goes to the court.

Sources

  1. Introducing Astra for Law (OpenAI, 17 September 2026)openai.com
  2. Legal Research Bench (Vals AI)vals.ai
  3. Coverage of opinions in CourtListener (Free Law Project)courtlistener.com
  4. CourtListener is now available inside Claude (Free Law Project, 12 May 2026)free.law
  5. Standing Order Regarding the Use of Generative Artificial Intelligence in Court Filings, Judge Nina Y. Wang (US District Court, District of Colorado)cod.uscourts.gov
  6. Order on Artificial Intelligence, Judge Stephen Alexander Vaden (US Court of International Trade)cit.uscourts.gov

How we use AI