A law firm needs more from AI than a convincing memo. It needs citations an attorney can verify, analysis that fits the facts, and a process that protects client information. Comparing AI models for law firms starts with those requirements.
OpenAI's September 17 announcement introduces Astra for Law: GPT-6 Astra paired with a legal search index and legal-specific instructions. Initial access is limited to selected firms; API access is forthcoming. OpenAI announcement
For a managing partner choosing where to invest, the useful question is which setup will improve a particular job. Researching precedent, reviewing a contract, and preparing an intake summary require different tests. Here's how to evaluate the options in the announcement—and what its comparisons actually establish.
What the research comparison shows
OpenAI reports testing 200 U.S. legal research questions from Vals AI's private validation set. At the highest reasoning effort, Astra for Law passed the overall correctness check on 54.0%, versus 38.7% for GPT-6 Astra with web search alone. That is a 15.3 percentage-point gain, or roughly 40% relative improvement. Evaluation details
Treat those figures as results for the tested configurations. They do not mean every answer is 54% accurate, or that a firm will save 40% of its research time. They also do not establish performance across every jurisdiction and practice area.
The result makes a useful case for testing research access alongside reasoning ability. When evaluating a system, ask attorneys to score both the authorities it finds and how accurately it applies them. A plausible argument supported by an irrelevant case should fail that review.
Comparing AI models for law firms
The comparison below is our suggested evaluation framework, not a ranking from a hands-on Flying Dog Media benchmark.
- Astra for Law: Start with research tasks involving difficult precedent searches. Measure relevant authorities, citation support, treatment of adverse cases, and attorney correction time.
- General-purpose GPT-6 Astra: Start with document preparation, research support, and internal workflows. Measure fidelity to supplied material, instruction following, and usefulness after review.
- Claude: Start with playbook-based contract review, document analysis, and drafting. Measure missed provisions, unsupported edits, source traceability, and consistency across documents.
Astra for Law: OpenAI also presents selected comparisons with Claude Fable 5.1, reporting a reversed holding in one Claude example and a missed precedent in another. These vendor-selected examples warrant investigation; they do not establish a universal winner. Comparison examples
General-purpose GPT-6 Astra: OpenAI describes the base model as supporting complex reasoning, research, and document creation. Those capabilities make it a candidate for a broader operational pilot. A sensible test might ask it to turn approved matter notes into a structured internal update, with every statement linked back to the notes. GPT-6 Astra documentation
Claude: Anthropic's legal guide describes uses including contract review against a playbook, diligence summaries, deposition preparation, and drafting. That supports including Claude in a document-workflow trial. It does not supply a comparable score against Astra for Law. Evaluate the exact Claude model and product configuration available to your firm. Anthropic's legal guide
For example, give each candidate the same sample agreement and approved negotiation playbook. Ask it to identify deviations, cite the relevant clauses, and propose edits. Then have an attorney assess whether those edits preserve the business deal. That is a more useful purchasing exercise than asking which assistant writes the smoothest paragraph.
The product around the model matters
A model comparison leaves several buying questions unanswered. Which documents can the system access? Can it preserve matter boundaries? Can an attorney open the source behind a statement? Where does the reviewed output go?
Record the complete setup for every trial: model version, product, search access, connected data, instructions, and settings. Changing any of those can change what you are evaluating. A general assistant with public web search and a legal application with licensed research access are different configurations, even when their underlying models match.
Confidentiality also needs a product-level review. Before using client information, establish:
- Whether inputs and outputs are used for training.
- How long prompts, files, outputs, and logs are retained.
- Who, what systems, and which connected services can access them.
- Whether permissions follow the firm's matter restrictions.
- What deletion, export, and audit options the contract provides.
The ABA's Formal Opinion 512 addresses competence, confidentiality, communication, and other duties when lawyers use generative AI. Its guidance reinforces the need to understand a tool's limitations and review its work. Firms should also check applicable jurisdictional requirements and client instructions. ABA Formal Opinion 512
Run a pilot that measures usable work
Start with one recurring task and a small set of public, synthetic, or appropriately approved materials. Choose examples with known answers, including a few difficult exceptions. Use the same assignment and scoring criteria for each candidate, and record any differences in available tools.
Have a lawyer review the outputs without model labels where practical. Score factual accuracy, material omissions, citation quality, and the time needed to produce an acceptable result. For contract work, include cross-references and exceptions; for research, include adverse authority and jurisdictional relevance.
Calculate the cost per accepted result. Include subscriptions or usage charges, setup, attorney review, and correction time. An inexpensive answer that takes substantial repair may cost more than a slower, more useful draft.
Decide the acceptance criteria before the trial. A research assistant might need verifiable support for every material proposition. An intake summarizer might need to preserve names, dates, and uncertainties without inventing facts. Those standards should reflect the task's consequences.
For firms in Charlottesville and across Virginia, we recommend starting with a narrow workflow, comparing the available candidates, and expanding only after the results justify it. Astra for Law deserves consideration for research; general-purpose Astra and Claude deserve testing for the work your team repeats every week.
Flying Dog Media helps firms plan and build AI workflows for intake, drafting, and automation. Bring us one process your team wants to improve. We can help define the evaluation and turn the findings into an implementation plan.
Product information checked September 23, 2026. This comparison draws on vendor documentation, not an independent hands-on benchmark.
Ready to put AI to work?
Tell us what you're trying to solve. We'll get back to you within one business day.
Start a Conversation →