How Frontier AI Models Perform on Real Legal Work

A legal-only bench marking study designed and graded by qualified attorneys across various legal task types

Executive Summary

Percipient designed and administered a structured benchmark to measure how leading large language models perform on the work of lawyers and legal professionals. The study was loosely based on the GDPVAL evaluation framework and adapted for the specific requirements of legal practice, where reasoning, citation accuracy, and faithful adherence to controlling authority matter.

The benchmark exercise included legal task types across various practice areas: litigation, contract redlining, employment, and insurance coverage. Each task asks the model to produce an actual deliverable a lawyer would deliver to a client or senior partner: a memo, a coverage opinion, a redlined contract, or a coded document review set with a privilege log.

Outputs from frontier models across Anthropic, OpenAI, Google, xAI, Moonshot AI, and DeepSeek were collected and blindly graded by qualified legal professionals using detailed, pre-determined rubrics. The results show meaningful and consistent differentiation between models, with reasoning-focused and extended-thinking variants outperforming standard configurations on the most demanding analytical work.

This article describes the study design, methodology, scoring framework, models evaluated, and the results.

Many public benchmarks for artificial intelligence models may not optimally grade how AI performs on legal work. This is so because the LLM benchmarking exercises are often combined with other disciplines, are self graded or rely on bar-exam style questions.

While these benchmarking exercises are thoughtful and valuable, the nuance of true legal work is hard to master and measure. A practicing lawyer is not asked to choose between four multiple choice options. Lawyers are asked to read a complaint, an insurance policy or employment file, identify the issues that matter and apply the controlling authority to produce a deliverable that a client or senior partner may rely on. Whether a model can do this, and how reliably, is the question Percipient set out to answer.

The benchmark design was based on four pillars:

    • Real-world legal issues. Every task is grounded in a realistic fact pattern, with reference materials a practicing lawyer would actually use: insurance policies, contract playbooks, employment files, and binding case law. Material for each task also includes deliberate noise: facts or documents that look topically relevant but are not germane to the legal task. This tests whether a model is applying genuine legal judgment or pattern-matching on surface features.

    • Multimodal deliverables. For the benchmarking exercise. the models are not asked to just summarize legal documents or provide high level legal analysis. The benchmark prompts require an output a lawyer would actually deliver–a coverage opinion, an employment law memo, a coded document review with privilege log, or a redlined contract with comments. Accurate format is part of the score, because in practice a messy deliverable has diminished value.

    • Blinded, rubric-based grading. Rubrics were created by legal professionals and finalized before output review. Graders scored each deliverable without knowing which model produced it. Where two graders scored the same set, their passes are kept independent and reconciled only after both were complete.

    • Qualified legal professionals. All model output and deliverables were graded by attorneys with an average of over 25 years of legal experience. Domain experts were used for each corresponding area of law.

Each task includes a structured prompt, supporting reference materials, and a defined deliverable format. The legal tasks are summarized below.

Task Description Primary Grading Dimensions
Insurance Coverage Memo Analyze a duty-to-defend scenario under Illinois law involving a construction defect / water-infiltration suit and produce a coverage memorandum  Eight-corners application; occurrence analysis; subcontractor cases; business-risk exclusions; reservation of rights; legal accuracy and absence of hallucinations; format and writing
Employment Law Memo Analyze a former employee’s potential discrimination claims under the ADA, Title VII, and the ADEA, including a circuit split on reassignment as accommodation. Recommend a motion-stage posture and a settlement strategy responsive to the client’s framing. Procedural and jurisdictional issues; ADA accommodation analysis; termination claims; motion-stage and trial likelihood; settlement strategy; quality control and noise resistance; format and writing
Litigation Document Review Code 78 simulated litigation documents based on a real world legal matter for responsiveness, privilege, and issue tags. Produce a privilege log for privileged items. Coding accuracy across responsiveness, privilege and issue tags; reasoning quality; privilege log format and sufficiency
Contract Review and Redline Review a vendor agreement against a defined Vendor Contract Playbook. Identify each clause that deviates from playbook position, generate redline language, and provide reasoning for each change. Deliver a redlined .docx with comments. Edit completeness; edit precision; reasoning aligned with playbook; absence of hallucinations; formatting of redlined deliverable and comments

Models Evaluated

The evaluation measured models from Frontier AI companies. As each task was administered, the then-current production version of each model was used. Where extended-thinking or reasoning variants were available, they were included alongside their standard counterparts.

The employment and contract review benchmarks, which were administered after the coverage and document review work, used updated model versions. The master dataset preserves these distinctions in a private key so that grading remains blind by memo number and is not biased by knowledge of which specific model produced a given output.

Provider Model Family Variants Evaluated
Anthropic Claude Opus 4.6, Opus 4.6 (extended), Sonnet 4.6, Opus 4.7 (Adaptive)
OpenAI GPT / ChatGPT 5.4 Pro (standard), 5.4 Pro (extended), 5.4 Thinking (extended), GPT 5.5 Thinking
Google Gemini 3.0 Pro, 3.1 Pro
xAI Grok 4.20, 4.3 (beta)
Moonshot AI Kimi K2.5, K2.5 Thinking, K2.6 Thinking
DeepSeek DeepSeek Standard

Methodology

Evaluation methodology followed a structured workflow designed to minimize bias at every stage.

# Phase Description
1 Prompt drafting A two-person team took ownership of each task. One person drafted the prompt, supporting reference materials, and the deliverable specification.
2 Peer review Person Two reviewed the prompt for clarity, completeness, and any ambiguity that might skew model outputs. Comments are returned to Person One.
3 Prompt finalization Person One finalized the prompt used uniformly across all model evaluations.
4 Rubric development Person Two independently developed the scoring rubric.
5 Blind model evaluation The prompt was run through each model in the roster and outputs compiled  and anonymized.
6 Independent scoring Graders scored deliverables using the rubric, blind to authorship. 

Grading Framework

Each rubric was task-specific and used partial-credit scoring at the sub-task level. Sub-task grading roll up into section subtotals, and section subtotals roll up into a 100-point overall total. Rubrics were anchored to specific authority and specific facts in each prompt so that grading reflects whether the model did the legal work the task required, not whether the writing was generally sound.

Coverage Rubric

Section Max Representative Criteria
1. Duty to Defend / Eight Corners 15 Eight-corners rule; scope of duty; potential coverage standard
2. Occurrence 10 Acuity v. M/I Homes; own work vs. other property; complaint scope
3. Property Damage / Subcontractor Cases 15 JP Larsen; West Van Buren / Metropolitan Builders; 950 W. Huron
4. Business Risk Exclusions 15 Coverage A exclusions (j)/(k)/(l)/(m); subcontractor exception; PCOH
5. Reservation of Rights 15 Duty-to-defend conclusion; RoR analysis; Ehlco estoppel and DJ option
6. Legal Accuracy 15 Hallucinations; irrelevant or incorrect arguments
7. Format, Professionalism, Writing 15 Memo format; conservative tone; balanced authority; citations

 

Employment Rubric

Section Max Representative Criteria
1. Procedural and Jurisdictional 10 Administrative exhaustion; state law claims; circuit analysis
2. ADA Failure to Accommodate 20 Prima facie framework; qualified individual; reassignment / EEOC v. UAL; interactive process; Dark v. Curry County noise test
3. Termination Claims 15 ADA termination; ADA retaliation; Title VII race and gender; ADEA
4. Motion Stage and Trial Likelihood 10 Motion to dismiss; summary judgment; trial outcome assessment
5. Settlement Strategy and Risk 15 Direct recommendation; damages exposure; cost of litigation; intangible costs; negotiating posture
6. Quality Control 15 Hallucinations; irrelevant arguments and noise resistance; assumptions and missing information
7. Format, Professionalism, Writing 15

Memo format; outline / bullet structure; Bluebook citations; tone; organization

 

Document Review Rubric

Section Max Representative Criteria
1. Coding 80 Thoroughness; issues identified; completeness ;accuracy; reasoning; redlined .docx; comments
2. Privilege Log 20 Format; reasoning; sufficiency

 

Contract Review

Criterion Max What is Evaluated
1. Edit Complete 20 Whether the redline captures every change the playbook requires for each clause
2. Edit Precise 20 Whether the redline language is precisely worded and tracks the playbook position
3. Reasoning Aligned with Playbook 20 Whether the stated reasoning for each redline matches the playbook rationale
4. No Hallucination 20 Whether the output is free of invented contractual terms, authorities, or facts
5. Format — Redlined .docx 10 Whether a usable redlined .docx was generated
6. Format — Inline Comments 10 Whether inline comments were generated and free of hallucinations

 

Model Results

Coverage

Model Expert 1 / 100 Expert 2 / 100 Average / 100 Rank
Claude Opus 4.7 (Adaptive) 79.0 87.5 83.25 1
ChatGPT 5.4 Pro 76.0 84.5 80.25 2
Grok 4.20 70.5 76.5 73.50 3
Kimi K2.5 Thinking 47.5 71.0 59.25 4
DeepSeek 58.5 53.5 56.00 5
Gemini 3.0 Pro 34.5 57.5 46.00 6

 

Document Review

Rank Model Coding / 80 Priv. Log / 20 Total / 100
1 Claude Opus 4.6 (extended) 77.8 19.0 96.8
2 Claude Sonnet 4.6 75.5 18.5 94.0
3 Claude Opus 4.6 75.4 18.5 93.9
4 ChatGPT 5.4 Pro (extended) 76.0 17.5 93.5
5 ChatGPT 5.4 Pro (standard) 74.7 18.5 93.2
6 DeepSeek 75.9 16.5 92.4
7 ChatGPT 5.4 Thinking (extended) 75.3 16.5 91.8
8 Gemini 3.0 Pro 73.7 17.0 90.7
9 Kimi K2.5 74.4 15.0 89.4
10 Grok 4.20 52.6 11.0 63.6

 

Employment

Rank Model S1/10 S2/20 S3/15 S4/10 S5/15 S6/15 S7/15 Total
1 Kimi K2.6 Thinking 5.0 15.5 5.5 5.0 8.0 9.0 8.5 56.5
2 Claude Opus 4.7 (Adaptive) 8.0 7.5 6.0 8.0 13.0 9.0 10.0 61.5
3 DeepSeek 7.0 9.0 9.5 8.0 11.5 8.0 5.0 58.0
4 GPT 5.5 Thinking 7.0 11.0 13.5 5.0 7.5 9.0 8.5 61.5
5 Gemini 3.1 Pro 6.0 1.5 3.0 2.5 10.0 9.0 9.0 41.0
6 Grok 4.3 (beta) 5.0 7.5 8.0 2.5 10.5 8.0 10.0 51.5

 

Contract Review

Model Expert 1 / 100 Expert 2 / 100 Average / 100 Rank
Claude Opus 4.7 (Adaptive) 88.00 90.25 89.13 1
GPT 5.5 Thinking 88.00 87.75 87.88 2
DeepSeek 86.50 75.25 80.88 3
Grok 4.3 (beta) 77.50 81.00 79.25 4
Kimi K2.6 Thinking 86.00 69.00 77.50 5
Gemini 3.1 Pro 80.50 68.00 74.25 6

 

The completed tasks support several preliminary observations.

    • Models cluster tightly on routine work and deviate on analytical work. On document review, nine of ten model variants scored between 89 and 97 out of 100. On the coverage memo, the same model families spread across a 37-point range. The harder the task, the more the rubric distinguishes between systems that genuinely apply controlling authority and systems that merely sound legally fluent.

    • Reasoning-mode models lead the analytical tasks. Across the coverage, employment, and contract benchmarks, tasks that require sustained multi-step legal reasoning, the top scorers were consistently extended-thinking or reasoning-mode variants. Claude Opus 4.7 (Adaptive) and GPT 5.5 Thinking were the top two on contract review and tied for the top score on employment, and Claude Opus 4.7 (Adaptive) also led the coverage benchmark.

    • Extended thinking helps where the analysis is hard, not where the work is large. The extended-thinking variant of Claude Opus 4.6 was the top scorer on document review, but standard variants of the same family were within roughly three points. On the harder analytical tasks, the gap opens up and reasoning-mode variants pulled away from their standard counterparts.

    • Format and writing are not free points. Across both expert passes on the coverage benchmark, the lowest-scoring model lost roughly half of its points in the legal-accuracy and writing sections combined — driven by hallucinated authority, irrelevant arguments, and a deliverable that did not satisfy the memo specification. On the contract task, format alone is worth twenty points, and the lowest format scores came from models that produced substantive analysis but failed to generate a usable redlined .docx or properly attached inline comments.

    • Two-grader passes are worth the cost. On the coverage benchmark, the average difference between Expert 1 and Expert 2 scores was about 11 points; on contract review, the average gap was about 6 points but several individual outputs differed by more than 15 points. Both pairs of passes nonetheless agreed on the top of the field, suggesting the rubrics are doing real work — but it also suggests that a single-grader pass would have produced a less reliable ordering in the middle of the field.

    • Noise resistance is a real differentiator. Each task includes deliberate noise, facts and documents that look topically relevant but are not legally responsive. On the employment memo, where the fact pattern includes both hallucination traps and a Dark v. Curry County noise test, the spread between the highest and lowest scorers on the Quality Control section was meaningful, even though the rubric weighting on that section is only fifteen points.

How These Results Should Be Used

This benchmark is not a vendor scorecard. It does not declare a winner. In the real world, models change, prompts vary, and the same model can perform differently on fact patterns. The point of the work is methodological: to show what a defensible legal AI evaluation looks like, to provide a repeatable framework for measuring real legal performance, and to give legal teams a basis for asking sharper questions when they evaluate a tool.

Ultimately, the most critical takeaway from this benchmark isn’t the final score, but the fact that veteran attorneys with decades of domain expertise still varied when grading complex legal problems. This proves that true legal work is deeply nuanced, subjective, and there is still much room for improvement with model outputs. 

The full data set may be found on Hugging Face.

Additional Suggested Reading: What Does it Really Cost to Use AI for Legal Work?

Percipient Logo

Ready to Streamline Your Legal Operations?

Let’s Talk. Learn  how we can help shift your team’s focus from busywork to strategic work.

Related Posts

Percipient helps legal teams efficiently and accurately handle legal matters with human in the loop technology.

Services
  • Contract Review
  • Managed Review
  • Subpoena Compliance
  • EDiscovery & Digital Forensics
Resources
  • Articles
  • Technically Legal Podcast
Company
  • About us
  • Privacy Policy

© Percipient LLC. All Rights Reserved.