LEGORA RESEARCH

Introducing the Legora BAR

Category

Legora Research

published

July, 2026

September 9, 2025

author

Jacob Lauritzen
Emil Sjölander
Ebba Helfer
Firas Danil

Our mission at Legora has always been to put the most advanced AI in the hands of lawyers, in the most intuitive ways possible. Since our founding in 2023, we've been advancing the state of the art of what's achievable with frontier models in the legal context.

Delivering on that mission means we have to know, rigorously and continuously, how well our platform actually performs on real legal work. We've run internal evaluations since our earliest days, but we believe that visibility belongs with our customers too, so they can contextualize performance in a fast-evolving model landscape.

Today, we are introducing the Legora Benchmark for Agentic Reasoning (BAR), a benchmark for complex legal work in the agentic era. Legora BAR evaluates AI models across a comprehensive set of end-to-end legal tasks drawn from real-world cases, purpose-built to track how Legora performs as both the underlying models and our own harness advance over time.

Legora BAR is built in collaboration with:

Why we built this benchmark

Thousands of legal teams do their most important work on Legora every day. That trust comes with a responsibility to put the best models to work, and to focus our efforts on making the Legora product measurably better over time. To do this successfully, we need to measure ourselves in a repeatable, reliable way. The Legora BAR is how we do this.

To improve the quality of our product, we have to measure it the way our customers actually use it. Many legal benchmarks are designed as technical evaluations that run in synthetic environments with simplified harnesses built specifically for benchmarking. Those environments are functional but don't fully reflect what a legal team experiences in practice. Additionally, open-sourced cases enable labs to train their models on them, which can create artificially high scores on public benchmarks.

We're taking a different approach. The quality of a legal AI product is about more than the underlying model. It’s the combination of the model, the harness, the tools it has access to, and the system it operates within. Our benchmark evaluates models on real legal cases across practice areas, inside the Legora harness and the Legora aOS™. This is the same harness and system our customers use every day. The environment is real, so the results represent what our customers actually experience. This is the standard we are committed to measuring.

Our intention with the Legora BAR is not just to measure progress, but to accelerate it. We run the benchmark continuously: every evaluation helps our team identify where to improve next, and each improvement ships directly into the Legora platform. This creates a constant loop of evaluation, identification, and improvement that results in a measurably better product and, ultimately, more value for our customers over time.

The Benchmark Fundamentals

When legal teams evaluate AI, they care about one thing: the quality of the work it produces. That quality comes from the whole system working together, which is why we benchmark the full system. It’s the only way to measure what a legal team actually experiences.

  1. The Legora harness equips the model with the tools, skills, legal sources, and workflows needed to perform legal work.

  2. The input prompt is the task you give Legora. Goal-oriented, well-scoped, and written with the variability of a real lawyer.

  3. The case environment is the matter the agent works in. It holds the documents, playbooks, precedents, templates and other materials lawyers rely on.

  4. The model is the reasoning engine. It supplies the core AI capability, and we run the best models from the leading labs.

The Legora BAR

We maintain a growing evaluation corpus of over 5,000 cases across 28 practice areas (see below), each reviewed by a legal professional. The Legora BAR is a carefully selected subset made up of hundreds of cases drafted by our law firm partners and in-house legal engineers. It reflects the same mix of practice areas, task types, and difficulty levels as the full suite, so the benchmark gives a representative view of Legora's performance across real legal work.

Practice Areas

28

Every major area, from high-volume review to end-to-end drafting and advisory work.

Cases

5,161

Each a complete task set inside a real matter, graded case by case.

Each a complete task set inside a real matter, graded case by case.

Source Documents

11,075

The files that make up the cases, from source material to precedents and prior drafts.

M&A
850
IP & tech
700
Real estate
600
Governance
345
Litigation
286
Capital markets
275
PE & VC
200
Trusts & estates
155
Funds
145
Tax
135
Commercial
125
Data & cyber
120
Life sciences
105
Patents
100
Competition
95
ESG
90
Restructuring
85
Energy
80
Lending
80
Insurance
75
White-collar
75
Employment
75
Arbitration
70
Structured fin.
65
Immigration
60
Trade
60
Transfer pricing
55
Sanctions
55
<120
120–274
275–599
≥600 cases

Each evaluation case, or eval for short, consists of four elements that mirror how legal work is created and executed:

  1. Case description — The input prompt: an end-to-end workflow a lawyer might run in Legora. Each is classified as short, medium, long (see below).

  2. The matter file room — The folder structure, and corresponding documents, that holds all documents relevant to the case. This includes, but is not limited to, firm-specific templates, playbooks, precedents, and case-specific documents.

  3. The Legora answer — The output from Legora. This can be a single document or a coordinated set of PDFs, Word documents, spreadsheets, redlines, or an in-chat answer.

  4. Rubrics — The criteria against which we compare the Legora answer. They are expert-written, binary, individually verifiable criteria covering facts, analysis, citations, and recommendations. They represent the gold standard output.

Each case is classified based on several categories, including its difficulty level. Legora needs to perform well on both simple and more complex prompts. Therefore, we added the difficulty dimension to our benchmark, where its definition is rooted in something very familiar to legal professionals: how much time the work would take a legal expert. It’s the unit firms scope and bill by, and it translates across practice areas. Examples of corresponding prompts can be found below.

Level

Senior-associate effort

Senior-associate
effort

What it looks like

Short

A few hours, up to ~half a day

A bounded, well-specified piece of work: a discrete question, a single document checked against a few points.

Medium

One to three days

A complete work product an associate would be assigned: a memo, a redline against a playbook, a pleaded section, a disclosure schedule.

Long

A week or more, up to hundreds of hours

An end-to-end delivery that moves the whole matter forward: a full position paper, a complete pleading, a deal's document set. Sustained senior judgment across a large, messy record.

Why the cases stay private

The benchmark runs on real law firm use cases supplied by recognized partners with synthesized and/or anonymized data. As we use real know-how and IP from participants, we do not open-source them. That constraint is also a strength. When benchmark cases are made public, they end up in the training data of the next model generation. This means a high score can reflect familiarity with the test as much as capability. Our cases stay private. Every score reflects the agent seeing the work for the first time, the way a lawyer does.

For transparency, we have open-sourced one of the cases, co-created with AfterQuery. It is a fully-synthetic case that mirrors the private cases inside BAR, and shows exactly how we build and score a case. AfterQuery reviewed it to confirm our data-curation methodology meets industry standards.

Defining quality – how we score

Every task comes with a defined list of criteria: the rubrics, covering the items a great answer must get right, written by a legal professional. We score the work based on how much of that checklist it actually gets right. That percentage is the Quality Rubric Score. We run our evaluations in isolated sandbox environments within the Legora platform. This enables an environment to be functionally identical to our client's experience, one where it operates in isolation and restricted to only the accessible data. Each case is run three consecutive times to ensure statistical relevance, and the output gets scored by an LLM-as-a-judge, which then gets compared to the corresponding rubrics. This procedure is repeated for all models. A more in-depth blog post on the technical set-up of our evaluation platform will follow shortly.

01

A lawyer writes the checklist

The must-haves for a great answer: the facts, the analysis, the conclusions, and the sources.

02

The system does the task

It works through the matter and produces the deliverable, the same way a lawyer would be asked to.

03

We tick what it got right

The share of the checklist it satisfies is the score. Important items count for more than minor ones.

Agreement review
90%
Quality rubric score
Identifies every material change-of-control provision
High
Flags the termination rights correctly
High
Explains the commercial impact of the key risks
Medium
Proposes appropriate redline language
Low
Notes the secondary filing deadline
Low

This answer got the high-importance points right and missed only a minor one, so it scores well. Miss something marked high and the score falls much further than missing a low item.

This answer got the high-importance points right and missed only a minor one, so it scores well. Miss something marked high and the score falls much further than missing a low item.

Some agent benchmarks grade all-or-nothing: a deliverable either satisfies every criterion or fails outright. All-or-nothing grading collapses every result into a single pass/fail and hides where systems actually differ, on the hard work where the differences matter most.

We weight our grading instead. Weighted scoring mirrors how a partner reviews work: omitting a defined term and omitting a material change-of-control provision are both drafting errors, but no partner treats them as equally serious. The weighting procedure is an integral part of the legal engineer's work when reviewing cases. They draw knowledge from their expertise, and mark rubrics as high/medium/low importance dependent on how much risk the deliverable exposes its stakeholders to if the rubric was not met. Since high-importance misses are penalized heavily, a materially incomplete deliverable still scores like one.

Key findings

The headline comparison for the benchmark are the models themselves; each model is run through the Legora harness on the same cases and tasks, scored using our quality judge. Our evaluation consists of a number of dimensions including overall output quality, cost, latency and citation coverage. While output quality is a critical component of what we measure, a model scoring high on quality does not necessarily represent the complete picture of performance. For repetitive short tasks, or bulk operations in Tabular Review, speed and cost might be the key metrics legal teams might optimize for.

Legora is model-agnostic. We deliberately evaluate across a range of models, some live in the Legora platform and some not. We collaborate closely with the frontier labs to evaluate new models and provide feedback on how they can be optimized to perform better in the Legora harness and in the legal vertical. This allows us to constantly stay at the frontier of AI technology and deploy the best performing models to our customers in real-time. In this benchmark, we ran our evaluations across a range of models from Anthropic, OpenAI and SpaceXAI.

The graph below showcases the relative output quality scores across various models and difficulty levels, with the black horizontal line representing the average. The models perform similarly on shorter legal work and begin to show variance as the work gets longer and more complex, where sustaining judgment across a full matter is critical. On the longest running tasks, Fable 5, Grok 4.5, Opus 4.8 and Sonnet 5 all significantly outperform the average across the full suite.

Quality by difficulty

Each difficulty panel is normalized to the average of the seven shown models within that difficulty.

Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)Average of the 7 shown models per difficulty (1.00×)0.75×1.00×1.10×ShortThe Legora Benchmark for Agentic Reasoning 5.6 LunaThe Legora Benchmark for Agentic Reasoning 5.6 SolThe Legora Benchmark for Agentic Reasoning 5.6 TerraThe Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Fable 5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Sonnet 5MediumThe Legora Benchmark for Agentic Reasoning 5.6 LunaThe Legora Benchmark for Agentic Reasoning 5.6 TerraThe Legora Benchmark for Agentic Reasoning 5.6 SolThe Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Fable 5LongThe Legora Benchmark for Agentic Reasoning 5.6 TerraThe Legora Benchmark for Agentic Reasoning 5.6 LunaThe Legora Benchmark for Agentic Reasoning 5.6 SolThe Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Fable 5

Quality vs latency and cost

When we look at the trade-off between quality and speed, we see a spread across the model set with Grok 4.5 showing up with lowest time per case, at a high quality output.

Quality vs median time per case

Overall quality relative to the seven-model average, against median minutes per case.

Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)Best: Fast and highest quality0.90×1.00×1.10×Median time per case (further right = slower)The Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Fable 5The Legora Benchmark for Agentic Reasoning GPT-5.6 SolThe Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning GPT-5.6 LunaThe Legora Benchmark for Agentic Reasoning GPT-5.6 Terra

On Quality vs. Cost, models like Grok 4.5, Sonnet 5, Opus 4.8 deliver high quality and lower cost compared to the average of the set.

Quality vs cost per case

Overall quality relative to the seven-model average, against average cost per case at public list prices.

Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)Best value: Strong and cheap0.90×1.00×1.10×Average cost per case (public list prices) (further right = pricier)The Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Fable 5The Legora Benchmark for Agentic Reasoning GPT-5.6 SolThe Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning GPT-5.6 LunaThe Legora Benchmark for Agentic Reasoning GPT-5.6 Terra
Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)Best value: Strong and cheap0.90×1.00×1.10×Average cost per case (public list prices) (further right = pricier)The Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Fable 5The Legora Benchmark for Agentic Reasoning GPT-5.6 SolThe Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning GPT-5.6 LunaThe Legora Benchmark for Agentic Reasoning GPT-5.6 TerraThe Legora Benchmark for Agentic Reasoning GPT-5.5The Legora Benchmark for Agentic Reasoning Grok 4.3

In addition to model-specific evaluations, The Legora BAR is a critical tool for informing the performance improvements we make to the Legora harness. Since launching the Legora aOS, we have used the BAR results to improve our harness, delivering a 5% increase in relative output quality across all models that are live in production.

Same models, one month apart - the harness improves

QualityMay average≈ +5% average upliftJune to July, same modelsJUNE 2026JULY 2026

The line shows average benchmark quality across the same set of models, evaluated on our harness in June and again in July. Because the models themselves did not change, the uplift reflects advances in the Legora harness and platform rather than in any individual model. The chart shows direction and approximate scale, not absolute scores.

Citation performance

Leveraging AI for legal analysis requires a high degree of trust, and that trust can not be earned without transparency around information sources. Lawyers must have the ability to fact check every answer Legora produces, which is why we back all of our answers with citations.

Because the Legora BAR runs inside a real matter file room, it can also measure one of the most important indicators of output quality: the citation score. Legora requires every answer it produces to be backed by a citation, so we measure two things:

  • The Cited Answers score checks whether each claim carries a citation at all.

  • The Grounding score checks whether the source documents the AI consulted to generate an answer, are reflected in the citations. In other words, did it cite everything it relied on?

In the chart below we see that OpenAI models rank highest in Cited Answers scores. On the Grounding Scores, Fable 5 and Opus 4.8 show up particularly strong.

Claude (Anthropic)GPT (OpenAI)Grok (SpacexAI)Average across all models (1.00×)0.90×1.00×1.10×Cited answersThe Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Fable 5The Legora Benchmark for Agentic Reasoning 5.6 LunaThe Legora Benchmark for Agentic Reasoning 5.6 TerraThe Legora Benchmark for Agentic Reasoning 5.6 SolGroundingThe Legora Benchmark for Agentic Reasoning Grok 4.5The Legora Benchmark for Agentic Reasoning 5.6 SolThe Legora Benchmark for Agentic Reasoning Sonnet 5The Legora Benchmark for Agentic Reasoning 5.6 LunaThe Legora Benchmark for Agentic Reasoning 5.6 TerraThe Legora Benchmark for Agentic Reasoning Opus 4.8The Legora Benchmark for Agentic Reasoning Fable 5

The black line is the average across all models and all difficulty levels. Each bar shows how far a model sits above or below the average(colors identify the labs, see the key above).

Next steps

We are co-creating the Legora BAR with a small number of leading law firms and model labs. The benchmark corpus grows every month, with the fastest growth in the hardest cases: full end-to-end matters where agents are only now capable enough for the measurement to mean something. We will soon publish a more technical blog post on how the evaluation platform works and how we draft our cases, so the wider legal and research communities can learn from what we are building.

If you are a firm or a lab that wants to help define how agentic legal work is measured, we would like to hear from you. Reach out to your Legora representative to get involved.

We are co-creating the Legora BAR with a small number of leading law firms and model labs. The benchmark corpus grows every month, with the fastest growth in the hardest cases: full end-to-end matters where agents are only now capable enough for the measurement to mean something. We will soon publish a more technical blog post on how the evaluation platform works and how we draft our cases, so the wider legal and research communities can learn from what we are building.

If you are a firm or a lab that wants to help define how agentic legal work is measured, we would like to hear from you. Reach out to your Legora representative to get involved.

Example prompts

Transactional · SHORT · a few hours

We received the counterparty's draft NDA for Project Falcon (in the matter folder). Redline it against our NDA playbook, with particular attention to the term, the definition of Confidential Information, and the non-solicit. Mark anything that has to stay out of policy, and give me a short cover note on the points that need a decision from us.

Expected output
A clean redline applying the playbook, with the three priority points addressed and any out-of-policy positions clearly marked. A short cover note listing only the points that need a decision.

Litigation · Medium · one to three days

Draft the Statement of Claim in the ICC arbitration against Meridian Components under the 2019 Supply Agreement. Work from the matter documents: the agreement, the termination correspondence, and the delivery and payment records. Plead our primary case on wrongful termination and run the unpaid-invoices claim in the alternative, keeping the pleading to what the record supports.

Expected output
A pleaded Statement of Claim in which every allegation is grounded in the matter's evidence, with the primary and alternative cases properly structured. Claims the record cannot support are left out rather than pleaded and hedged.

Tax · LONG · weeks of senior-associate work

Using the master file template provided in the matter, draft a transfer pricing Master File for the Freshworks Group for the financial year ended 31 December 2023, based on the documents in the matter folder.

A complete Master File following the template, with the entity characterizations and TP methods grounded in the audited accounts, ledger and benchmark studies which are all part of the matter. Known inconsistencies in the record are flagged and routed for confirmation rather than silently resolved.

Full case and environment is open sourced and can be downloaded here

Product

Solutions

Certified

Company

Legal

Resources

Social

© 2026 Legora. All rights reserved.