Why we built this benchmark
Thousands of legal teams do their most important work on Legora every day. That trust comes with a responsibility to put the best models to work, and to focus our efforts on making the Legora product measurably better over time. To do this successfully, we need to measure ourselves in a repeatable, reliable way. The Legora BAR is how we do this.
To improve the quality of our product, we have to measure it the way our customers actually use it. Many legal benchmarks are designed as technical evaluations that run in synthetic environments with simplified harnesses built specifically for benchmarking. Those environments are functional but don't fully reflect what a legal team experiences in practice. Additionally, open-sourced cases enable labs to train their models on them, which can create artificially high scores on public benchmarks.
We're taking a different approach. The quality of a legal AI product is about more than the underlying model. It’s the combination of the model, the harness, the tools it has access to, and the system it operates within. Our benchmark evaluates models on real legal cases across practice areas, inside the Legora harness and the Legora aOS™. This is the same harness and system our customers use every day. The environment is real, so the results represent what our customers actually experience. This is the standard we are committed to measuring.
Our intention with the Legora BAR is not just to measure progress, but to accelerate it. We run the benchmark continuously: every evaluation helps our team identify where to improve next, and each improvement ships directly into the Legora platform. This creates a constant loop of evaluation, identification, and improvement that results in a measurably better product and, ultimately, more value for our customers over time.
The Benchmark Fundamentals
When legal teams evaluate AI, they care about one thing: the quality of the work it produces. That quality comes from the whole system working together, which is why we benchmark the full system. It’s the only way to measure what a legal team actually experiences.
The Legora harness equips the model with the tools, skills, legal sources, and workflows needed to perform legal work.
The input prompt is the task you give Legora. Goal-oriented, well-scoped, and written with the variability of a real lawyer.
The case environment is the matter the agent works in. It holds the documents, playbooks, precedents, templates and other materials lawyers rely on.
The model is the reasoning engine. It supplies the core AI capability, and we run the best models from the leading labs.
The Legora BAR
We maintain a growing evaluation corpus of over 5,000 cases across 28 practice areas (see below), each reviewed by a legal professional. The Legora BAR is a carefully selected subset made up of hundreds of cases drafted by our law firm partners and in-house legal engineers. It reflects the same mix of practice areas, task types, and difficulty levels as the full suite, so the benchmark gives a representative view of Legora's performance across real legal work.
Practice Areas
28
Every major area, from high-volume review to end-to-end drafting and advisory work.
Cases
5,161
Source Documents
11,075
The files that make up the cases, from source material to precedents and prior drafts.
Each evaluation case, or eval for short, consists of four elements that mirror how legal work is created and executed:
Case description — The input prompt: an end-to-end workflow a lawyer might run in Legora. Each is classified as short, medium, long (see below).
The matter file room — The folder structure, and corresponding documents, that holds all documents relevant to the case. This includes, but is not limited to, firm-specific templates, playbooks, precedents, and case-specific documents.
The Legora answer — The output from Legora. This can be a single document or a coordinated set of PDFs, Word documents, spreadsheets, redlines, or an in-chat answer.
Rubrics — The criteria against which we compare the Legora answer. They are expert-written, binary, individually verifiable criteria covering facts, analysis, citations, and recommendations. They represent the gold standard output.
Each case is classified based on several categories, including its difficulty level. Legora needs to perform well on both simple and more complex prompts. Therefore, we added the difficulty dimension to our benchmark, where its definition is rooted in something very familiar to legal professionals: how much time the work would take a legal expert. It’s the unit firms scope and bill by, and it translates across practice areas. Examples of corresponding prompts can be found below.
Level
What it looks like
Short
A few hours, up to ~half a day
A bounded, well-specified piece of work: a discrete question, a single document checked against a few points.
Medium
One to three days
A complete work product an associate would be assigned: a memo, a redline against a playbook, a pleaded section, a disclosure schedule.
Long
A week or more, up to hundreds of hours
An end-to-end delivery that moves the whole matter forward: a full position paper, a complete pleading, a deal's document set. Sustained senior judgment across a large, messy record.
Why the cases stay private
The benchmark runs on real law firm use cases supplied by recognized partners with synthesized and/or anonymized data. As we use real know-how and IP from participants, we do not open-source them. That constraint is also a strength. When benchmark cases are made public, they end up in the training data of the next model generation. This means a high score can reflect familiarity with the test as much as capability. Our cases stay private. Every score reflects the agent seeing the work for the first time, the way a lawyer does.
For transparency, we have open-sourced one of the cases, co-created with AfterQuery. It is a fully-synthetic case that mirrors the private cases inside BAR, and shows exactly how we build and score a case. AfterQuery reviewed it to confirm our data-curation methodology meets industry standards.
Defining quality – how we score
Every task comes with a defined list of criteria: the rubrics, covering the items a great answer must get right, written by a legal professional. We score the work based on how much of that checklist it actually gets right. That percentage is the Quality Rubric Score. We run our evaluations in isolated sandbox environments within the Legora platform. This enables an environment to be functionally identical to our client's experience, one where it operates in isolation and restricted to only the accessible data. Each case is run three consecutive times to ensure statistical relevance, and the output gets scored by an LLM-as-a-judge, which then gets compared to the corresponding rubrics. This procedure is repeated for all models. A more in-depth blog post on the technical set-up of our evaluation platform will follow shortly.
01
A lawyer writes the checklist
The must-haves for a great answer: the facts, the analysis, the conclusions, and the sources.
02
The system does the task
It works through the matter and produces the deliverable, the same way a lawyer would be asked to.
03
We tick what it got right
The share of the checklist it satisfies is the score. Important items count for more than minor ones.
Some agent benchmarks grade all-or-nothing: a deliverable either satisfies every criterion or fails outright. All-or-nothing grading collapses every result into a single pass/fail and hides where systems actually differ, on the hard work where the differences matter most.
We weight our grading instead. Weighted scoring mirrors how a partner reviews work: omitting a defined term and omitting a material change-of-control provision are both drafting errors, but no partner treats them as equally serious. The weighting procedure is an integral part of the legal engineer's work when reviewing cases. They draw knowledge from their expertise, and mark rubrics as high/medium/low importance dependent on how much risk the deliverable exposes its stakeholders to if the rubric was not met. Since high-importance misses are penalized heavily, a materially incomplete deliverable still scores like one.
Key findings
The headline comparison for the benchmark are the models themselves; each model is run through the Legora harness on the same cases and tasks, scored using our quality judge. Our evaluation consists of a number of dimensions including overall output quality, cost, latency and citation coverage. While output quality is a critical component of what we measure, a model scoring high on quality does not necessarily represent the complete picture of performance. For repetitive short tasks, or bulk operations in Tabular Review, speed and cost might be the key metrics legal teams might optimize for.
Legora is model-agnostic. We deliberately evaluate across a range of models, some live in the Legora platform and some not. We collaborate closely with the frontier labs to evaluate new models and provide feedback on how they can be optimized to perform better in the Legora harness and in the legal vertical. This allows us to constantly stay at the frontier of AI technology and deploy the best performing models to our customers in real-time. In this benchmark, we ran our evaluations across a range of models from Anthropic, OpenAI and SpaceXAI.
The graph below showcases the relative output quality scores across various models and difficulty levels, with the black horizontal line representing the average. The models perform similarly on shorter legal work and begin to show variance as the work gets longer and more complex, where sustaining judgment across a full matter is critical. On the longest running tasks, Fable 5, Grok 4.5, Opus 4.8 and Sonnet 5 all significantly outperform the average across the full suite.
Quality by difficulty
Each difficulty panel is normalized to the average of the seven shown models within that difficulty.
Quality vs latency and cost
When we look at the trade-off between quality and speed, we see a spread across the model set with Grok 4.5 showing up with lowest time per case, at a high quality output.
Quality vs median time per case
Overall quality relative to the seven-model average, against median minutes per case.
On Quality vs. Cost, models like Grok 4.5, Sonnet 5, Opus 4.8 deliver high quality and lower cost compared to the average of the set.
Quality vs cost per case
Overall quality relative to the seven-model average, against average cost per case at public list prices.
In addition to model-specific evaluations, The Legora BAR is a critical tool for informing the performance improvements we make to the Legora harness. Since launching the Legora aOS, we have used the BAR results to improve our harness, delivering a 5% increase in relative output quality across all models that are live in production.
Same models, one month apart - the harness improves
The line shows average benchmark quality across the same set of models, evaluated on our harness in June and again in July. Because the models themselves did not change, the uplift reflects advances in the Legora harness and platform rather than in any individual model. The chart shows direction and approximate scale, not absolute scores.
Citation performance
Leveraging AI for legal analysis requires a high degree of trust, and that trust can not be earned without transparency around information sources. Lawyers must have the ability to fact check every answer Legora produces, which is why we back all of our answers with citations.
Because the Legora BAR runs inside a real matter file room, it can also measure one of the most important indicators of output quality: the citation score. Legora requires every answer it produces to be backed by a citation, so we measure two things:
The Cited Answers score checks whether each claim carries a citation at all.
The Grounding score checks whether the source documents the AI consulted to generate an answer, are reflected in the citations. In other words, did it cite everything it relied on?
In the chart below we see that OpenAI models rank highest in Cited Answers scores. On the Grounding Scores, Fable 5 and Opus 4.8 show up particularly strong.
The black line is the average across all models and all difficulty levels. Each bar shows how far a model sits above or below the average(colors identify the labs, see the key above).
Next steps
Example prompts
Transactional · SHORT · a few hours
We received the counterparty's draft NDA for Project Falcon (in the matter folder). Redline it against our NDA playbook, with particular attention to the term, the definition of Confidential Information, and the non-solicit. Mark anything that has to stay out of policy, and give me a short cover note on the points that need a decision from us.
Expected output
A clean redline applying the playbook, with the three priority points addressed and any out-of-policy positions clearly marked. A short cover note listing only the points that need a decision.
Litigation · Medium · one to three days
Draft the Statement of Claim in the ICC arbitration against Meridian Components under the 2019 Supply Agreement. Work from the matter documents: the agreement, the termination correspondence, and the delivery and payment records. Plead our primary case on wrongful termination and run the unpaid-invoices claim in the alternative, keeping the pleading to what the record supports.
Expected output
A pleaded Statement of Claim in which every allegation is grounded in the matter's evidence, with the primary and alternative cases properly structured. Claims the record cannot support are left out rather than pleaded and hedged.
Tax · LONG · weeks of senior-associate work
Using the master file template provided in the matter, draft a transfer pricing Master File for the Freshworks Group for the financial year ended 31 December 2023, based on the documents in the matter folder.
A complete Master File following the template, with the entity characterizations and TP methods grounded in the audited accounts, ledger and benchmark studies which are all part of the matter. Known inconsistencies in the record are flagged and routed for confirmation rather than silently resolved.


