Leading on quality
1.06x
Fable 5.1, against the 15-model average
Best to worst
17%
Best on citations
Opus 5
1.05x on cited answers
Quality by difficulty
Each difficulty panel is normalized to the average of the 15 models within that difficulty.
Source: Legora
On easy work the field is essentially flat, with nearly every model within 3% of the average. The separation happens on long cases, where the spread runs from 0.78x to 1.15x. Opus 5 and Fable 5.1 lead the long tier, with Muse Spark 1.3 and Gemini 3.8 Flash close behind at around 1.09x to 1.10x. The lesson is the same as in our last update: short tasks no longer separate the models, long multi-step matters do.
Quality vs median time per case
Overall quality against median agent time per case, both relative to the 15-model average (1.00x). Open-weight models are grayed out. Frontier models are served by their own labs, so their latency is a stable reference point. Open-weight models can be hosted by many providers, and the time we measured reflects the host we used rather than the model alone.
Source: Legora
Muse Spark 1.3 is the fastest model in the field at 0.36x the average median time per case while still landing above average on quality (1.02x), which makes it the speed point on the frontier. Gemini 3.8 Flash (0.75x time, 1.04x quality), Opus 5 (0.96x, 1.05x) and Fable 5.1 (1.06x, 1.06x) complete the Pareto curve: each step up in quality costs time. The open-weight models spread widely along the time axis, with the DeepSeek Flash models at 2.3x to 2.5x, which reflects the host we used as much as the model.
Verbosity: combined output per run, split by kind
Combined output per run relative to the 15 model average (1.00x).
Source: Legora
Verbosity is a recent addition to BAR. We measure it because preferences differ: a lawyer may want a detailed explanation on some tasks and a short, direct answer on others, and this shows how much each model writes rather than how well. GPT-5.6 Terra and DeepSeek V4 Pro are the most concise at 0.76x to 0.77x, while GPT-6 Astra, Opus 5 and Fable 5.1 are the most expansive at 1.20x to 1.46x, with most of the difference sitting in documents rather than chat responses. Gemini 3.8 Flash is the exception, at 1.91x on chat responses against 0.50x on documents: it tends to answer inside the chat rather than produce documents.
Quality vs cost per case
Overall quality against average cost per case at public list prices, both relative to the 15-model average (1.00x)
Source: Legora
Among frontier models, Gemini 3.8 Flash is currently the best price-to-quality point on the Pareto curve: 1.04x quality at 0.29x the field's average cost per case. Fable 5.1 and Opus 5 are the only two models above it on quality, and they remain the right choice where the last few points on hard matters matter more than cost. The open-weight models redraw the bottom of the curve: DeepSeek V4 Flash delivers 1.02x quality at 0.12x cost and GLM 5.3 Flash 0.96x at 0.07x.
Citation performance
Cited answers: share of runs with a citation. Grounding: share of documents read that were cited. Each panel normalised to the 15 model average.
Source: Legora
Opus 5 tops cited answers at 1.05x, and the top of that panel is crowded: eleven models sit at or above the average. Grounding is where the models diverge. Gemini 3.8 Flash leads at 1.18x, followed by Kimi K3 at 1.14x and Opus 5 at 1.10x.
Which model for which kind of work
Each task family (column) is normalised to the 15-model average for that family, so a column shows which models are stronger or weaker at that kind of work rather than how hard it is.
| Analyse / Answer | Extract | Compare | Draft | Review | Edit | |
|---|---|---|---|---|---|---|
| 1.02× | 1.05× | 1.12× | 0.97× | 1.04× | 1.07× | |
| 1.04× | 0.98× | 1.11× | 1.09× | 0.98× | 1.11× | |
| 1.01× | 1.03× | 1.03× | 1.02× | 1.16× | 0.92× | |
| 1.05× | 0.99× | 1.03× | 1.10× | 0.99× | 1.08× | |
| 1.03× | 0.98× | 1.04× | 1.02× | 1.03× | 1.07× | |
| 1.03× | 1.00× | 1.02× | 1.08× | 1.03× | 1.02× | |
| 1.04× | 1.02× | 0.90× | 1.04× | 1.06× | 1.01× | |
| 1.01× | 1.01× | 0.97× | 1.03× | 0.94× | 1.04× | |
| 1.01× | 0.99× | 0.97× | 1.03× | 1.00× | 1.00× | |
| 0.95× | 1.01× | 1.12× | 1.00× | 1.00× | 1.01× | |
| 0.90× | 0.99× | 1.09× | 1.08× | 1.09× | 1.10× | |
| 1.00× | 0.99× | 0.95× | 0.92× | 0.94× | 0.87× | |
| 1.00× | 1.01× | 0.90× | 0.87× | 0.98× | 0.96× | |
| 0.95× | 1.00× | 0.87× | 0.88× | 0.90× | 0.91× | |
| 0.96× | 0.95× | 0.89× | 0.88× | 0.84× | 0.83× |
Source: Legora
No model wins every column, which is the strongest argument for staying model-agnostic. Fable 5.1, GPT-6 Astra and Opus 5 lead on Compare at 1.11x to 1.12x, while Kimi K3, Opus 5 and GLM 5.3 lead on Draft. Gemini 3.8 Flash stands out on Review at 1.16x, the single largest lead anywhere in the table.

