Evaluating legal AI
The question
Does the evidence support the decision?
Our ambition is a legal function that stays informed as an organisation changes and contributes while its decisions are taking shape. That sets a practical question for evaluation: does the system help people reach a sound decision, with current evidence and an acceptable amount of review?
A benchmark is a set of test tasks. Its score depends on what counts as passing. We examined public accounts, methods and code from Harvey and Legora. The findings below explain what their scores measure and what an organisation should establish before relying on them.
Before the score
The tools around the AI matter.
Harvey describes a system for checking contracts against a firm’s written review rules, often called a playbook. An AI coordinator divides the review among several AI workers. Each proposes changes separately. The system combines changes that fit together and sends conflicting edits back to the coordinator. It also retains information about the deal and earlier work for follow-up requests.
Legora describes an AI system that plans the work, uses tools to carry it out, checks the result and presents it for the user’s approval. Its wider platform, called aOS, selects models, retains information from earlier work and assigns tasks to specialist AI workers. Firm instructions and matter history help guide the work.
These are the companies’ descriptions of their products; we have not independently tested them. The underlying AI model is only one part of the service. Sources, tools, instructions and human review also affect the result. Testing a model alone answers a different question from testing the product a lawyer actually uses.
Harvey LAB
What does the AI reviewer see?
Harvey’s Legal Agent Benchmark, or LAB, uses an AI reviewer to assess the work. In the published code we inspected, that reviewer receives the task’s title, the answer and one item from the marking checklist, called a rubric. The code does not separately supply the original instructions, case documents or source references attached to that checklist item.
Requirement met or not met
The reviewer checks whether the answer meets the written requirement. A carefully prepared checklist can contain accurate facts, citations and important omissions to look for.
The quality of that checklist therefore matters. This grading step does not independently examine the full case file to check every claim. That limitation alone does not establish that a particular score is wrong.
Harvey LAB
A higher score for unchanged answers.
Harvey reports a small test in which ten AI answers were scored twice: first by one AI reviewer, then by two. The answers stayed the same, but the overall score rose from 10% to 15%.
Why? Each reviewer checks whether an answer meets every requirement in the test. The first reviewer passed one answer out of ten: 10%. The second passed that answer and one more: 20%. Averaging their scores gives 15%. Only one answer passed both reviews.
Harvey says using two reviewers reduces dependence on any one model’s judgment. It publishes the scoring rule and each reviewer’s results. The lesson is that a score can rise because the assessment changed. To know whether the work improved, we also need to know whether the same test was used.
Legora BAR
What does a high score actually tell you?
Legora reports a collection of 5,161 cases and uses a selected group of hundreds for its BAR benchmark. Its quality charts compare models with the group average. They award more weight to important checklist requirements than to minor ones.
These comparisons help show which model did better in the test. They do not directly tell us how many completed pieces of work were ready to use.
Legora also measures whether claims have citations and whether the documents the AI consulted appear among those citations. Those measures leave two further questions: does each cited document support the claim, and did the AI find all the documents it needed? Other parts of Legora’s checklist may address some of these questions.
in the scanned accounts.Pass
could not be confirmed.Pass
Either answer can pass this rule.
In this example, the checklist accepts either the specified figure—23 average full-time-equivalent employees—or a clear statement that the figure could not be confirmed from the supplied documents. The note explains that this alternative accommodates work prepared without converting the scanned page into readable text. Acknowledging uncertainty is useful. But this pass alone does not tell us whether the system successfully read the accounts.
Harvey Tenet
Better tools can raise the same model’s score.
Harvey reports that Kimi K3, before Harvey’s additional training, scored 58.8% in published APEX Agents results and 67.5% when run with Harvey’s different tools and access to files. This shows why the setup matters. To measure what additional model training achieved, results need to be compared under the same conditions.
Harvey also reports a separate assessment, called APEX, run independently by Mercor on 100 tasks that were not shared with Harvey. That is meaningful external evidence. APEX and APEX Agents are different tests, and we have not inspected the private tasks or results of individual attempts.
From measurement to use
Ask to see the work behind the result.
- 01 / The work
- What was requested? What evidence was available? What did the AI produce, miss or get wrong? What did a professional have to correct?
- 02 / The test
- Which tasks were chosen, who checked the answers and what counted as passing? Include unsuccessful attempts and explain how they affected the score.
- 03 / The benefit
- How long did it take to produce work that could be used? Include software and professional fees, together with the organisation’s time preparing information, reviewing answers and correcting mistakes.
For ongoing legal support, we would also test what happens when important evidence changes. Does the system prompt the right review? Can another adviser understand the record and continue the work? We propose applying these standards to our own service too.
A useful benchmark can justify investigating a tool. A decision to rely on it requires evidence relevant to the actual work and the people responsible for it.
Accountability
Understanding and accountability
An AI’s explanation needs examination too.
Developers know how they build, train and run their models. That knowledge does not give a complete, human-understandable explanation of the internal process behind each answer. Research has begun to explain parts of that process. Anthropic’s 2025 circuit-tracing paper is one example, but its authors identify important gaps: some computation remains unexplained, and the simplified model they study may not follow the original model’s internal process.
A model’s written reasoning is a separate question. In controlled experiments with Claude 3.7 Sonnet and DeepSeek R1, researchers found that supplied clues sometimes influenced an answer without appearing in its written reasoning. These were particular models, selected clues and multiple-choice tasks; the study does not establish how often this happens in legal work or make all reasoning-based monitoring useless.
NIST distinguishes transparency about a system’s use from explanation of its operation and interpretation of its outputs. Its framework also makes clear that transparency alone does not establish accuracy. This is voluntary guidance, not a certification of any service.
Testing improvement
OpenAI / 6 September 2026
Measure the result, including the review.
OpenAI reports increased coding, experimentation and task success inside its research organisation. Its account describes research assistance under human direction; it does not establish a corresponding gain in the quality or pace of scientific discovery.
OpenAI uses another AI to assess whether a task succeeded, checking later evidence of the outcome. It reports agreement with human judgment on 25 tasks assessed by people. The main chart excludes uncertain outcomes; OpenAI says the upward trend remains if these count as failures. Estimated human task durations are also AI assessments. These are preliminary, internal measures.
More than half of successful tasks estimated to take a person four to eight hours in the six months reported involved human intervention. The article also acknowledges unresolved problems in safely advancing fully autonomous recursive self-improvement.
Applying the standard to our own approach
Test how improvement is judged.
We want AI to help improve the methods behind legal work, including the checks used to assess it. Three primary sources inform that approach. They offer methods and reasons for caution; they do not establish our service’s effectiveness.
AI can help improve an AI workflow. The 2023 DSPy paper describes selecting and improving instructions and examples against a chosen measure of performance. Its experiments include one program helping to improve another. These are results on tasks such as mathematics and answering questions, not evidence of effective legal services. Improving a workflow’s instructions is also distinct from changing an AI model’s underlying parameters.
The checking method needs review too. NIST’s 2023 AI Risk Management Framework calls for reviewing measures and controls, evaluating how well testing works, and using feedback to improve systems. This is voluntary guidance, not proof of a system’s quality. It supports examining the way we detect errors as well as correcting the errors we find.
A familiar test can give false confidence. Research by Dwork and colleagues shows how repeatedly using the same data to choose what to test next can undermine the validity of the findings. This is a general statistical concern, not a finding about Harvey or Legora. It is a reason to reserve independent examples for checking proposed changes, and to keep examining whether our tests represent the work that matters.
Our approach is to compare old and new methods, retain the reasons and results, and look for useful gains in accuracy, timeliness and the organisation’s effort. If the test changes, that change must be visible too. These are design choices informed by the sources above, not performance we have demonstrated. How we put this into practice.
Connected knowledge
Organising knowledge as it grows
Keep track of what depends on what.
Anthropic reports that initial attempts to formalise Fermat’s Last Theorem lost track of shared work. It attributes its successful effort to Prove2Me, which separated theorem statements from proofs, tracked dependencies and made descriptions searchable. We have read the account, not reproduced the proof.
The Prove2Me paper distinguishes formal validity from whether a statement captures its intended meaning. People audit the goals, definitions and milestones; the system permits other intermediate steps without that human review of meaning because Lean checks proofs of the audited statements. The proof repository also distinguishes exact formal statements from automatically generated English summaries that need interpretation.
Sources, method and limits
Checked on 6 September 2026. Every empirical finding above links to the inspected primary documentation or code. Code references are pinned to a particular revision; web publications and pull-request discussions may change.
We read public documents and inspected published code without running it. We did not repeat the benchmarks, test the paid products or inspect private evaluation records. The public code does not show how every private test was conducted. This research does not establish that Rule & Reason provides a better service.
This is a targeted examination of measurement and explanation, not a complete product assessment. It is published by Rule & Reason with a commercial interest in how legal work is delivered. The same evidential standard should be applied to our claims.
Inspected source revisions
- Harvey LAB
- a2b429eb6c9683c4fdeced3bc6b3af36edf239a6
- Legora public case
- f5032cb7ce0c713379057b5fb9f5b944b6163b04