Choose one business decision to support
A pilot should answer a practical question, such as whether AI assistance makes first-pass supplier review more useful for the legal team. It should not begin with a promise to automate all contracts. Choose a document type, risk playbook, business unit, and reviewer group.
For a UK legal department, record the governing-law assumptions and escalation route to the relevant specialist. A UK office address does not establish that every agreement uses English law. Put the approved data boundary and the conditions for using real matter material in place before the first upload.
Create a reference set with known review decisions
Ask experienced reviewers to assess a small, representative set before looking at the tool's output. Include an ordinary agreement, a heavily negotiated version, missing schedules, an amendment, and a poor scan. Preserve disagreements between human reviewers as issues to resolve, rather than pretending the reference set is infallible.
Record each expected issue, its source clause, the playbook position, and the reason it matters. Use anonymised or synthetic documents only if they still reflect the structure and difficulty of the real task. Easy examples can make almost any pilot look successful.
Use a scorecard that exposes failure modes
| Measure | What to record |
|---|---|
| Missed issues | Reference issues absent from the output, with severity |
| Unsupported findings | Findings not supported by the cited wording |
| Source accuracy | Whether the citation points to the right clause and version |
| Reviewer effort | Preparation, checking, correction, and export time |
| Handover quality | Whether the final result supports the intended business decision |
Do not collapse the results into one accuracy number without explaining its denominator. Missing a critical limitation issue and misstating a heading are different outcomes.
Test the workflow beyond the first answer
Have a reviewer open source references, correct a finding, compare related documents, and export the result. A compelling screen demonstration may still leave the team reformatting every report. Include those steps when measuring effort.
For example, a tool may identify a liability clause correctly while its proposed wording conflicts with an indemnity elsewhere. Record the dependency and whether the reviewer could resolve it efficiently. Keep the pilot instructions stable long enough to compare results. When the playbook changes, identify which earlier results need to be rerun or reviewed.
Make a bounded rollout decision
Agree acceptance criteria before interpreting the scores. A sensible outcome may be approval for a limited first-pass review with mandatory checks, a revised pilot, or rejection for that use case. Record the reasons and the person responsible for the next step.
Keep client and business expectations aligned with what the pilot established. Report observed results from the actual test rather than turning them into a universal speed claim. After rollout, sample work and revisit the decision when the document mix, instructions, model behaviour, or contractual arrangements materially change.
Worked example and decision record
Illustrative scorecard, not a Judicio performance result: suppose a lawyer labels 12 contract questions before a pilot. Nine require a substantive finding and three should return no relevant provision. The tool flags ten, including eight of the nine required findings. It therefore has eight correct flags, two extra flags and one missed finding.
| Measure | Calculation in this invented example | Meaning |
|---|---|---|
| Precision | 8 / 10 = 80% | Of the flags returned, how many were correct? |
| Recall | 8 / 9 ≈ 88.9% | Of the required findings, how many were found? |
| False negatives | 1 / 9 required findings missed | Inspect the missed issue and its consequence |
| Completion | 12 questions attempted; record failures separately | Do not remove failed documents from the denominator |
The numbers describe a tiny teaching example. They do not establish accuracy for other documents, languages or vendors. If the missed issue is an express consent requirement, a seemingly attractive overall score may still fail the team’s release criteria. Write severity rules before seeing the results so they cannot be adjusted to make a preferred tool pass.
Run a reviewable workflow
Build a reference pack using approved, non-sensitive sample documents or material your team is authorised to process. Include a clean agreement, an amendment that changes the main text, a poor scan and a document with no relevant clause. Have a lawyer label expected findings without reading the AI answer first; resolve disputed labels before scoring.
Run the same question set on the frozen pack. Record the product version or run date, selected mode, prompt, document manifest and unsuccessful attempts. Measure preparation, processing wait, source verification and correction separately. A faster first answer can still create more total work if the reviewer must reconstruct missing context.
Repeat difficult cases after changing a template, retaining the earlier run as a comparison. The SRA’s responsible-use guidance is the professional context for England and Wales; the proposed scoring method and acceptance thresholds here are editorial recommendations. Report raw counts and exceptions alongside percentages, and keep confidential evidence inside the approved workspace.
The measurement download records actual observations in blank fields. Do not substitute the invented numbers above for a real trial or present a development diagnostic as an independent customer benchmark.
Checklist and acceptance criteria
Use this checklist at handover. Record the reviewer, date, source version and unresolved items beside each answer; a tick without evidence does not close the issue.
- Freeze the file manifest, questions and reference labels before the run.
- Count correct flags, extra flags, misses and processing failures.
- Review each critical miss, irrespective of the aggregate score.
- Record total reviewer effort and cost, including corrections.
- Approve only the tested workflow and document types.
Download the editable score a uk in-house ai contract review pilot checklist (Markdown). It includes blank fields for your matter record and can be opened in a text editor or copied into your team’s document system.
A useful pilot may conclude “suitable for first-pass extraction with specified review,” “restricted to clean English-language contracts,” or “not accepted.” Record the chosen boundary and what new evidence would justify expanding it. Passing a small pilot does not justify removing supervision.
Sources and next steps
This is an editorial workflow guide for legal professionals. The suggested checks are our practical recommendations, not a statement that a regulator requires a particular software workflow.
Explore Document Review and Review Matrix, or review Judicio's regional coverage and limitations. Check the underlying source and your organisation's approved process before relying on an output.