Legal services
Defensibility was the requirement nobody had written down
A litigation team wanted a large language model to cut a disclosure review. The binding requirement was not accuracy but the ability to explain and reproduce the cut to a court. That ruled the model out of the decision and into a narrow assisting role.
A non-functional requirement decided which technology was allowed.
Problem
A dispute in the Business and Property Courts produced a document population far beyond manual review. Junior fee-earners were being assigned to first-pass relevance. A partner asked whether an LLM could do the first pass instead.
Context
A mid-sized firm with an established e-discovery platform and no prior use of generative models in a live matter.
Current architecture
A hosted e-discovery platform with technology-assisted review already licensed and unused, plus manual review workflows.
Constraints
- Procedural: disclosure must be reasonable and proportionate, and the approach is disclosable to the other side and the court.
- Reproducibility: the method used to exclude documents must be explainable and repeatable months later.
- Confidentiality: documents cannot leave the jurisdiction or the firm’s tenancy.
- Time: the disclosure deadline is fixed by order.
Evidence
Each statement placed on the ladder before it was used.
Practice Direction 57AD governs disclosure in the Business and Property Courts and has been in force since 1 October 2022. It requires legal representatives to cooperate to promote reliable, efficient and cost-effective disclosure, including through the use of technology.
The Practice Direction defines technology-assisted review as all forms of document review undertaken or assisted by technology, including but not limited to predictive coding and computer-assisted review. Treating it as the established route for large populations is practitioner commentary, not the wording of the Practice Direction itself.
A method that cannot be re-run to the same result is hard to defend if the cut is challenged.
An LLM would be cheaper per document than trained reviewers. Never costed against re-review and sampling overhead.
Generative output is non-deterministic by default; that is a feature in drafting and a liability in a disclosure exercise.
Questions that changed the answer
- If the cut is challenged, can we describe the method and reproduce it exactly?
- What sampling regime would give a defensible recall estimate, and who signs it?
- Which decisions in this workflow are judgement, and which are classification?
- Does an LLM in this matter engage any obligation we would have to disclose to the other side?
Options
Including the one nobody wanted to discuss.
LLM first-pass relevance
Prompt a model to mark relevance across the population.
What it costs: Non-deterministic, hard to reproduce, and puts the exclusion decision on a method the firm cannot yet explain to a court.
Established TAR for the cut, LLM for triage only
Use the licensed predictive-coding workflow for the relevance cut. Use a model only to flag likely privilege and likely key documents for human attention, never to exclude.
What it costs: Less headline saving, and two tools to run instead of one.
Manual review with more reviewers
Staff the review conventionally.
What it costs: Defensible and slow, and arguably not proportionate at this population size.
Economics
Four horizons, not one estimate.
- Build
- Near zero: the TAR capability was already licensed and never switched on. The real cost was the protocol and the sampling design.
- Run
- Per-document model cost is small; the significant run cost is the human sampling that makes any automated cut defensible.
- Change
- Each new matter needs a fresh protocol. That is a fixed cost of doing this properly, on either route.
- Exit
- Low. The work product is a documented protocol and a reviewed set, both portable.
Decision
Use the existing technology-assisted review for the relevance cut. Restrict the language model to surfacing likely privileged and likely key documents for human review. No document is excluded by a model.
Why
The requirement that decided this was never on the original list: whatever is used must be explainable and reproducible to a court. TAR has an established, describable method. An exclusion produced by a non-deterministic model did not meet the bar, but nothing stops a model pointing a human at something.
Why not the alternatives
- LLM first-pass relevance: Not reproducible to the standard a challenged cut requires, and the saving was never costed against the sampling it would demand.
- Manual review with more reviewers: Defensible but disproportionate at this population, which the Practice Direction speaks to directly.
Trade-offs accepted
- A smaller cost saving than the partner hoped for, in exchange for a method that survives challenge.
- Two tools and two workflows for the review team to learn in one matter.
Reversibility
The asymmetry that drove the decision. A disclosure cut, once made and served, cannot be quietly redone, and an unexplainable method is discovered at the worst possible moment.
Non-functional requirements
- The method must be describable in a disclosure review document
- The same inputs must produce the same cut on re-run
- Documents remain within the firm’s tenancy and jurisdiction
- Every model-surfaced privilege flag carries a named human sign-off
Decision gate
GO on the review as designed: technology-assisted review makes the cut, and the model only surfaces privilege and key documents for a human. What stays PAUSED is the model making the cut itself. That needs a reproducible configuration and a sampling regime counsel will sign.
Implementation
TAR configured and validated with a control set. Model-assisted privilege triage run separately, with every flag reviewed by a human and logged. A written protocol produced before the review started, not after.
What the decision was expected to achieve
A proportionate review completed to the order, with a method the firm can describe. Recall against the control set is the measure, not documents per hour.
No outcome is claimed. This is an illustrative example, so there is nothing measured to report, and a real engagement would state what happened and how it was verified.
Lessons
- Ask what happens when the method is challenged. That question reorders the options faster than any accuracy benchmark.
- Non-determinism is a feature when drafting and a defect when excluding.
- A licence the firm already owned beat the technology everyone wanted to talk about.

