Testing AI's Promise For Large-Scale Document Review

| | Law360

Can generative artificial intelligence make document review more efficient and accurate without increasing risk, and if so, how?

Our recent test of a generative AI tool offers early evidence that it can satisfy legal obligations in responsiveness review — the process of identifying which documents must be produced in discovery.  Using documents from a prior regulatory matter, we compared responsiveness and issue coding decisions by attorney reviewers against predictions generated by Relativity's aiR Review. 

Excluding documents the tool labeled "borderline," the AI achieved 83.9% recall and 84.7% precision in identifying responsive documents.  In plain terms, the tool found roughly 84 of every 100 documents that the human reviewers deemed responsive, and when it flagged a document as responsive, it was right about 85% of the time. 

These findings compare favorably to traditional technology-assisted review workflows — the machine-learning approach that courts have accepted for over a decade — which often treat recall in the 70% to 75% range as an acceptable benchmark, depending on the matter.  Further, precision in technology-assisted review workflows typically drops significantly as recall increases, meaning that far more nonresponsive documents are mixed in with the responsive ones. 

These results offer a practical starting point for assessing where generative AI technology can help with responsiveness review now — and where lawyers still need guardrails.

Putting Generative AI to the Test

Our sample included 1,600 custodial documents, primarily emails and common attachments, and a smaller set of workpaper documents.  Documents that could not be processed because of limits on text size or a lack of extracted text were excluded, affecting about 6% of the population.

Before running the full sample through the tool, the team built and refined prompts based on the review protocol from the original, attorney-reviewed matter.  The initial prompts for the AI included nearly identical information provided to attorney reviewers in that protocol, with only minor, nonsubstantive revisions.  The team then iteratively tested the prompts on samples of documents, refining them after each round.

For example, in early test runs, the tool overrelied on certain industry terms and a particular keyword as signs of responsiveness, which led to false positives.  The team adjusted the prompt to make clear that those terms alone did not make a document responsive.

After completing this iterative testing, the team ran the tool across the full 1,600-document control set.

Where the AI Delivered

The strongest results came from the responsiveness review of the custodial documents, e.g., emails and common attachments, such as Microsoft Word documents.  Specifically, in the control set of 1,600 custodial documents, the tool predicted 554 documents were responsive.  The gold standard review, conducted by two attorneys who had led the original review, determined that 559 documents were actually responsive.

Of the documents that the AI predicted were responsive, 469 (84.7%) were confirmed responsive (true positives), while 85 (15.3%) were not responsive (false positives).  Of the 600 documents it predicted were not responsive, 590 (98.3%) were confirmed to be not responsive; the tool missed only 10 (1.7%) actually responsive documents. 

Excluding the documents the tool identified as borderline from the analysis, these figures work out to a recall rate of 83.9%, with 469 of 559 responsive documents found, and a precision rate of 84.7%, with 469 of 554 predictions correct.  In sum, the tool found most of the documents that mattered without flooding the review with false positives. 

When the tool's borderline responsive predictions were treated as responsive, recall rose to 96.2%, but precision dropped to 58.4%.  Such a trade-off may be acceptable in some matters where the priority is to reduce the risk of missing responsive information.  In other matters, however, it could create new burdens, such as driving up review costs by sending marginal documents to attorney reviewers or increasing business risk by producing sensitive documents that are not responsive.  These findings suggest that borderline documents should be excluded from review.

Recognizing the Limits of Generative AI

Mixed Results With Issue Coding

While the responsiveness results for the custodial documents were compelling, the issue coding results were mixed. 

The study team ran the AI tool against the same 1,600 custodial documents for the 10 issue tags most frequently applied in the underlying human review, and the results varied by issue. 

For some issues, the tool's predictions aligned well with the final review determinations.  For other issues, performance depended on issue complexity, document type and the amount of context available from text alone.

For example, one issue called for information limited to a specific year.  The AI's predictions for that issue resulted in high recall but low precision, caused in large part by the tool's inability to consider document metadata or other indicia of date.  As a result, the AI predicted that documents from other years were responsive as well. 

By contrast, another issue involved documents that were largely spreadsheets and charts, and their recall was relatively low.  In that format, relevance was more apparent in native or image view than in extracted text.

Responsiveness review often asks a direct question: Does this document fall within a defined request or subject area?  Issue coding can require more nuance, especially when the issue turns on timing, metadata, document relationships, or information outside the four corners of the extracted text.

In short, the AI performed better when the right coding decision was apparent from the plain text of the document.  It struggled when the decision depended on structure, context or visual information. 

Challenges With Workpaper Documents

The workpaper documents in our test presented a separate challenge.  These were a specialized, industry-specific document type with characteristics that proved challenging for the AI.  In addition, the workpapers were structured as a family of documents that included a cover page providing key context for the human reviewers, which the AI could not consider. 

Ultimately, the study team concluded that generative AI was not an effective tool for coding this type of document and did not run the analysis on the full 200-document sample.  The primary limitation is that current generative AI tools typically operate on the extracted text of a single document, so they cannot consider family relationships or the context provided by family documents, which often includes critical information for a responsiveness decision. 

Practitioners should account for these limitations when designing an effective discovery workflow.  The study shows that current AI document review tools work best on primarily text-based documents (other AI tools focus on image review). 

At the outset, then, practitioners should assess the dataset and exclude from the AI workstream any documents that do not meet text-extraction and other technical requirements.  They should also look for document categories that are prone to incomplete or inaccurate extracted text, such as templates, scanned documents, images, graphs or charts.  

Where document family relationships are critical, practitioners should consider whether workarounds would be feasible and economical.  And where information beyond the four corners of a document, such as metadata, is key, practitioners should consider excluding those documents from the AI process or addressing them with a tailored workflow, such as modifying the prompt or applying a metadata filter before running the tool.  

AI in Discovery Workflows

Our findings provide empirical validation for the accuracy of AI tools in responsiveness review and suggest several practical ways legal teams could incorporate generative AI into discovery workflows.

For example, the results support the use of AI to replace or supplement first-level responsiveness review in appropriate matters.  AI tools could also be used to prioritize documents for second-level human review, route borderline documents to experienced reviewers or perform quality control by identifying disagreements between human coding and AI predictions. 

AI tools may be a defensible and efficient component of the discovery workflow in appropriate cases.  They belong where they add value and their risks can be managed.  Practitioners should consider evaluating the data set, testing and refining their prompts, and identifying when AI might not be appropriate for particular subsets of data. 

The Bottom Line for Lawyers

The study suggests that AI tools can deliver review results that compare favorably with accepted benchmarks for technology-assisted review.  Lawyers still need to understand their tool's limits, validate its outputs, and make reasoned decisions about how the AI's predictions will affect production, issue analysis and quality control.

The study shows that AI analysis can be a reasonable and reliable method for document review — including as a basis for culling nonresponsive documents from the review set. Future studies should examine how its performance changes with different prompts and review protocols, and how results generalize beyond a single closed matter.

Robert Keeling is a co-managing partner and the leader of the second request practice at Redgrave LLP.

Amy Hanke is counsel at the firm.

The opinions expressed are those of the author(s) and do not necessarily reflect the views of their employer, its clients, or Portfolio Media Inc., or any of its or their respective affiliates.  This article is for general information purposes and is not intended to be and should not be taken as legal advice.