Industry newsAi evidence rules
A Federal Judge's 10-Point Rulebook Puts the Proof Burden on Your AI Review Tool, Not Just Vendor Marketing
On July 2, 2026, ACEDS reported on a special edition of the Journal of Technology Law & Policy, tied to the 13th Annual UF Law E-Discovery Conference, in which Chief Magistrate Judge William Matthewman proposed ten rule changes for handling AI evidence in court, including mandatory pretrial notice whenever AI is used and a rewrite of Rule 901's authentication standard. The same edition carries a practitioner benchmark from Robert Keeling, Ray Mangum, Amy Hanke and Alyssa Ogden, who tested a generative-AI review tool against human reviewers on a 1,600-document matter and reported 83.9% recall and 84.7% precision. A companion paper by Aron Ahmadia, Nathan Reff, Matthew Reiber, Chad Roberts and Cristin Traylor sets out a statistical playbook for defending AI-assisted review in court, Judge Xavier Rodriguez examined who controls data on an employee's personal phone, and a fourth paper by Ralph Artigliere, David Horrigan and Rose Hunter Jones argues AI amplifies both skill and error. The next conference runs February 10-11, 2027.
What are the 10 rules a federal judge is actually proposing?
Mandatory pretrial notice whenever AI touches evidence, plus a rewritten Rule 901 authentication standard built for AI-generated and AI-altered material.
Matthewman's article opens with cases he's already seeing cross his docket: deepfaked evidence, AI-generated victim impact statements, and fabricated expert citations that slipped past opposing counsel entirely. That's the backdrop for the ten proposals. None of this is adopted procedure — it's one judge's blueprint published in a journal — but it signals the direction courts are leaning, and litigation platforms that can't produce a clean AI-use record when asked will be the ones caught flat-footed.
Does the 1,600-document benchmark actually prove GenAI review works?
Not by itself. The 83.9% recall and 84.7% precision figures come from one matter, one unnamed tool, and no disclosed funding relationship.
ACEDS's summary doesn't say which tool was tested or who paid for the study, and a single 1,600-document matter is a data point, not a track record — precision and recall move with document mix, issue complexity, and how "responsive" was defined by the reviewers who built the gold standard. That's precisely why the companion paper exists: a statistical playbook for defending AI-assisted review implies courts will want to see the methodology, not just the final percentages, before they credit them.
What does the phone-data ruling debate mean for collection scope?
Rodriguez found both existing legal tests for who controls BYOD data broken, so discovery plans built on either one are on shaky ground.
Most litigation collections now include an employee's personal phone — chats, call logs, recordings. If the legal tests courts use to decide who "controls" that data are unreliable, the collection and chain-of-custody documentation around personal-device evidence becomes the thing opposing counsel challenges first, not the review itself.
What should buyers ask their own e-discovery vendor now?
Ask for the sampling methodology behind any accuracy number, error rates broken out by document type, and a per-document audit trail.
If mandatory AI-use notice becomes standard practice, an aggregate recall/precision score won't satisfy a judge or opposing counsel — they'll want to know which specific coding calls were AI-made, which were attorney-reviewed, and whether each answer traces back to a cited exhibit. Vendors who can only offer a headline percentage, with no visibility into how it was measured, are offering exactly the kind of unverified claim Matthewman's proposed rules are aimed at.
Frequently asked questions
Are Matthewman's 10 rules already in effect?
No. They're proposed in a journal article tied to a law school conference, not adopted by any federal rulemaking committee — a forecast, not current procedure.
Which GenAI tool scored 83.9% recall and 84.7% precision?
ACEDS's summary doesn't name it, describing only "a leading GenAI review tool" — a gap worth closing before citing similar numbers back to a client or court.
Is the benchmark result representative of AI review generally?
Not necessarily. It's one matter with 1,600 documents; results vary with document type, issue tagging definitions, and how the human baseline was built.
Sources: ACEDS, reporting on the Journal of Technology Law & Policy special edition tied to the 13th Annual UF Law E-Discovery Conference.
The original report
Industry headlines from other publications. Each links to the original reporting on the publisher's own site.