
Quick Answer
An AI document review is reproducible for court when four records survive each run: the instructions with every revision, the model version, the exact retrieval set and a fixed random seed.
A human check of the output does not replace them. A compliance practitioner's rule fits here too: keep an audit trail of every query and response. Ask a vendor to show all four records for one finished pass, before a matter rests on the tool. Start with the instructions; they are revised before the run, and each revision counts.
Key Points
- A Rule 26(g) challenge asks how a discovery response was reached, and four records answer it: the instructions, the model version, the retrieval set and the random seed.
- Quinn Emanuel's Business Litigation Reports of March 19, 2026 says there is "no set standard" for the level of disclosure a generative AI review requires.
- ComplexDiscovery's August 2024 buyers guide put generative AI recall "usually in the 70% range " out of the box and in "the 90% range " only after further testing and iteration.
A review pass that can be repeated can be tested: two runs, laid side by side.
A firm usually comes to AI review with one question: how can the cost of document review come down? The savings can be large. In 2025 a law firm's director of electronic discovery reported increased accuracy and savings estimated at over $5 million for one matter with generative AI tools. In 2024 an employee of one AI review vendor put large-scale human review at between $1.25 and $1.65 per document. Cost is the easy half.
The harder half is whether the cheaper pass can be shown to someone who doubts it. In a February 2026 thread in Reddit's AMLCompliance community, one commenter said regulators usually focus on data residency, auditability, and the ability to trace every finding back to the original document. A commenter in the same thread advised: "Keep an audit trail of every query and response." I take that as a near neighbor of the litigator's problem, because a challenge from opposing counsel turns on the same thing: where did this result come from?
This article is about that second half. An agentic review is one in which the software plans and carries out its own steps, and a reproducible run is one a second party can repeat and get the same result. I name the four records that make the difference; I show why a run is hard to repeat once the tuning is done; and I set out the questions to put to any tool before a matter rests on it. It begins with a question from the other side of the case.
Generative AI now does routine work in defensive document review and production, and an attorney still signs each discovery response under Rule 26(g). One question sits beneath a challenge from the other side. Could this pass be run again?
The price is low. In the Winter 2025 eDiscovery Pricing Survey from ComplexDiscovery and EDRM, GenAI-assisted review was most often priced at $0.26 to $0.50 per document, and nearly one-third of respondents reported no familiarity with or use of it. I read the two figures together: a low price can draw a firm in before anyone has asked what the tool keeps.
The technology has begun to earn its place without overturning how discovery is done, and a familiar comfort lingers beside it: a person checked the output, so the review must be defensible. Put one question to that reviewer: what was the tool told to look for? A run is made of small things that pass quickly: the instructions drafted and redrafted; the model that sat beneath the product that week; the documents pulled for each question; the chance left in every answer. In a January 2026 r/ediscovery thread on review platforms, a commenter praised a platform's AI tooling for "discrete investigative purposes" but could not imagine using it for complex matters, least of all anything "high-stakes where consistency is demanded," and consistency is made of those four small things. Lose one, and nobody can test it.
What Makes an AI Document Review Reproducible for Court?
A reproducible agentic review rests on four records: the instructions as written, the model version, the exact retrieval set and a fixed random seed, each kept in an append-only audit trail.
Reproducible means that another person, given the same materials, can run the same pass and arrive at the same coding calls. An agentic review is one where the software plans and carries out its own chain of review steps, without a person starting each one. Rule 26(g) is the federal rule under which an attorney signs a discovery response, and a challenge under it asks how that response was reached. Quinn Emanuel's Business Litigation Reports of March 19, 2026 says two things that bear on the answer: there is "no set standard" for the level of disclosure required, and the prompts that drive a generative AI review "closely resemble the instructions that would be given to a human reviewer in a document review protocol."
Check for these four before a run begins:
- The instructions as written. The exact prompt text and the issue definitions (the written description of what makes a document responsive to each issue in the case), with every revision kept in order.
- The model version. The exact identifier of the model that read the documents on the day of the run, which is a different thing from the name of the product.
- The exact retrieval set. The list of documents each pass actually read, each tied to its original by a content hash, a fingerprint computed from the file that changes if the file changes.
- A fixed random seed. The starting number that pins down the otherwise random choices a model makes as it writes, so that the same run can be asked for twice.
Think of what a run leaves behind when it is over: the tags lying on every document, thick as seed on a broad field; the rationales, full and fluent, one after another; the summary counts; and under all of it, unless someone thought to keep them, nothing of the conditions that made them. At Relevant e-Discovery our own record starts from the ground: originals are immutable and content-hashed, the audit trail is append-only, and our outputs link back to the exact exhibit they came from, so an attorney can check a call against its source before filing. The second record is the one most easily lost, a point taken up in our piece on how a review vendor can swap models mid-matter.
In 2024, an active Relativity aiR user on r/ediscovery called "the rationale, insights, and highlighting" the "hidden gems" of that product, and another commenter passed on a description of the attorney's part changing "from someone actively making the decision to someone double checking accuracy." A rationale explains a call. It does not let anyone make the call again. The common assumption is that a readable reason and a human check close the matter; they show that one answer looks right, and they leave open whether the run could be repeated.
A review of 4 sources suggests a pattern: the talk is of explaining a result and of checking it, far more than of running it again. This suggests the weak point is the run, not the reviewer. In practice, that means asking for the four records before the first document is coded.
One candid aside. Nothing I have read measures how far two runs of the same instructions drift apart, so I cannot say how often a missing seed changes a coding call; only that without the seed nobody can rule it out. Which brings the harder matter into view: why a tool that reads so fluently may still be unable to repeat itself.
What do courts ask to see when an AI review is challenged?
Ask what a judge needs to repeat an AI document review and the case law gives a different answer. Courts have mostly asked how a review was built and how it was checked.
No court has yet set a rule for generative AI review. In March 2026, the litigation firm Quinn Emanuel wrote that there is "no set standard" for how much a party must disclose. The existing decisions cover technology-assisted review (TAR), the machine-learning method that preceded generative AI, and the firm says their common principle is that transparency is crucial.
The clearest list comes from Kaye v. N.Y.C. Health & Hospitals Corp., decided in the Southern District of New York in January 2020, where four disclosed items were held sufficient. Later that year, in Livingston v. City of Chicago, the city disclosed the TAR software it planned to use and how it planned to validate the review.
Set beside the four records we keep from every AI run, Kaye's list matches on three items and misses on one.
| What Kaye accepted | Generative AI counterpart | Among the four run records? |
|---|---|---|
| Collection criteria | The retrieval set: the exact documents the run read | Close. It shows what was read but not why those documents were gathered |
| Name of the software | The tool and its pinned model version | Yes |
| Review workflow | The written prompts and every revision | Yes |
| How results were validated | Recall and precision tests on the prompts | No |
| Not on Kaye's list | A fixed random seed | Yes, though no court in our sources asked for one |
The workflow match is the strongest. Quinn Emanuel explains that generative AI prompts "closely resemble the instructions that would be given to a human reviewer in a document review protocol." Prompts can be revised and tested, but "generative AI review is not an iterative process once the prompts have been run over the corpus." After the run, the saved prompt text and its revisions are the only written record of the decision rules a judge would question. Writing on JD Supra in February 2025, a document review services provider described the same order of work: "You identify your review set, develop prompting, validate prompting, and apply validated prompting to your review set."
The same provider expects the models themselves to change: "The LLM best suited for eDiscovery review today may or may not be the best 6 months from now." Kaye counted the software's name toward transparency. When the underlying model can be swapped within months, the version number does the job the product name used to do.
Validation, the fourth Kaye item, has no match among the four records. The recall numbers suggest it may matter most.
Testing produces the gap between those bars. A run's final settings show where the review landed; only validation shows the result can be trusted, and Kaye and Livingston both asked about it.
No decision in our sources mentions a random seed. Federal Rule 26(g) requires the signing attorney to certify a response after "a reasonable inquiry," which is a reasonableness test. We still keep the seed as a working tool, because it helps a team repeat a pass when a result needs confirming. Our sources show courts asking for the confirmation, not the seed itself.
Timing matters as much as content. In In re Valsartan in 2020, the District of New Jersey faulted a party that did not disclose it might use TAR when doing so was "objectively reasonable and foreseeable." Quinn Emanuel's minimum is short: "disclosing the use of AI to the opposing party." An industry analysis of the Winter 2025 eDiscovery Pricing Survey, run by ComplexDiscovery and EDRM, found that nearly one-third of respondents had no familiarity with or use of GenAI-assisted review, and it listed auditability among the reasons legal teams remain cautious. We build append-only audit trails into our own product, so we have a stake in how this standard develops.
Three of the four records explain how a review was built. In every decision we found, the court also wanted evidence that the review worked, and none of the four records supplies that unless the validation results are saved with them.
Five checks before you certify an AI review
- Write your disclosure in the terms Kaye accepted: collection criteria, the tool and model version, the workflow, and the validation method.
- Freeze the final prompt text before the full run and keep every earlier revision, dated, beside it.
- Save the validation results with the four records: the sample tested, and the recall and precision measured on each prompt revision.
- Record the model version when the run happens, and ask your vendor whether that exact version can be called again later.
- Tell opposing counsel you are using AI review once that use becomes foreseeable, without waiting for a dispute.
How we checked this
We read a March 2026 law firm analysis of the TAR disclosure cases, an industry buyers guide on review methods, an article by a document review services provider, an industry analysis of a 2025 pricing survey, and the text of Federal Rule 26. We relied on the law firm's descriptions of Kaye, Livingston and Valsartan and did not read the opinions ourselves. The recall ranges are typical figures from one buyers guide, not results from a controlled study. The services provider sells review work, and we sell review software with an append-only audit trail, so both of us have a stake. Still unknown: what courts will require specifically for generative AI. A June 2026 Northern District of California decision, Schulte v. LinkedIn Corp., deals with generative AI tools in discovery, but we have not reviewed the opinion.
- Quinn Emanuel, analysis of AI in defensive document discovery, March 19, 2026.
- ComplexDiscovery, buyers guide on manual review, TAR and AI, August 17, 2024.
- Document review services provider, article on choosing generative AI for review, JD Supra, February 28, 2025.
- Industry analysis of the Winter 2025 eDiscovery Pricing Survey, 2025 eDiscovery review update, JD Supra, June 25, 2025.
- Legal Information Institute, Federal Rule of Civil Procedure 26, retrieved October 8, 2026.
- JD Supra, document review news listing citing Schulte v. LinkedIn Corp., as of October 8, 2026.
- Relevant e-Discovery, product description of append-only audit trails and content hashing, 2026.
Why Is It So Hard for an Agentic Review Tool to Repeat Its Own Run?
A run is hard to repeat because its result comes from tuning, and because the market divides custody records from cited answers.
Recall is the share of the truly responsive documents that a review finds. A 2024 industry comparison of review methods put generative AI recall "usually in the 70% range" straight out of the box, on a level with first-generation TAR (technology-assisted review, in which software learns from documents attorneys have already coded), and in "the 90% range" only after further testing and iteration. The difference between those two figures is not in the software. It is in the rounds.
Picture the rounds as they go: the instructions drafted, run on a sample, read against the attorneys' own calls, altered; run again, read again, altered again; slow rounds, each leaving the wording a little changed, as a path across a field is changed by every crossing, until the recall climbs and the team is content and nobody can say afterwards which wording did the work. That is the first reason the records go missing. The review a firm ends with is the last of several, and the temptation is to keep only the last.
The second reason lies under the product. Writing in 2025, a Purpose Legal author told co-panelists that the generative AI review tools then on offer were "basically the same", sharing a workflow of identifying the review set, developing prompts, validating them and applying them; the colleagues pushed back, and the author granted that the tools differ in model, prompting, interface, results and price. The same author warned: "The LLM best suited for eDiscovery review today may or may not be the best 6 months from now." So the parts that differ, and the part that changes, are the very parts a second run needs held still.
The third reason is how the market is built. Our analysis of it is plain: enterprise platforms give Bates numbering, privilege handling and chain of custody; the newer AI entrants give cited question-and-answer and chronologies; and we built Relevant e-Discovery as the combination, not one or the other. Agentic tools add a further layer. In a 2025 ACEDS article, Exterro's Kousik Chandrasekaran described agentic AI as something that "moves beyond static prompts to orchestrate tasks across identification, review, and production, while maintaining auditability and human oversight." That is a fair statement of the aim. It does not say which records leave the tool.
More steps, more places where a record must be taken. The implication is that an agentic run needs the four records at every step it orchestrates, not once for the whole matter. A cited answer shows where a fact came from, which is why I treat source-linked review as the defensibility bar; it still does not show the conditions under which the answer was produced.
Set two accounts of the same review side by side.
- Before: The AI coded the set, and our attorneys checked a sample of its calls.
- After: The final pass ran the last approved wording of the instructions, on the named model version, over the hashed document list attached, with the seed on record; every earlier wording, and the test result that retired it, is attached as well.
The first account asks to be believed. The second can be tested. And the test a small firm can run for itself, before a matter is committed to any tool, is where this goes next.
How Should a Small Firm Test an AI Review Tool's Audit Trail Before Committing a Matter?
Bring a messy collection and a hard question to the demo, run one pass, then ask the tool to show each of the four records for that run.
The records are hard to rebuild after the fact, and some cannot be rebuilt at all. A model that has since been replaced, a wording that was overwritten, a document list nobody saved: none of these comes back when opposing counsel asks. So the time to secure them is before the review starts, and the place to do it is the demo.
The demo we offer at Relevant e-Discovery stands on that footing: bring a messy collection and a hard question, and watch the tool read, code and cite your own evidence instead of reading a case study about someone else's. Bring the collection as it truly lies: the long email threads, the scanned pages gone grey, the chat exports, the duplicates heaped deep on one another; bring the question that has been troubling the case; and let the work be done in the open, slowly enough to follow. Then, of that one pass, and of any tool you are weighing, ours included, ask four questions.
- The instructions. Show me the exact wording that just ran, and the wording it replaced. Can both be exported, with the date of each?
- The model. Which model version answered, by its identifier and not the product's name? Will I be told if it changes while my matter is open?
- The retrieval set. Which documents did this pass read, and which did it leave out? Can the list be exported with a content hash for each file?
- The seed. Was the random seed fixed and written down? If the same pass is run again now, do the same calls come back?
The fourth question is the one to press. Run the pass a second time while you are still in the room. A tool that can repeat itself will show it. A tool that cannot will show that too.
Listen, as well, for the kind of answer you are given.
- A weak answer: It is all kept in the audit log, and support can pull it if you ever need it.
- A strong answer: Here is the export for the pass you just watched; the four items are in it, and the file is yours to keep.
Buyers far from litigation are asking for the same thing. In a February 2026 thread on Reddit's AMLCompliance forum, an analyst at a mid-size bank, whose team runs 15-20 active investigations monthly with 50-200 documents in each, listed a "complete audit trail of what was searched and found" among the requirements, and set aside enterprise platforms as too expensive and "overkill for our investigation volume." One reply put the rule shortly: "If an AI can't show where it found an answer, I wouldn't trust it for compliance work." Small volume does not lower the standard. It only lowers the budget for meeting it.
Two things about our own platform bear on the test. First, our processing runs single-tenant, or inside your own AWS account under your own keys, with no vendor retention and no model training; the nuance is that where the vendor retains nothing, the firm itself is the keeper of the run records, so ask where each record is stored and who can still read it after the matter closes. Second, we built for a solo or small firm on a single messy matter, the segment that enterprise platforms cannot serve without an administrator; for such a firm the records must come out of the tool as a matter of course, not through a console that someone has to be hired to run. Cost belongs in the same conversation, and our comparison of per-matter and annual platform pricing for a small firm takes up that side of the choice.
Write the four answers down, with the date and the name of whoever gave them. And if the second run comes back different from the first, note which documents moved.
What Should a Litigation Team Keep From Every AI Review Run?
Keep four records from every run: the instructions and each revision, the model version, the exact documents retrieved and the random seed, saved as the pass runs.
The courts have the subject before them now. A WilmerHale note of July 21, 2026 describes Schulte v. LinkedIn Corp., a Northern District of California decision of June 30, 2026, as addressing parties' use of generative AI tools in discovery. I cite it for its subject and its date, and I make no claim about its holding. The date is enough. I expect the argument to move, over the next year or two, from whether a tool may be used to what record of each run a party can put on the table.
A human check tests the output. It does not reach back to how the set was gathered: in 2024 a commenter on Reddit's e-discovery forum noted that each of the major review platforms was adding self-collection tools for O365, with automated workflows to process the data, so keep the collection and processing log beside the run. Left alone, a run leaves little behind it: a tag on a document; a line of reasoning; a count at the foot of a report. The slow work of saying how it all came about then falls to the attorney who signs. So start small. Pick one finished pass, ask the vendor for all four records, and note in the file which of them came back.
Summarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently Asked Questions
What Else Do Attorneys Ask About Reproducing an AI Review?
Most questions return to one point: keep the instructions, model version, retrieval set and seed for each run, and disclose AI use early.
How do I make sure my discovery process withstands a challenge from opposing counsel?
Prepare for the challenge before the run, because that is when the record is cheap to keep. Disclose that AI is in use, test and revise the instructions before they touch the review set, and save the four records for every pass. Then an objection can be met with a document. Memory will not serve.
Do I have to tell the other side that I used AI for document review?
The rules are unsettled on how much detail is owed, but disclosure of the use itself is the floor. In 2020 a federal court in New Jersey, in In re Valsartan, faulted a party for failing to disclose that it might use TAR, the older technology-assisted review in which software learns relevance from human coding. I would raise the subject with opposing counsel early, and in writing.
What is a random seed, and why would a court care about it?
A random seed is the number that fixes the choices a model would otherwise make by chance. With the seed fixed and the other three records unchanged, a second pass has the best chance of matching the first. No source I rely on measures how far two unseeded runs drift apart, so I give no figure. The point holds without one: a result that moves cannot be tested.
Is a human check of the AI's tags enough to make the review defensible?
A human check is necessary, and it is not sufficient. A 2025 ACEDS article by a self-described skeptic with three decades in the legal field found some applications "beginning to show promise and move the needle," while nothing had "completely upended e-discovery processes." Read plainly, people still carry the process. Their check confirms the tags; the run records show how the tags were made.
Can I re-run an AI review after the vendor changes the model?
You can run it, but you will not be repeating the same pass. A different model version is a different reader, and its tags may differ on the same documents. Ask the vendor whether an earlier version can be pinned for the life of a matter, and record the version on every run either way.
What would a reproducible-run report contain?
It would cover one pass and hold four things: the instructions with each revision and its date; the model name and version; an identifier and content hash (a fingerprint computed from a file's contents) for every document retrieved; and the seed with the other run settings. I have no example from a real matter to print here, so I describe the contents and give no figures.
What does AI-assisted review cost compared with manual review?
At Relevant e-Discovery, our AI-assisted review runs cents per document, against dollars per document and roughly $19K/GB for manual review. You can see the difference on your own matter. Price and record are separate questions, so ask both.
How do I see the four records on my own documents?
Bring a messy collection and a hard question to a demo, and ask to be shown each record for the pass you watch. Ask it of any tool, ours included. To book a demo with us, use the contact page at relevantediscovery.com/contact-us.
Written by
Michael
Kansky
Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.
Connect on LinkedIn

