
Quick Answer
Unless your AI review contract pins a model version per matter and requires advance notice of updates, the vendor may swap models mid-matter, and March's coding may not reproduce in June.
Ask for three things in writing: a pinned model version, advance notice with the right to pause and revalidate, and preservation of the version, prompts, outputs and human decisions behind each call. The stakes are rising. In a November 2025 practitioner thread on AI review, one commenter predicted that where the technology is used it "will almost certainly reduce the total number of documents requiring human review, possibly substantially." Fewer human eyes means the model carries more of the defense.
There is a trade-off no clause removes. A pinned model is a model frozen out of its own improvements, and a firm that insists on March's reviewer in June also forgoes whatever June's reviewer might have caught; I would accept that bargain on most matters, though never silently. Which is why validation, and not the pin alone, has to carry the weight when opposing counsel starts asking.
Key Points
- Quinn Emanuel's March 2026 report says generative AI review "is not an iterative process once the prompts have been run over the corpus," so its validation measures one model.
- A 2025 decision in the New Jersey insulin pricing litigation linked "some level of transparency and validation of TAR methodologies" to Rule 26(g) of the Federal Rules of Civil Procedure.
- A November 2025 practitioner thread reported that Relativity's aiR was scheduled to join Relativity's hosting environment beginning in April 2026 , making it effectively a "free" part of the platform.
The model-update clause sits in the vendor agreement, not in the review screen.
Let AI cut the number of documents that need human eyes, but buy the cut on terms that keep the same model coding the matter from first production to last.
The savings are real. Practitioners now describe the gain less as faster review than as understanding the case sooner, with timelines, key players and document categories surfaced at the outset rather than discovered in the eighth week of a linear review.
The arithmetic is not new, though the machinery is. After the ILTA conference in August 2024, a practitioner on Reddit reported that in some sessions "lots of folks raised their hands" when asked whether they were using a generative AI product to auto code documents in review. Another member of the same thread objected that generative AI "literally generates new things, like summaries of documents," which made it a different instrument from the predictive coding models litigators had long since learned to defend. The distinction matters. A predictive model learns from your reviewers, whereas a generative model arrives already taught, by a vendor, and can be retaught by that vendor between one production and the next.
What I want a small firm to notice is that the cheapest review is the one you never have to run twice. Sirion, a contract-management vendor, claims in its own review guide that AI "can often be more consistent than human reviewers, especially over long periods or large document sets," yet that consistency outlasts a long matter only under version pinning, a contract term that fixes which model release codes a given matter until the parties agree otherwise. Without it, a low per-document price buys the work but not the worker. Cheap review is not cheap twice.
This year's ILTACON brought another round of AI announcements from the large review platforms, as LawNext's round-up recorded, and from the buyer's side of the table each announcement reads as a question about which model will be coding your documents next quarter. Whether you are weighing Relativity against its smaller rivals or a newcomer against both, the answer that decides the real cost sits in the agreement, in a clause most buyers have never been shown.
GenAI review has moved into daily eDiscovery work while pricing frameworks have not kept pace, and contract structure remains an unresolved question that buyers must settle before a vendor settles it.
The Winter 2026 eDiscovery Pricing Survey, run by ComplexDiscovery OÜ with the Electronic Discovery Reference Model and drawing 53 completed responses across 25 questions, found buyers weighing the promise of AI efficiency against "unresolved questions around exception handling, quality control, and contract structure." Law firms made up 43.4% of its respondents. The uncertainty, in other words, belongs to the people who sign.
Contract structure is where I would begin, because it conceals the question that matters most in litigation. Ask the vendor, in writing, which model will read these documents at the first production, and who decides whether the same model is still reading them at the last. The documents will not have changed. The reader will have.
Courts do not ask for a motionless machine. They ask for transparency, and some have gone so far as to expect the producing party to name the software and describe how its results were validated, which is a reasonable demand until one realizes that a validation performed against one model says very little about the model that replaced it, so the first document worth reading is not the vendor's brochure but the change clause folded quietly into its terms.
Top questions this article answers
- Can my AI review vendor update the model without telling me? Bundled hosting and per-document pricing both leave the switch with the vendor.
- What happens if my AI review vendor changes its model mid case? Validation stops describing the reviewer.
- What contract terms keep AI document review reproducible? Pinning, notice and records, negotiated before the first prompt.
What does your AI review agreement actually say about model updates?
Most agreements fix the price, which can run cents per document against roughly $19K/GB for manual review, yet leave the choice of model to the vendor.
I would read the modification clause first, because it is usually the shortest paragraph in the agreement and the one with the longest reach. Pull the master agreement, the order form and any AI addendum, and check four things:
- Find the language that lets the vendor modify, update or replace the service, and see whether it mentions models at all.
- Confirm which service tier you bought, since AI terms can shift from one tier to the next.
- Look for any promise to warn you before a change takes effect.
- See whether the rate is tied to a named model or only to a document count.
Tier matters more than most buyers suppose. Even vendors concede, in their own presentations, that AI terms differ from provider to provider and from one tier of service to the next, so that two firms on the same platform may hold quite different rights over what the machine does with their documents and when it may be changed beneath them.
The platforms are changing shape as well. A commenter on the r/ediscovery forum reported in November 2025 that Relativity's aiR was scheduled to be integrated into Relativity's hosting environment beginning in April 2026, making it effectively a "free" part of the platform rather than a potentially expensive add-on. A feature you did not buy separately is a feature you cannot easily refuse. Its upgrade schedule becomes the host's schedule. There is something almost feudal in the arrangement, for the tenant pays a settled rent while the lord keeps the right to rebuild the house around the tenant, who learns of the new staircase only by falling down it.
Pricing tells the same story from the other side. The Winter 2026 eDiscovery Pricing Survey by ComplexDiscovery OÜ and the Electronic Discovery Reference Model found that hybrid models and per-document billing each account for roughly 28% of reported primary GenAI review pricing models, with $0.11 to $0.50 per document emerging as the competitive zone. Conventional wisdom treats a locked rate as a locked service. The rate fixes what each document costs and says nothing about which model reads it.
In our analysis, the gap between cents and dollars per document is a delta a buyer can see on their own matter, which is exactly why the cheaper reviewer has become worth guarding. A review of 4 sources suggests the pattern is structural rather than accidental: tiered terms, bundled features and per-document prices all move the upgrade decision toward the vendor. Anyone weighing per-matter against annual platform pricing should read the model clause beside the rate card, as one reads the small print of a treaty beside its preamble, and should ask, before the first batch is coded, what becomes of a review already validated when the vendor throws the switch.
Why does a mid-review model change threaten a validated AI review?
Validation describes one model's coding, so a later swap leaves the numbers vouching for a reviewer that has gone, and an append-only audit trail is what shows where the seam fell.
Validation, in this setting, means testing a review's output against human judgment on samples and reporting how often the two agree. Quinn Emanuel's March 2026 publication on AI in defensive discovery explains that generative AI review is trained with written prompts closely resembling a human review protocol, and that it "is not an iterative process once the prompts have been run over the corpus." The prompts are tested, revised and validated with standard metrics before the full run. Then the run happens once.
Every figure that comes out of that validation is a statement about a particular model reading a particular corpus under particular instructions. Sirion's vendor guide to document review describes quality control as senior attorneys sampling batches of reviewed documents to find inconsistencies in coding, and after a swap the practical step is to draw those samples from both sides of the change date, since that seam is where any inconsistency will gather. The very title of an article in the Journal of Technology Law & Policy, on statistically defending artificial intelligence in document review, tells you where the defense of such a review lives. It lives in the numbers. So ask the vendor, before the first batch runs, whether its logs will tie every coding call to the model version that made it.
Contrary to a widely held view, the danger is not change itself. Continuous active learning refines its model in real time as reviewers code, and courts accepted it all the same; in Rio Tinto PLC v. Vale S.A., the Southern District of New York called permission to use TAR "black letter law." What courts have never been asked to bless is a change that nobody designed, measured or wrote down, which is the precise character of a vendor upgrade landing in the middle of rolling productions, like a new clerk taking over a ledger halfway down the page without initialing a single entry.
The practitioners who handle this best keep the record close to the document. One describes a workflow in which the reviewer opens and cites the underlying document before marking anything responsive or privileged, routes low-confidence or novel items to an exception queue, and requires human sign-off before production; the characteristic failure, in that account, is a fluent summary that drops its qualifiers. A model swap can multiply that failure quietly. Fluency survives an upgrade. Qualifiers may not.
Our own approach rests on immutable originals with content hashing, append-only audit trails and a documented chain of custody, so that the process holds up if opposing counsel challenges it, and that is the same discipline a firm should demand of any vendor whose model might move. I have come to think the question worth asking is not whether a model changed but whether anyone can prove when it did. As models converge, a theme taken up in what happens when every legal AI has a good model, a vendor will feel less and less hesitation in exchanging one for another, and the contract is the only place where that hesitation can be manufactured, which brings us to the clauses themselves.
What is actually at risk when the model changes partway through your review, and what will you need to prove?
One practitioner warns that a model swap can change a review's coding quickly. The harder question comes later, when you have to show which reviewer coded your documents and how you knew it worked.
Jeffrey Fleming gave a blunt example on an Electronic Discovery Reference Model webcast about AI quality and validation, published October 2, 2026. If a team on a private Claude deployment sees its default model change from Sonnet to Opus, it will get "markedly different answers pretty quick," he said. The team must then explain why the outputs shifted, or go back to the old model.
Explaining a shift is hard because of how a generative AI review earns trust. Young Yu, another panelist, described how much human coding comes before the full run.
What gets checked before the full run
- Test and revise prompts on sample sets of about 50 documents.
- Have humans code 500 or 1,000 documents to train the model.
- Code a control set of anywhere from 385 to 3,000 documents.
- Code an elusion sample of about 1,500 documents.
- Measure precision, recall and elusion, then run the prompts across the full set.
Drawn from Young Yu on an EDRM-hosted webcast (2026) and a practitioner on r/ediscovery (2025).
The litigation firm Quinn Emanuel notes in a March 2026 report that generative AI review "is not an iterative process once the prompts have been run over the corpus." The validation figures therefore measure one model reading one set of prompts. Swap the model and the samples stay on file, but they now describe a different reviewer.
Courts care about those figures. Quinn Emanuel finds "no set standard" for disclosing AI use in discovery, but says courts consistently stress transparency, which "at a minimum, includes disclosing the use of AI to the opposing party." Some courts expected parties to name the software and describe how they validated it. A 2025 decision in the New Jersey insulin pricing litigation linked "some level of transparency and validation of TAR methodologies" to Rule 26(g). Under that rule, a lawyer signing a disclosure certifies, after "a reasonable inquiry," that it is "complete and correct as of the time it is made." Most of this case law concerned older technology-assisted review, and the firm calls the expected level of disclosure "constantly evolving." Still, a disclosure naming one model and its validation could describe a reviewer that did not code every document.
Can a contract simply freeze the model? Partly. Fleming said vendors' AI terms differ for every provider and can even differ by service tier, and he listed "model updates" among the things legal teams should keep watching. A pricing survey of 53 respondents by ComplexDiscovery and EDRM, published in March 2026, describes buyers facing "unresolved questions around exception handling, quality control, and contract structure."
A pin also has a limit that neither you nor the vendor controls. Anthropic, the maker of Claude, says on its model deprecation page: "As safer and more capable models launch, Anthropic regularly retires older ones." A vendor can keep a version pinned only as long as its maker keeps it running. Self-hosting, Fleming said, moves that control in-house, along with "a lot more drift monitoring."
Even a pinned model varies. Yu said asking an AI the same question 100 times will likely produce some variance, and he counts understanding that variance as part of validation. You need to know the normal variance before you blame the model when two runs disagree.
Each kind of proof here points at one model: the validation numbers, the disclosure a court may ask for and the certification a lawyer signs. The vendor decides which model runs, the model's maker decides how long it stays available, and even the same model gives slightly different answers from run to run. A pinned version buys you time. Proof comes from the record: which model, which prompts and which validation sample stand behind each coding decision. We sell AI review, so we have a stake in that conclusion, and our platform keeps append-only audit trails for this reason.
| Risk | What the evidence shows | Record or clause that answers it |
|---|---|---|
| Output shift after a swap | "Markedly different answers pretty quick" | Model version logged per batch, plus a right to revert |
| Validation that describes the old model | Samples are coded before a run that is not repeated | Right to revalidate on the retained control and elusion sets |
| Disclosure gap | Some courts expect the software and its validation to be named | A dated record of model, prompts and results |
| Retirement by the model's maker | The maker regularly retires older models | Advance notice clause and a known retirement date |
| Variance between runs | The same question asked 100 times will likely vary | Measured variance on a sample, kept as a baseline |
- Ask your vendor in writing which model and version will code your matter, and whether that default can change before the matter closes.
- Ask when that version is due to be retired upstream, and whether the vendor will pass the model maker's notice on to you.
- Keep your coded control set and elusion sample, so any new model can be measured on the same documents before it codes anything else.
- Log the model version, prompt version and date against every coding batch, so a disclosure of software and validation matches the documents it covers.
- Run the same prompts on a sample more than once to learn the normal variance before treating a change in coding as a model problem.
How we checked this
We drew on a transcript of a practitioner webcast, a law firm's summary of the case law on AI in discovery, the text of Federal Rule 26, a foundation model maker's deprecation page, an industry pricing survey and a practitioner forum thread. The webcast was presented by an eDiscovery services provider, and its panelists spoke from experience, not from measured studies. We relied on Quinn Emanuel's account of the court decisions and did not read the opinions ourselves. The survey reflects 53 self-reported responses. No figures here come from our own data. We sell AI review software, so we have a commercial stake in how buyers handle model changes. Still unknown: we did not review any vendor's actual terms, and none of our sources describes a court ruling on a model that changed partway through a review.
- Jeffrey Fleming and Young Yu, EDRM webcast transcript on AI quality and validation, JD Supra, October 2, 2026.
- Quinn Emanuel, report on AI in defensive document discovery, March 19, 2026.
- Legal Information Institute, text of Federal Rule of Civil Procedure 26, retrieved October 6, 2026.
- Anthropic, Claude model deprecations page, retrieved October 6, 2026.
- ComplexDiscovery and EDRM, Winter 2026 eDiscovery Pricing Survey analysis, March 6, 2026.
- r/ediscovery, practitioner thread on AI review pricing and prompt testing, July 17, 2025.
Which contract clauses keep an AI review reproducible?
Before signing, bring a messy collection and a hard question to the demo, then ask in writing for a pinned model, advance notice and the right to revalidate.
I ask for three things, and I ask for them before the first prompt is written, because a clause negotiated after the coding has begun is a clause negotiated from weakness:
- Pin the model version for the life of the matter. The order form should name the version that will code your documents and bind the vendor to run it until the matter closes or you release it.
- Require advance written notice before any change, with the right to stay on the current version or to revalidate, at the vendor's expense, before the new version touches a single document of yours.
- Preserve the record of every coding decision: the model version, the prompt or configuration, the inputs, the outputs, the timestamps and the human decisions that accepted or overrode them.
The third clause is where the usual advice stops, and the common assumption is that preservation is enough. It is not. A perfect log of a model you can no longer run is an epitaph, not a reproduction; it tells the court what the reviewer said and cannot make the reviewer say it again.
The economics of reruns sharpen the point. In a July 2025 thread on r/ediscovery, practitioners explained that nobody fed a whole collection to the machine at the outset; they tested and refined prompts on small sets of about 50 documents, and one described AI review as "pay per play," unlike general assistants that allow unlimited retries. A tuned prompt is tuned to one reader. Change the reader and the tuning becomes folklore. If a vendor's upgrade forces that work to be redone, the contract should say who pays, and the answer ought to be the party that chose to change.
Set side by side, the gap between a silent agreement and a negotiated one looks like this:
| Term | What a silent or general change clause leaves you with | What to negotiate | What it protects |
|---|---|---|---|
| Model version | The vendor chooses the model and may replace it | A named version pinned for the matter's duration | Validation stays attached to the reviewer that earned it |
| Notice | Changes take effect when deployed | Advance written notice, with the right to stay or to revalidate first | Time to pause, test and decide |
| Revalidation cost | Each rerun is billed like any other run | Vendor-initiated changes revalidated at the vendor's expense | A rerun does not become a penalty for someone else's choice |
| Records | Logs kept, or discarded, at the vendor's discretion | A per-document record of version, prompt, inputs, outputs, timestamps and human decisions | An answer ready for opposing counsel |
| Training and retention | Depends on the tier purchased | No training on matter data and no vendor retention | Privilege, and the stability of the model itself |
Before: a general clause that, in substance, lets the provider modify or update the services from time to time. After: For each matter, the provider will run the model version named in the matter order, will not substitute another version without advance written notice, and will, at the client's election, either continue the named version or revalidate the new one at its own expense before applying it. The first sentence governs a subscription. The second governs a case.
Two further terms deserve a place, and on both I can speak for our own design. At Relevant, processing runs single-tenant or inside your own AWS account under your own keys, with no vendor retention and no model training, so nothing you feed it trains anyone's model or waives privilege; ask any vendor to put the equivalent in writing. And because our outputs link back to the exact exhibit they came from, an attorney can verify each answer in one click before filing, which is the habit a records clause should make contractual rather than optional. Ask, too, how the vendor reports the documents its pipeline could not process, a question taken up in the processing exceptions your vendor isn't showing you, because a version log that omits them is a log with a hole in it.
The demonstration is the natural moment to raise all of this. A buyer who watches a tool read, code and cite their own evidence, rather than a case study about someone else's, has every reason to ask whether the reader at work that afternoon will still be at work in June, when the second production goes out and the first is already in opposing counsel's hands.
How do you defend an AI review when opposing counsel learns the model changed?
Answer with validation, not with the model's name: show the coding was tested before and after the update, and that the contract let you pause, compare and revalidate.
Volume is about to make the question unavoidable. A reviewer posting on Reddit in September 2026 wrote: "Our company reviewed 1M documents a month last year at this time. Now we are averaging more than 1M per day." The same commenter added, "The floodgates are opening." When one provider codes in a day what it once coded in a month, any update to the model will land in the middle of somebody's matter, and the forecast I draw from that figure is a plain one: within the next year or two, the date on which a vendor changed its model will become a fact a litigator must be ready to explain, as surely as the date a custodian was interviewed.
The reported decisions ask for candor rather than for a frozen model. Disclosing the AI's use to the opposing party has been treated as the minimum, and some courts have gone further, requiring the producing party to name the specific software and describe how its results were validated. Predictive models had already been coding documents for more than a decade by 2024, retraining as reviewers fed them new calls, and their acceptance never rested on their standing still. It rested on the numbers. Apiko, a software developer, praises one legal AI tool in its buyer's guide because each answer shows the agent's decision path, so you can "audit or tweak logic mid-stream"; a firm that accepts that invitation has made a model change of its own, and should date and validate it as carefully as any vendor's. And yet a model swapped between March and June is a witness replaced by a sibling between direct examination and cross, the face familiar and the voice close, with nobody left in the room who can vouch for the testimony.
Were I the one standing up to answer the challenge, I would want the record to show three moves, made in this order, after the vendor's notice arrived:
- Pause new coding on the matter, leaving every document already coded under the old version exactly as it was.
- Compare the updated model against the old one on documents the old version already coded, and keep every disagreement, since the disagreements are what opposing counsel will ask about.
- Roll back or revalidate: return to the pinned version if the calls diverge, or rerun the validation if you adopt the new one, and log the date, the version and the result either way.
None of that sequence can begin until someone tells you the model changed. Under standard terms, the telling is left to the vendor's discretion, which makes the next renewal, rather than the next meet-and-confer, the moment to ask for it in writing.
Summarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently Asked Questions
What else do litigators ask about AI review model updates?
Most questions concern drift, disclosure, cost and unpinned vendors, and the answers share one rule: record the model version and date behind every coding decision.
What is model drift in AI document review?
Model drift is a change in what an AI review tool decides that comes from an update to the model or its settings, not from any change in your documents or your prompts. It matters because validation measures one model's behavior on one collection. Once the model moves, those measurements describe a reviewer who has left the building. The documents stayed put. The reviewer did not.
Do I have to tell opposing counsel which AI model coded the documents?
No reported standard yet names the model version as a required disclosure. Courts have, however, tied transparency about review methods to Rule 26(g), the federal rule that requires a producing party to make a reasonable inquiry in discovery, and the expected level of disclosure keeps evolving. Keep the version and the date behind each call in the file, so that if the question is asked, the answer already exists.
How do firms combine TAR with generative AI review?
One recommended approach sends the documents a generative model found to be a close call into a continuous active learning workflow, the form of TAR that refines its model in real time as human reviewers code. Practitioners have compared the arrival of generative review to the way CAL and TAR eventually became standard practice, and some clients still ask for linear or CAL review. The pairing carries a contractual consequence: two models, two version histories, and both belong in the change log.
What does AI-assisted review cost compared with manual review?
At Relevant e-Discovery, AI-assisted review runs cents per document rather than the dollars per document of manual review, a difference you can see on your own matter. Across the wider market, the winter 2026 pricing survey described earlier found document review rates stable but per-document billing still opaque, and it called outcome-based pricing "nascent." I would read an opaque price list as a reason to read the model terms twice. To price your own collection, reach the team through the contact page at relevantediscovery.com/contact-us.
Does it matter whether AI review comes bundled with my hosting platform?
Bundling changes the price, not the question. A feature included in hosting is harder to decline, or to hold steady, than an add-on bought separately, so the commitment to a pinned version belongs in the hosting agreement itself. Practitioners have also cautioned that these tools cannot yet handle every review scenario or document type, which keeps human judgment, and human quality control, in the loop.
Can I still use AI review if my vendor refuses to pin a model version?
Yes, though the file must then carry what the contract will not. My own rule is plain: validate before each production, log the version and date the vendor reports, and obtain written notice of updates even where you cannot obtain a freeze. The weaker the contract, the heavier the file. And a vendor that will promise neither notice nor a version history is telling you, in the only language a contract speaks, how it expects the next challenge to go.
Written by
Michael
Kansky
Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.
Connect on LinkedIn

