Home › Insights › Reviewers Who See the AI's Tag First Rarely Overrule It

Reviewers Who See the AI's Tag First Rarely Overrule It

Close-up editorial photo of a litigation reviewer

Quick Answer

Yes. A reviewer who sees the AI's code first can anchor on it, so tag-first QC mostly measures agreement. Only a random sample coded blind measures the AI's real error rate.

Blind means the reviewer reads and codes the document before the AI's tag is visible. The evidence behind this article contains no measured override rate for either order, which is to say your own QC log has to produce that number. Cheap AI review only stays cheap if the sample survives opposing counsel's first question about how you checked it.

Did this answer your question?

Key Points

  • A QC sample coded with the AI's tag visible measures agreement, not accuracy; only a random sample coded blind gives an error rate the AI didn't shape.
  • In a 2025 r/ediscovery thread on gen AI first-pass review, one practitioner said how much a team QCs by humans is "entirely up to them."
  • Relevant e-Discovery's AI-assisted review runs cents per document against roughly $19K/GB for manual review, which shrinks the human pass to QC.
Three things litigation teams believe about AI review QC. Myth or fact?
Call each one, then see how other readers called it.
1 A QC reviewer who confirms the AI's issue code has independently checked it.
2 Recall and precision for gen AI review come from a random sample.
3 Faster AI review means less human judgment is needed.
Close-up editorial photo of a litigation reviewer's wooden desk at dusk: a small colored tag clipped to a stack of bound exhibit boards, a row of round color-coded stickers lined up along the desk edge with one covered

A blind QC pass: the reviewer codes the document before the AI's tag is visible.

A tag-first review is one where the reviewer sees the AI's issue code before reading the document. A blind review is the same document, coded by a person who never saw that code. Same document, same issue list, different order. The order is the whole argument.

If you landed here trying to reduce the cost of document review in litigation, or comparing Relativity against other eDiscovery platforms, or searching for the best eDiscovery companies for law firms, you are probably weighing price, speed and features. Fine. Those matter. But every one of those comparisons eventually runs into a quieter question, which is how anyone will know the AI coded the collection correctly once the human pass shrinks to a sample. (It is kind of the least glamorous line in any vendor comparison, which is exactly why it gets skipped.) Honestly, it should be the first one you ask.

Here is what the rest of this piece covers, in order:

  • Why a visible AI tag changes what a reviewer notices, and what anchoring means in plain terms.
  • How to pull a random QC sample that reviewers code without the label, and how recall and precision come out of it.
  • What to ask a platform before you buy, and what to say when opposing counsel asks how you checked the AI.

One caveat before any of that. The evidence behind this article includes no measured override rate comparing tag-first and blind reviewers, and no sample size for such a comparison. I'm not going to invent either. What I can do is show why the blind number is the only one worth defending, and what a team would need to log to produce it on their own matters, starting with the moment a reviewer opens the first document and the tag is either sitting there or it isn't.

Gen AI first-pass review gets validated on a random sample for recall and precision, but the common workflow never says whether the QC reviewer should see the AI's tag first.

The gap is the whole article. Coding blind means a reviewer reads and codes a document without seeing the AI's issue code. Coding tag-first means the AI's answer is on screen before the first line is read. Those sound like the same job. They are not. One produces an independent judgment, and the other produces a reaction to somebody else's.

Practitioners arguing about gen AI review on price make the stakes obvious. In 2025, one buyer said quotes for Rel aiR and Everlaw seemed "more expensive than offshore review," and a commenter explained the pricing model in five words: "You pay per prompt per doc." Honestly, I get the temptation that creates. When every extra look costs money, a quick confirm click on a visible tag feels like thrift (it is, right up until someone asks how you measured your error rate).

So here's my position, stated up front. A QC sample coded with the AI's tag visible measures agreement, not accuracy. Only a sample coded blind gives you an error rate that belongs to the review rather than to the AI. Whether you're weighing Relativity against other platforms or running one messy matter on something smaller, that's the number you'll need when opposing counsel asks.

Fair warning: the evidence doesn't include a measured override rate for legal review, tag-first versus blind. The argument here is about design, and it holds without one, starting with why the order of the tag matters more this year than it did before.

Questions this article answers

Why does the order of the AI tag matter so much right now?

Relevant e-Discovery's AI-assisted review runs cents per document against roughly $19K/GB for manual review, which means the human pass shrinks to QC, and QC has to be honest.

Before trusting any QC pass on AI issue codes, ask three plain questions:

  1. Did the QC reviewer see the AI's code before reading the document?
  2. Was the sample drawn at random, or pulled from documents the AI already flagged?
  3. Is the reviewer's call logged apart from the AI's, so the two can be compared later?

Quick definitions, since the whole argument hangs on them. An issue code is the tag that sorts a document under a claim or defense in the case. A QC sample is the slice of AI-coded documents a human re-reviews to estimate how often the AI got it wrong.

A yes to the first question is the bad one. It means the QC pass measured something narrower than accuracy: how willing one person was to argue with a label already sitting on the screen. (Think of grading an exam with the answer key stapled to the front. You'd catch the howlers, at least.)

The money is why this is urgent now. At a spread like that, AI first-pass review stops being an experiment and becomes the default for any small firm that can do arithmetic, which is to say all of them. Vendors have noticed. One mass-tort vendor case study says AI tools "cut medical record review from days to minutes" and spot connections "that human reviewers often miss." Fine. Maybe. The pitch never asks what happens to the human who reads the AI's answer before reading the record.

The common assumption is that faster AI review means less human judgment is needed. I'd argue the reverse. The small amount of human judgment left in the process is now the only independent check standing between the AI's codes and your signature on a production, so it has to be clean.

On the record side, our platform keeps immutable originals with content hashing, append-only audit trails, a documented chain of custody and a fail-closed privilege gate on production, so the process holds up if opposing counsel challenges it. We also deploy inside your own AWS account, under your own keys, which settles who holds the data. There's a fuller breakdown of who controls what in a managed service versus your own AWS account if that's the part keeping you up.

None of that tells you whether the QC reviewer was looking at the answer key. An audit trail proves what happened. It cannot prove independence. Read together, these 4 sources show the same shape: cost, speed, custody and control all have answers now, and reviewer independence still doesn't.

In practice, your QC sample is where defensibility now lives. If the sample is contaminated, nothing downstream cleans it, and the next question is exactly how a visible tag does the contaminating.

Does seeing the AI's coding first bias human document reviewers?

Yes, it can. One-click links back to the exact exhibit make an AI code easy to verify, but a reviewer who starts from the tag is checking it, not coding independently.

The mechanism has a name. Anchoring is the tendency to lean on the first piece of information you see when you make a judgment. In AI-assisted review, the first thing a reviewer sees is often the AI's issue code, sitting in the coding panel before they've read a single line. The question in their head quietly changes from what is this document about? to is there any reason to disagree? Those are different questions. The second one is much easier to answer no.

Picture the same document, coded two ways:

  • Tag first: The reviewer opens an email already marked responsive to a notice-of-defect issue. They skim for notice language, find a sentence that sort of fits, and confirm. Elapsed time, very little. Reasoning, borrowed.
  • Blind: The reviewer opens the same email with no tag showing. They read it top to bottom, notice it's about a different product line entirely, and code it not responsive to that issue. Then the AI's call is revealed, and the disagreement goes in the log.

Same reviewer. Same document. Different answer, because the first version never asked the question that mattered. (Nobody did anything wrong in the first version, which is honestly the worst part.)

The field hasn't settled how much of this checking should happen at all. In a 2025 r/ediscovery thread on gen AI first-pass review, one practitioner said it plainly: "How much and what a team chooses to QC by humans is entirely up to them." The thread described no standard for how often reviewers check or overrule the AI's output. That silence matters. It leaves the method of QC, including whether the tag is visible, to whoever happens to set up the review.

I'd love to hand you a number here. The evidence doesn't include a measured override rate for tag-first versus blind coding in legal document review, so I won't invent one. The comparison is cheap to run on your own matter, and the result would come from your own QC logs, which is where it belongs anyway.

Here's where the tension gets real. Our outputs link back to the exact exhibit they came from, so an attorney can one-click verify before filing. That's the structural answer to the Mata v. Avianca sanctions risk, and it's why source-linked review is becoming the defensibility bar. But verifying a citation and independently coding a document are two separate jobs. Contrary to how it feels, a fast confirm click on a visible tag is not a second opinion.

We tell buyers to bring a messy collection and a hard question to the demo, then watch the platform read, code and cite their own evidence. That's the right way to decide whether to trust a tool. It's the wrong way to measure a tool's error rate on a live matter, because by then you've watched it work and you're kind of rooting for it. We also built this for the solo or small firm on a single messy matter, the segment enterprise platforms price out, and in a firm that size the person doing QC may be the same attorney who watched the AI code all week.

Small teams can't buy independence with headcount. They have to build it into the order of operations, starting with what the reviewer is allowed to see first.

How can I reduce the cost of document review in litigation and still trust the QC?

Let AI do first-pass coding on private infrastructure with no vendor retention or model training, then spend part of the savings on a small blind QC sample instead of tag-visible re-review.

Start with the sequence practitioners already use. In that same 2025 r/ediscovery discussion, one commenter laid out the gen AI workflow: test the prompts on a small set of about 500 documents, make sure the results are what you want, then "run it on a random sample for recall/precision metrics, then across the overall set (or segments of the set)." The commenter compared it to TAR 1.0, except that "you just tell AI what you want to find" instead of training the system. Another commenter there argued that the savings come from total project cost, not hourly or per-document rates. If you're coming from a TAR mindset, it's worth knowing that predictive coding isn't the only defensible AI review, and a validation sample is how the newer methods show their work.

Two terms carry the weight here. Recall is the share of truly relevant documents the review actually found. Precision is the share of documents coded relevant that really are relevant. Both numbers come from that random sample. Which is the whole problem, because if the sample is coded with the AI's tag on screen, both numbers inherit the AI's answer.

The fix doesn't add a step. It changes how one existing step is run:

  1. Draw the sample at random from the full set before anyone reviews AI results for those documents.
  2. Hide the AI's issue codes in the QC view, so the reviewer sees only the document.
  3. Have the reviewer code each document from scratch, with a one-line reason for each call.
  4. Unblind, compare the reviewer's call to the AI's, and log every disagreement before anyone resolves it.

Recall and precision then come from the blind calls, not the confirmed ones. Everything else in your QC plan can stay exactly as it is. The blind sample is the only number in the file the AI didn't help write.

On cost, I agree with the total-project framing, and it cuts both ways. Under per-prompt, per-document pricing, re-running a full tag-visible QC pass is the expensive way to learn very little. A blind random sample is a bounded slice of the set, coded once, and it produces the one error-rate figure you can stand behind.

Where the sample lives matters too. Our processing runs single-tenant, or inside your own AWS account under your own keys, with no vendor retention and no model training, so nothing your reviewers code in the blind sample trains anyone's model or waives privilege. We also keep the defensible spine (Bates, privilege, chain of custody) and the RAG side (cited Q&A and chronologies) in one tool rather than two. A QC result stored apart from the production record is one more thing somebody can ask you to reconcile.

Then there's the disagreement log, which is where this gets interesting, and also the part most QC plans quietly leave out.

Which platform question should you ask before you buy AI review?

Before you compare platforms or price per-prompt plans, run a blind QC sample on your own messy matter and see whether the AI's issue codes survive a reviewer who never saw them.

Whether you're weighing Relativity against other platforms or asking which e-discovery companies suit a small firm, put one question to every vendor: can my reviewers code a random sample without seeing your AI's answer first? Short answers are fine. Vague ones aren't.

The next 12-24 months, scored

Where AI review QC sampling heads next

Forecasts on how litigation teams will validate AI issue coding, pay for that validation, and staff the human check over the next two years.

1 sources analyzed1 community discussion
A

What changes in AI review validation

Use each forecast to plan sampling, budgets and reviewer workflow before your next AI-assisted document review.

Contrarian signal
48/100
Low confidence 12-24 months

Gen AI review that skips TAR-style training will not shrink human involvement. It will push reviewer hours into prompt testing and independent QC. Vendors claiming review drops from days to minutes will increasingly be asked for blind-sample results to back that claim.

Faint signals worth tracking: Practitioners already describe a sequence for gen AI review. They test prompts on about 500 documents, validate the results, then run a random sample to get recall and precision metrics before scaling to the full set. Practitioners note that gen AI review resembles TAR 1.0 except that it needs no training. Vendor case studies claim review time falls from days to minutes without sacrificing quality.

B

Sources behind the review QC forecasts

Each source below is listed with the specific line that supports a forecast on AI-assisted document review and its quality control.

Source What it states Forecasts it backs
Gen AI for First Pass Review - More expensive than human review [Community / Forum] Run on a random sample to get recall and precision metrics. “Its kind of the same process as for TAR 1.0 but whereas in machine learning you train the system, you just tell AI what you want to find. There's less…”
Commenter 3 says the workflow resembles TAR 1.0, except that gen AI needs no training. The user describes what to find instead.
Blind random samples become the standard AI review check
No-training AI moves human effort into validation
Where the forecasts come from: every public source, the line it contributes, and the calls it supports.
C

What would shift these review QC forecasts

These scenarios would weaken or reverse the forecasts on blind sampling, validation effort and AI review pricing.

A note on uncertainty

Treat these scores as weights, not verdicts: “Blind random samples become the standard AI review check” leads on evidence, and “No-training AI moves human effort into validation” is where the sources disagree most.

  • Courts or opposing counsel could routinely accept vendor accuracy claims without sample-based recall and precision figures.
  • Per-run pricing could make repeated validation samples too costly for most matters.
  • Review platforms could also build independent sampling into the workflow, which would remove the need for a separate blind check.
Methodology Each forecast is scored 0-100 from the public sources shown for it: how many there are and how authoritative they are.

What will matter most in AI review QC over the next 12 to 24 months?

Random-sample recall and precision checks will become the expected test of gen AI review, and blind-coded samples will be the ones that survive a challenge.

PredictionWeak signal todayWhy it mattersSource
Blind random samples become the standard check before anyone relies on AI issue codes.Practitioners already test prompts on a small set, then validate on a random sample for recall and precision before running the full set.A tag-visible sample can overstate how often the AI is right. A blind one gives a figure that holds up if opposing counsel challenges the process.2025 practitioner thread on gen AI first-pass review
AI review without a training step moves human hours into validation instead of removing them.Practitioners describe gen AI review as TAR 1.0 minus the training: you describe what to find rather than teaching a model.Staffing plans that assume reviewers disappear will miss the QC work that makes an AI-coded production defensible.Same 2025 thread
Pricing models decide how often a team can afford to re-validate.In 2025, pricing had already split between per prompt, per document, per run charges, bundled prompts and unlimited prompting.Under per-run pricing, every fresh validation sample costs money, so the blind sample has to be planned and budgeted up front.Same 2025 thread

The middle row is the one people will argue with. The pitch for gen AI review is that it needs no training, so humans shrink out of the picture. I think the opposite happens. The training step goes away, and the human work reappears downstream as prompt testing and QC. Which is fine, honestly, as long as somebody budgets for it.

I'll be plain about the uncertainty. Nothing in the evidence measures how much a visible tag suppresses overrides in legal review. The first prediction rests on design logic and on where practitioners already are, not on a published rate.

Three developments would change the forecast:

  • Courts or opposing counsel routinely accepting vendor accuracy claims without sample-based recall and precision figures.
  • Per-run pricing making repeated validation samples too costly for most matters.
  • Review platforms hiding AI codes by default in QC views, which would settle the question without anyone having to argue it. (I'd enjoy being wrong that way.)

Until one of those happens, the sample is yours to design, and a reviewer can't unsee a tag.

What belongs in every AI-review QC plan from here on?

A random sample coded blind, with every disagreement logged, belongs in every plan, because it's the only error-rate figure in the file that the AI's own tags didn't shape.

Here's where I think this goes. Practitioners already describe running a random sample for recall and precision before trusting gen AI across a full set. The next step is obvious, and cheap: code that sample blind. I expect blind sampling to become the checkpoint opposing counsel asks about first, because a tag-visible sample simply can't answer the question they'll ask, which is how you know the AI was right.

When first-pass review costs cents a document instead of dollars, the review stops being where the money goes. The QC sample is where the credibility goes. Even this year's eDiscovery Day is staging a live debate that puts AI on trial before a sitting federal judge. The bench is clearly curious. (So, presumably, is the other side.)

So whatever platform you land on, Relativity or one of the smaller e-discovery companies built for firms like yours, hide the tag on the sample. Log the fights. Keep the number.

Written by

Michael

Kansky

Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

What do litigation teams ask about blind QC for AI review?

Most questions come down to cost, defensibility and logistics: what a blind sample is, who should code it, and how it holds up when opposing counsel asks how you validated the AI.

What is a blind QC sample?

A blind QC sample is a random set of AI-coded documents that a reviewer codes without seeing the AI's issue codes. Only after the reviewer commits to a call is the AI's answer revealed and compared. That order is the whole point.

How do I make sure an AI review holds up if opposing counsel challenges it?

Keep three things in the record: how the random sample was drawn, proof that it was coded blind, and the disagreement log. Report recall and precision from the blind calls only. When the other side asks how you know the AI was right, you can show them an answer the AI didn't help write.

Should the blind reviewer be someone other than the person who ran the AI review?

I'd prefer it, but a solo or small firm may not have a second attorney to spare, and that's fine. What matters most is that the reviewer doesn't see the AI's tag on the sampled documents before coding them. A fresh pair of eyes helps. A hidden tag helps more.

Can I measure the tag-first versus blind override gap on my own matter?

Yes. Log every QC call with a flag for whether the tag was visible, then compare how often each group overturned the AI. The evidence here doesn't include a published rate for legal review, so your own logs give you the first real number. Bring it to the meet-and-confer.

What should I ask when comparing Relativity with other e-discovery platforms?

Ask whether the QC view can hide the AI's codes on a sampled set, and whether the disagreement log stays with the production record. Per-prompt, per-document pricing matters too, because it decides how cheaply you can rerun a validation sample. Get the answer on your own collection, not on a slide.

What does Relevant e-Discovery charge, and how do I get started?

Relevant's AI-assisted review runs at cents per document, a difference from manual review you can see on your own matter. To scope a blind QC plan for your case, reach the team through the contact page at relevantediscovery.com/contact-us.

Read next

Stack of produced discovery documents under a desk lamp at night, one page showing faint hidden marks beside a laptop running AI document reviewEdiscovery

Can a Planted Document Fool Your AI Review Tool?

Yes. A planted document refers to a produced file carrying hidden text, white on white or tiny type, that the attorney never sees but the AI reviewer reads and may obey.September 30, 202630 min read
Laptop showing a long workplace chat thread beside a tall stack of printed pages with only a few flagged for reviewEdiscovery

What Share of a Chat Collection Is Actually Relevant?

Nobody has published a measured figure. A chat collection's responsive share refers to the fraction of gathered messages that a request for production actually calls for, and the only defensible number is one measured on your own matter.September 28, 202627 min read
Laptop running a homemade AI review workflow next to boxes of litigation documents and a sealed evidence boxEdiscovery

Building Your Own AI Review Agent: Where It Breaks

Yes, you can build one in an afternoon. A do-it-yourself AI review agent refers to a trigger, a container job and a model call, and it skips what makes review defensible.September 27, 202632 min read

See it on your matter

Bring us a messy collection - mailboxes, scans, phones, recordings - and watch it become one searchable, defensible record.