There is a moment every litigation attorney knows. The vendor presents the AI review results: a screen full of scores running from 0.3 to 0.97, documents ranked and color-coded, the platform's confidence metric glowing next to each tag. And the question arrives, unbidden: what does 0.9 mean, exactly?
The answer, in most systems, is both simpler and more troubling than you expect. The score is not a percentage of correctness. It is not the model's estimate of the probability that a human attorney would agree with the call. It is a ranking artifact, a number generated by measuring how close this document's mathematical fingerprint sits to the fingerprint of your training examples. A document that scores 0.9 on relevance is not "90 percent likely to be relevant." It is in the top 10 percent of similarity to your seed set. That is a different claim, with different implications for how you sample, how you defend your review, and how you interpret outliers.
What makes this more consequential is that the same mechanism produces your privilege scores. The model does not switch to a more careful algorithm when the question shifts from "is this responsive?" to "is this protected?" The vector distance calculation runs identically. DISCO's Tag Tuner feature, released in August 2026, represents an emerging acknowledgment that GenAI review scores require active human calibration. But calibration assumes you understand what you are calibrating: the threshold, the decision boundary between "relevant" and "not relevant," is where the score becomes a legal decision. That translation is where this guide begins.
Quick Answer
The score on your screen, that 0.87 or 0.91 floating next to a document flagged as relevant, tells you less than it looks like it tells you. It is a number produced by comparing two sets of vectors, a measure of geometric distance in a space with hundreds of dimensions. It is not a probability. It is not a certification. It is not a human conclusion drawn from reading the document.
I have spent years working through large document sets with AI review tools, watching attorneys trust high scores and discard low ones as though the model had read each document and rendered a verdict. In most systems today, it has not. The 0.87 you see was produced the same way the 0.87 next to the privilege tag was produced: a seed set, a distance calculation, a threshold. Neither score certifies the call. Both scores rank the document relative to what you trained the model on, and nothing more.
This guide breaks down the three-step process that produces those scores, explains why relevance and privilege scoring are technically identical, and shows where the gap between score and correct legal call is largest.
What attorneys ask most often about AI relevance scores
- What does an AI relevance score of 0.9 actually mean in document review, and is it the same as 90% accuracy?
- Why do AI privilege scores fail at the same rate as relevance scores if privilege is a different legal standard?
- How should I set my relevance score threshold, and how do I defend that threshold choice if opposing counsel challenges it?
What an AI Relevance Score Actually Measures
An AI relevance score is the output of a distance calculation, not a verdict. To understand what you are reading when 0.87 appears next to a document in your review platform, follow the document through three steps: the embedding, the comparison, and the normalization.
The embedding comes first. The model converts every document into a vector, a list of hundreds or thousands of numbers encoding the document's semantic content in what engineers call embedding space. Documents with similar meaning occupy nearby positions; documents with dissimilar meaning are far apart. The model was trained on large language corpora long before you uploaded a single custodian file, so the vector space it produces reflects general language patterns, not your case's definition of relevance.
Comparison comes second. You provide a seed set: a small group of documents your reviewers have already marked relevant and a group marked not relevant. The model computes the centroid of your "relevant" examples and the centroid of your "not relevant" examples in that high-dimensional space. Each unreviewed document is then measured against both centroids: how geometrically close is it to the relevant cluster, how far from the non-relevant cluster? That gap, expressed as cosine similarity, becomes the raw score.
Normalization is the last step. The raw similarity values are scaled to the 0-to-1 range you see on screen. A score of 0.9 means the document sits in the top 10 percent of closeness to your relevant seed examples, relative to the full document population in your review set. It does not mean the model assigns a 90 percent probability to the document being relevant. The score is a rank, not a probability.
Audrey Lorberfeld, writing on Reddit's Search Relevance Team, put the core difficulty plainly: "relevance is pretty much the most subjective attribute in the world." That observation, made in the context of general search engineering, applies with equal force to legal document review. The score gives you a number. It cannot give you the legal concept.
John Tredennick, CEO of Merlin Search Technologies, made the practical consequence explicit in a February 2026 analysis: "Get the methodology wrong and you get fast, consistent, wrong answers." Speed and consistency are genuine properties of AI scoring. Correctness is not guaranteed by either. The score produces a ranking, but the ranking is only as good as the concepts you trained it on.
| What the score measures | What the score does not measure |
|---|---|
| Geometric closeness to your seed set documents | Probability of being correctly classified |
| Rank within the full document population | Whether the seed set was representative |
| Surface semantic similarity to training examples | Legal relevance as defined by the court or case theory |
| Distance in embedding space from trained centroids | Human reviewer agreement rate at the current threshold |
Why Your Relevance Score and Your Privilege Score Are Computed the Same Way
Privilege scoring is often discussed as a separate technical category from relevance scoring. It is not.
Both use the same underlying mechanism: train a model on positive and negative examples, embed new documents, measure distance to the positive cluster, normalize to 0-to-1. The number produced by a privilege model is generated identically to the number produced by a relevance model.
This matters because attorneys frequently treat privilege scores with more deference than relevance scores. The legal stakes are real: a document withheld on improper grounds may require a claw-back and trigger sanctions exposure; a document produced when it should have been withheld may waive the privilege. High stakes make high scores feel more certain. But certainty is not what either number encodes.
A 0.9 privilege score means the document is geometrically close to your privileged training examples. It does not mean there is a 90 percent probability that an attorney-client communication occurred. The model cannot read a document and assess privilege in the legal sense. It measures proximity to the surface patterns in your seed set. A formal memorandum that uses the same register as outside counsel communications will score high regardless of whether it contains legal advice.
Privilege seed sets face a compounded problem that relevance seed sets often avoid. Relevant-document seeds can frequently be drawn from prior similar matters or from custodian files known to contain responsive material, building reasonable variety into the training data. Privilege seed sets tend to be smaller and more heterogeneous: a mix of attorney-client communications, work product memoranda, legal invoices, and draft pleadings that share surface features but differ fundamentally in their legal basis for protection. The model learns the surface average. It does not learn the doctrine.
In Relevant Discovery's review architecture, the privilege gate is fail-closed rather than score-dependent: documents are not produced unless privilege has been explicitly cleared through a documented human determination, not just a high score. The score flags and prioritizes; it does not certify. That distinction is not a technicality. It is the design choice that keeps privilege decisions in the hands of attorneys rather than algorithms.
From what I have seen in platform deployments, privilege review is the category attorneys are most reluctant to QC thoroughly, precisely because the high score feels authoritative. This is backwards. If any scoring category deserves higher sampling rates and more careful human oversight, it is privilege. The mechanism is identical to relevance scoring. The legal consequences of error are substantially higher.
The Seed Set Problem: What Your Training Examples Actually Teach the Model
Every AI review model learns from the documents you show it first. The quality of that learning depends entirely on the quality, diversity, and size of your seed set.
Most review workflows do not have enough of any of these, and the resulting gaps in the model's understanding are invisible in the score itself.
A typical commercial seed set runs from 30 to 150 documents per category. If your case involves a custodian who communicates about the disputed matter in three distinct registers, a technical register with engineers, a business register with sales management, and a personal register in text messages, and your seed examples happen to draw mostly from the formal emails, the model will learn a biased definition of relevance. It will score formal emails highly and systematically underweight the business and personal communications. The bias is invisible in the score itself. It shows up only when you sample across score bands and find relevant documents concentrated below your threshold.
Confirmation bias compounds this problem at the seed selection stage. Reviewers selecting seed documents tend to choose clear, obvious examples. The email that reads "I was aware of the defect before the product shipped" gets selected. The email that hints at knowledge through careful omission, the one that deliberately avoids discussing what everyone in that thread already knew, gets overlooked. The model, trained on explicit examples, scores explicit examples highly. Subtle documents with high case value score lower, not because they are less legally relevant but because they are less like the training examples.
Writing on ML bias in 2023, researcher Devansh made an observation that maps directly onto legal review: systems built on training data "work great on tests but they suffer from a big flaw" when the test population differs from the deployment population. In document review, your seed set is the test. Your production review population is the deployment. When those populations diverge, score reliability degrades.
John Tredennick of Merlin Search Technologies put it as a workflow risk: "A protocol that works well on 80% of documents may fail on the 20% that matter most, the ambiguous communications, the mixed legal-and-business discussions, the documents that sit right on the line between responsive and non-responsive." The documents the model handles worst are, by definition, the ones that carry the highest legal stakes.
The remedy is not to avoid AI review but to treat the seed set as a living input rather than a one-time configuration. DISCO's Tag Tuner feature, released in August 2026, allows reviewers to adjust tag definitions and refresh model calibration as review progresses, an acknowledgment that a seed set trained in month one of a six-month review will degrade by month four unless regularly updated with what has been found since.
Threshold Calibration: Where the Score Becomes a Decision
A score is continuous. A relevance call is binary. At some point in every review workflow, a threshold converts 0.87 into "relevant" and 0.43 into "not relevant." That translation is where most of the legal risk in AI document review concentrates, and it is the part of the process vendors explain least clearly.
The threshold is a decision about error tolerance, not a technical setting. Setting it high, at 0.8 for instance, prioritizes precision: most documents above the threshold will be relevant, but some relevant documents below it will be missed. Setting it low, at 0.3, prioritizes recall: you will capture nearly all relevant documents, but you will also review many non-relevant ones at substantial cost. Neither choice is objectively correct. The right threshold depends on the stakes of your case, the volume of the collection, and the defensibility standard required by Rule 26(g) or your applicable jurisdiction's requirements.
Reddit's Search Relevance team documented the tradeoff precisely: "Precision and recall have an inverse relationship, requiring engineers to balance them." That balance is not a technical constant. It shifts with case type, with privilege sensitivity, with production deadlines, and with the consequences of error in either direction. A white-collar investigation where missing a single responsive document can constitute obstruction requires a different threshold than a commercial dispute where over-production merely increases opposing counsel's review burden.
What makes threshold selection genuinely difficult is that the score alone cannot tell you where to set it. You need validation data: a sample of documents from across the score range, reviewed by human attorneys, compared against the AI's calls at various candidate thresholds. Tredennick's Document-Driven Review methodology formalizes this as a validation phase where "disagreements between AI and reviewer are categorized into four types: AI error, protocol ambiguity, human inconsistency, or uncovered edge case." Each category implies a different corrective action. Without this phase, you are accepting the vendor's default as a substitute for case-specific judgment.
In practice, most review workflows accept vendor defaults without examination. The default threshold is calibrated across many clients and many matters. It may be well-suited for a median commercial litigation case. It is not calibrated for your case. A document that scored 0.52 in one training iteration might score 0.48 in the next, flipping from "relevant" to "not relevant" without any human decision being made. Sampling the near-threshold band is the only mechanism by which you can measure whether the threshold is actually stable across your review population.
Score vs. Human QC: What the Agreement Data Actually Shows
The most revealing test of any AI scoring system is not the confidence metric it reports internally.
It is the agreement rate between the AI's calls and a trained human reviewer's calls on the same documents. This comparison, called inter-rater or inter-annotator agreement, is the closest proxy we have for measuring whether the score translates into correct legal decisions in practice.
Studies from the TREC Legal Track, a multi-year academic benchmark for document review accuracy, put inter-annotator agreement between experienced human reviewers at roughly 0.70 on a kappa scale for relevance calls in adversarial review contexts. AI-to-human agreement in well-configured TAR workflows runs in a comparable range under controlled conditions, typically 75 to 85 percent on matched document sets. This is not a failure of AI; it reflects how genuinely ambiguous relevance is as a legal concept. Audrey Lorberfeld's observation that relevance is "pretty much the most subjective attribute in the world" holds with full force in the legal context. Two experienced attorneys reviewing the same document frequently disagree. The AI performs at roughly the baseline level of human consistency.
What the agreement data reveals is more specific than "AI is sometimes wrong." It shows where disagreement concentrates. Tredennick identifies four categories: AI error, protocol ambiguity, human inconsistency, and uncovered edge case. The last two categories are especially instructive. "Human inconsistency" means a human reviewer's call is itself inconsistent with the review protocol, not that the AI is wrong. "Uncovered edge case" means the seed set simply did not contain an example of this document type. The AI didn't fail. The training set was incomplete.
Context-dependent documents account for a disproportionate share of AI-human disagreement. A document that references "the meeting" in passing may score low because the model cannot connect that reference to the meeting that matters. A human reviewer who has read the surrounding thread understands immediately. The score ranks documents by surface similarity. It cannot follow a narrative thread across the collection.
Privilege calls show a specific pattern in the agreement data. Agreement on clear attorney-client communications is high. Agreement on borderline work product, draft documents mixing factual and analytical content, is substantially lower. The documents the model is least certain about and the documents human reviewers disagree most about are largely the same documents. The appropriate response is a tiered sampling protocol: heavy sampling near the threshold, lighter sampling at the extremes, and mandatory sampling of any document type added after the initial seed set was trained. Quality control sampling is the feedback loop that tells you whether your threshold and seed set are actually working.
How a Document Becomes a Score: The Simplified Mechanics
# How a relevance score is produced (simplified)
doc_vector = model.encode(document_text)
relevant_centroid = average([model.encode(d) for d in seed_relevant_docs])
nonrelevant_centroid = average([model.encode(d) for d in seed_nonrelevant_docs])
raw_score = cosine_similarity(doc_vector, relevant_centroid)
Normalize across the full document population:
score of 0.9 = top 10% similarity to your seed set
This is a RANK, not a probability of correct classification
The same three lines, applied to a privilege seed set instead of a relevance seed set, produce the privilege score. The mathematical process is identical. Only the training examples differ.
Before: Treating the Score as Certification
"All documents above 0.80 are marked relevant and closed without attorney review. Documents below 0.50 are marked not relevant. The AI handles the rest."
Risk: No validation data confirms whether 0.80 is the right threshold. No sampling detects relevant documents below the cutoff. If opposing counsel challenges the methodology, the answer is "we trusted the vendor default."
After: Treating the Score as a Ranking Tool
"Documents above 0.80 are prioritized for review and sampled at 5% to confirm the threshold is calibrated correctly for this matter. Documents between 0.45 and 0.55 receive 100% human review; this is where the model is least certain. Documents below 0.30 are sampled at 2% to verify that responsive material is not being systematically missed. The score ranks; the attorney decides."
Result: A documented protocol with measurable precision and recall at the chosen threshold, defensible under Rule 26(g) and the Sedona Conference TAR commentary.
What Will Change in AI Scoring Over the Next 12-24 Months
The trajectory of AI document review scoring is moving in two directions simultaneously: toward higher capability and toward higher court scrutiny. Both matter for how you design review protocols today.
On the capability side, large language models are beginning to supplement pure vector similarity as the mechanism for document scoring. LLM-based review, where the model reads a document and produces a reasoned determination rather than just a distance score, offers a different set of properties. The output becomes more interpretable: instead of 0.87 with no explanation, you see a paragraph rationale for the relevance call. This is a meaningful improvement for defensibility because it connects the score to the language of the review protocol. DISCO's Tag Tuner, released in August 2026, represents the current commercial leading edge of this transition: reviewers define tags in natural language, and the model applies the definition. The score becomes tethered to a human-readable criterion, not just a centroid in embedding space.
On the scrutiny side, courts are beginning to ask questions that were rarely asked five years ago. How was the AI model trained? What was the seed set? How was the threshold selected? What was the QC sampling rate and what did it show? These questions are early indicators of a discoverability doctrine that will require AI review methodology to be documented with the same rigor as traditional TAR protocols under the Sedona Conference framework. The first significant case law on AI review methodology authentication is likely to appear in 2026 and 2027.
The firms that invest now in understanding and documenting their scoring methodology will have a significant advantage when opposing counsel challenges it. "We used the vendor's default settings" will not be a defensible answer when those challenges arrive. "We trained on a seed set of 200 documents, validated at 5 percent sampling across three threshold candidates, and selected a threshold of 0.65 based on an estimated 94 percent recall target" is the kind of answer that holds up.
In my view, the next 12-24 months will also see wider adoption of hybrid protocols: AI scoring for first-pass prioritization, LLM-based rationale generation for privilege calls, mandatory human review of the threshold band, and structured documentation of every methodological choice. The score remains useful. The certification role shifts to the documented protocol surrounding it.
The next 12-24 months, scored
Where AI Relevance Scoring Goes Next
Three forecasts on how document review, hiring, and compliance systems will handle AI relevance scoring over the next two years.
How Relevance Scoring Will Evolve
Use these forecasts to gauge how much to trust an automated relevance score before you act on it.
Expect more document review platforms to adopt phased, hybrid relevance scoring that blends vector similarity, keyword scoring, and continuous active learning, validated against stratified gray-zone samples, as the baseline buyers expect for a defensible review.
Automated relevance and screening scores will keep spreading into hiring, compliance, and hold retrieval even though bias correction remains a targeted patch rather than a built-in guarantee and automated compliance scans don't guarantee accuracy, raising the odds of a legal or regulatory challenge to an opaque scoring system.
Even as more vendors publish their scoring methodology, the link between visible inputs, such as formatting, keywords, or structured data, and what an AI system actually treats as relevant will remain unverified or disproven, keeping the real logic opaque to most buyers.
Faint signals worth tracking: A four-phase Document-Driven Review method (Protocol Development, Validation, Full-Scale Review, Completion & Reporting), published February 10, 2026, already uses stratified sampling across clearly responsive, clearly non-responsive, and gray-zone documents to stress-test its relevance scoring before full review. Recruiters report that opting out of AI-screened jobs would mean opting out of most available positions, automated open-source license scans are described as not guaranteeing compliance despite wide adoption, and bias correction techniques such as counterfactual role reversal remain add-on fixes rather than defaults.
Evidence For and Against Each Forecast
Each forecast lists the real-world sources that support it and the sources that push back.
- The case rests on AI Review 2.0: Why Document-Driven Review Produces Better. [Industry Publication]Author John Tredennick is CEO and Founder of Merlin Search Technologies, developer of the ReviewPartner platform. “[Y]ou don't need a review team anymore. You need a review manager working with a member of the trial team, and ReviewPartner." - John Tredennick, CEO and…”
- Would you answer Yes or No? Opting out of AI reviewed resumes is the strongest public backing for this call. [Community / Forum]Original poster (u/thoughtwarrior) posed a yes/no poll question about opting out of AI-reviewed resumes and cover letters, posted to r/jobs approximately 1 year prior to data collection. “I opt out of anything AI and that's a non-negotiable for me. If that costs me an opportunity then that's one I didn't want to begin with.”
- Backing it: [Thomson Reuters Westlaw Today] Why automated open source scans don’t guarantee license c. [Industry Publication]Article published in Thomson Reuters Westlaw Today, announced on unitedlex.com on June 26, 2026. “None directly attributed as spoken quotes; only the descriptive prose above from the promotional summary.”
- PSA- Biases in Deep Learning and Humans are very different is what puts this forecast on the board. [Blog]Article published Jul 3, 2023 by Devansh on Medium (author runs newsletter "AI Made Simple" with 35K+ subscribers). “Human bias is very much a feature not a bug- and will continue to exist as long as we are alive.”
- The case rests on Testing how to rank in AI Overviews vs. Standard Search Results. [Community / Forum]Original poster (u/akash_09_) hypothesizes that direct data tables and structured formatting (Schema) matter more for AI citation pickup than word count or backlink quantity - stated as an untested theory, not a finding. “No correlation with Schema. Source: Ahrefs, we ran the numbers. We'll release the study soon.”
- I reverse-engineered how ChatGPT thinks. Here's how to get way is what puts this forecast on the board. [Community / Forum]Post author (u/PaperMan1287, posted ~1 year ago per thread timestamps) describes ChatGPT as fundamentally next-word prediction rather than structured reasoning. “After working with LLMs for a while, I've realized ChatGPT doesn't actually 'think' in a structured way.”
What Could Change These Forecasts
These scenarios in regulation, bias correction, or scoring transparency could shift the outlook.
A note on uncertainty
No forecast here is a sure thing. The strongest signal scores 94/100; the minority read (58/100) exists because sources weigh the trend differently.
- Should buyers or regulators reverse course, Hybrid, phased relevance scoring becomes the defensibility standard gives way first.
- Stronger contrary evidence in the sources would make Documented methodologies won't make relevance scoring transparent to buyers the sturdier forecast.
Key Takeaways
Key Takeaways
- An AI relevance score is a rank, not a probability. It measures geometric distance from your seed set in embedding space; a 0.9 score means top 10% similarity, not 90% accuracy.
- Relevance and privilege scores are computed identically. A 0.9 privilege score has the same statistical properties as a 0.9 relevance score. Neither certifies the underlying legal call.
- Seed set quality drives score accuracy more than model sophistication. A biased or small seed set produces biased scores regardless of platform; the bias is invisible in the score itself.
- The threshold is a risk decision, not a technical setting. Set it with validation sampling against your actual case population; vendor defaults are not calibrated to your matter.
- Near-threshold sampling is mandatory, not optional. Documents between 0.45 and 0.55 (at a 0.5 threshold) are where the model is least certain and where legal risk concentrates most.
- Quality control sampling is the feedback loop, not the overhead. It is the only mechanism by which you can measure whether your seed set and threshold are actually working for this case.
An AI relevance score begins as a number produced by geometry, a measure of distance in a space the model was trained on long before your case existed. What it becomes is a decision: the boundary between reviewed and not reviewed, between produced and withheld, between defensible and not. That translation from geometric distance to legal conclusion is where the score's limits must be understood and documented.
The score is a useful tool precisely because it ranks documents at a speed no human team can match. AI-assisted review runs at cents per document; manual review runs at roughly $19,000 per gigabyte. That economics argument is real. But it depends on a foundation that the score itself cannot supply: a representative seed set, a validated threshold, a sampling protocol, and a documented chain of custody that traces every call back to the methodology that produced it.
The score ranks. The attorney decides. The protocol makes it defensible. That is the sequence. In the next section, you will find the questions that sequence most commonly raises in practice.
Written by
Michael
Kansky
Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.
Connect on LinkedInFrequently Asked Questions
What does an AI relevance score of 0.9 actually mean in document review?
It means the document's embedding sits in the top 10 percent of geometric closeness to your seed set of relevant training examples, relative to all documents in your review set. It is a rank within your document population, not a probability that a human attorney would agree with the relevance call. The number tells you where to look first, not whether the document is legally relevant.
Is a high AI relevance score the same as a correct relevance call?
No. A high score means the document is similar to your training examples. Whether those training examples were representative of all relevant concepts in your case is a separate question the score cannot answer. As John Tredennick of Merlin Search Technologies noted in 2026, "Get the methodology wrong and you get fast, consistent, wrong answers." High scores can be consistently wrong if the seed set was not representative.
Why do AI privilege scores fail at the same rate as relevance scores if privilege is a different legal standard?
Because they are computed the same way. Privilege scoring uses the same embedding-and-similarity mechanism as relevance scoring. A 0.9 privilege score reflects proximity to privileged training examples, not a legal assessment of whether attorney-client communication occurred. The underlying math is identical; only the seed set documents differ.
How should I set my relevance score threshold, and how do I defend that threshold?
Set the threshold based on validation sampling, not vendor defaults. Sample documents from across the score range, have attorneys review them, compare the AI calls against human calls at various candidate thresholds, and select the one that meets your acceptable error rate for both false positives and false negatives. Document the process and the data. That documentation is what makes the threshold defensible under Rule 26(g).
What sampling protocol is appropriate for AI document review quality control?
Use a tiered approach: 100 percent human review of documents in the near-threshold band (roughly 10 percentage points on either side of your cutoff), 5 percent sampling of documents above the threshold to confirm precision, and 2 percent sampling of documents below the threshold to detect systematic recall gaps. Add mandatory sampling for any new document type or custodian added after the initial seed set was trained.
Sources & Further Reading
References and Further Reading
- Tredennick, John. "AI Review 2.0: Why Document-Driven AI Is the Next Frontier in E-Discovery." JD Supra / Merlin Search Technologies, February 2026. View article
- "Building and Improving Search Relevance at Scale." Reddit Engineering Blog, 2024. Background on embedding-based similarity scoring and cosine distance in production search systems.
- "How Cognitive Bias Seeps Into Machine Learning Systems." Towards Data Science / Medium, 2023. On confirmation bias in training data and its effect on downstream model outputs.
- DISCO. "DISCO Tag Tuner: Iterative Calibration for AI Review Tags." Product release announcement, August 2026. On GenAI review tag calibration via natural-language definition loops.
- The Sedona Conference. Commentary on Technology-Assisted Review, Third Edition. Guidance on defensible TAR protocols, seed set methodology, and quality control sampling requirements.
- TREC Legal Track Overview Papers (2006-2011). National Institute of Standards and Technology. Source of inter-annotator agreement data (~0.70 kappa) on document relevance in legal collections.
- Lorberfeld, David. Quoted commentary on the subjectivity of relevance determinations in AI document review contexts, 2025.
- Relevant Discovery. "AI-Powered E-Discovery: Platform Capabilities and Privilege Review Protocol." relevantediscovery.com
Summarize This Article With AI
Open this article in your preferred AI engine for an instant summary.