HomeInsightsWhy AI Semantic Search Still Misses the Documents That Matter

Why AI Semantic Search Still Misses the Documents That Matter

Abstract visualization of a digital document retrieval system: rays of blue light scanning through a dark archive of floating documents, with some documents containing alphanumeric codes and contract identifiers

Key Points

  • A 15-30% recall gap separates hybrid search from pure semantic search on complex litigation sets, per Applied AI Research (2025); the gap is widest on account-number and contract-code queries.
  • LlamaIndex explicitly identifies exact-match queries - product codes, legal citations, account numbers - as the category where keyword search reliably outperforms semantic retrieval.
  • By 2027, AI search methodology challenges are expected to be as routine as custodian disputes; a semantic-only protocol with no keyword log cannot satisfy that standard.
Three things litigators believe about AI document search. Myth or fact?
Call each one, then see how other readers called it.
1 AI semantic search retrieves all relevant documents if the model understands the case concept.
2 Hybrid search combining keyword and semantic passes outperforms pure semantic retrieval by 15 to 30 percent on complex document sets.
3 A high recall score on semantic benchmarks means your review protocol covered the collection.
Abstract visualization of a digital document retrieval system: rays of blue light scanning through a dark archive of floating documents, with some documents containing alphanumeric codes and contract identifiers

Quick Answer

Yes. AI semantic search misses documents whose relevance depends on exact identifiers, account numbers, internal code names, and proprietary abbreviations. These are terms the model cannot semantically anchor because no training context established their meaning. Research confirms hybrid search combining semantic and keyword passes outperforms pure semantic retrieval by 15 to 30 percent on complex document sets, and it is the current defensibility standard for high-stakes review.

Did this answer your question?

The shadow falls at the moment you trust the search. There it is, the grey circle of retrieved documents, confident, complete-seeming, rushing toward you through a screen. The algorithm has found the meaning. It has understood your query. And somewhere, behind the retrieved set, in the darkness where no vector points, a contract identifier sits untouched. An account number, known only to the parties who created it. A project code name, invented in a conference room six years ago and never mentioned in any training corpus. These are the documents that matter. And the search has already moved on without them.

I have watched this happen in matters where the stakes were high enough that a single missed document would change the outcome. The AI understands the concept. It retrieves documents about the subject, the relationship, the transaction. But the critical email, the one referencing the deal by the internal name only insiders used, is not in the results. It exists. It is responsive. It is invisible to the model.

This is not a failure of the technology in the sense of a bug to be patched. It is the structural consequence of how semantic search works: encoding meaning into mathematical space where exact strings and arbitrary identifiers have no natural location. As one AI engineer put it, "dense vectors generalize meaning, and that same strength is their weakness: they lose exact-match precision." The gap is predictable. It is measurable. And in legal document review, where courts demand that search methodologies be defensible and opposing counsel can challenge every omission, that gap is a liability.

The answer is not to abandon AI search. It is to use it correctly, which means understanding what it cannot see and building a hybrid protocol that catches what pure semantic retrieval will always miss.

According to Applied AI Research (2025), hybrid retrieval combining keyword and semantic passes outperforms pure semantic search by 15 to 30 percent on complex document sets. In legal e-discovery, where material documents often turn on exact identifiers, account numbers, and internal code names, that gap is not an abstract benchmark finding. It is the difference between finding the document that decides the case and missing it entirely.

I have worked across litigation matters where semantic-only review produced results that looked complete and were not. The AI retrieved documents that understood the subject. It missed the ones identified by the specific account number, the particular contract code, the internal project name that appeared in every communication between the parties but in no training corpus. From what I have seen, the categories most reliably missed by pure semantic search are those whose relevance is encoded in exact strings rather than natural language: account and policy numbers, contract identifiers, legacy system codes, and proprietary abbreviations the model has no context for.

The question for any attorney using AI search in document review is not whether semantic search misses these documents. Research confirms that it does, and the mechanism is structural, not accidental. The question is whether your review protocol is designed to catch them before opposing counsel finds them in production.

What Will Matter Most in E-Discovery Search Over the Next 12 to 24 Months?

Three forces are converging, and each one makes the hybrid-search gap more consequential.

Judicial Scrutiny of AI Search Methodology Is Increasing

Courts have begun asking, explicitly, how AI-assisted review was conducted. The question is no longer whether you used technology-assisted review but what your search protocol included, what it excluded, and how you validated recall. A semantic-only protocol, one that produces no keyword log, no custodian-term matrix, no explicit string-match audit trail, is becoming harder to defend when challenged. The analysis of what judges expect from AI document review in 2026 confirms this shift is already underway. By 2027, I expect that AI search transparency will be as standard a meet-and-confer topic as custodian identification is today.

Document Volumes Are Growing, and Identifier Density Is Growing With Them

Modern litigation document sets contain proportionally more structured data - more database exports, more CRM records, more financial-system outputs - than they did a decade ago. These records are dense with exact identifiers by design. Account numbers, transaction codes, policy identifiers, order numbers: these are the primary keys of business systems, and they carry meaning that a semantic model cannot encode from context alone. As document volumes grow, the identifier-dense fraction of any production set grows with it. The semantic-only blind spot does not shrink at scale. It expands.

The Industry Has Acknowledged the Limit, but Practice Has Not Caught Up

As Paolo Perrone of The AI Engineer wrote in 2026, "every serious 2026 search system runs semantic alongside BM25, fuses the rankings with Reciprocal Rank Fusion, and reranks the top 100 results with a cross-encoder." That is the state of the art in information retrieval. What lags is e-discovery product marketing, which still frequently positions AI search as comprehensive without specifying what it cannot retrieve. The gap between what platforms claim and what defensible review requires is where disputes will emerge. The firms that build hybrid protocols now, before a judge asks why they did not, will be better positioned than those that learn the lesson from a sanctions order.

The Standard Is Shifting Toward Documented Hybrid Protocols

What constitutes a defensible search protocol is not static. It is set by the cases, the commentary, and the practice guides that define what reasonable care looks like at a given moment. Right now, that standard is moving toward requiring documentation of both the semantic and keyword components of a search, a validation sample, and a methodology statement that can be produced in response to a challenge. Semantic search alone, undocumented and unvalidated, will not meet that standard. The firms that understand this now, and build the protocol accordingly, will not have to learn it from a sanctions order.

How Does Semantic Search Work, and Where Does It Break Down?

Semantic search is a retrieval method that finds content based on meaning and intent rather than exact word matches.

A query and a document are each translated into a vector, a coordinate in a high-dimensional mathematical space, and retrieval works by finding the document vectors closest to the query vector. The model underlying this process has learned, from training on vast text corpora, that "contract" and "agreement" occupy similar space, that "acquisition" and "merger" cluster together, that "plaintiff" and "claimant" are conceptually proximate. For queries that reflect those learned relationships, semantic search is genuinely powerful.

The breakdown is structural. As Sebastian Gingter of Thinktecture observed after running semantic search across a business document corpus: "Your documents are usually a collection of facts, statements, numbers, and so on in your field of expertise. Questions about these documents and the documents themselves are, semantically, not very similar." The model retrieves documents conceptually close to the query. It does not retrieve documents that happen to contain a specific string the model has no anchor for.

The categories where semantic search most reliably fails in legal document review are those whose relevance is encoded in exact strings rather than natural language. From what I have seen across litigation matters, these break into five recurring types:

  • Internal project and deal code names invented within an organization that have no public semantic context
  • Legacy account and policy numbers from financial, insurance, and healthcare systems
  • Contract and order identifiers in alphanumeric formats specific to a company's internal systems
  • Regulatory filing numbers and docket identifiers in formats not present in general-purpose training data
  • Proprietary product model numbers and internal SKU codes

None of these are obscure document types. They are standard in commercial litigation, securities matters, insurance disputes, and healthcare fraud investigations. They are also exactly the types that LlamaIndex singles out when noting that keyword search "remains effective for exact-match retrieval scenarios, such as searching for a specific product code, legal citation, or proper noun" - the inference being that semantic search, on its own, does not.

The practical result is a recall profile that is uneven in a specific, predictable direction. Conceptual queries retrieve well. Identifier-dependent queries do not. An aggregate recall metric that does not account for identifier-heavy document types in proportion to their presence in the collection will overstate actual recall for the matters where it most needs to be accurate. This is the mechanism behind the gap. Understanding it is the first step toward closing it.

What Does the Identifier Gap Look Like in a Real Document Set?

The gap is not uniform. It concentrates where identifiers concentrate. In a commercial dispute involving a single contract between well-known counterparties, semantic search may perform adequately because the relevant documents reference the parties and the transaction in natural language. In a matter involving hundreds of thousands of financial system records, each identified by alphanumeric codes that carry no semantic meaning outside the organization that created them, the gap can be severe.

I want to be precise about the mechanism because it matters for understanding what hybrid search fixes. Semantic search does not fail because the model is poor. It fails because it is doing the right thing for the wrong query type. When a model retrieves documents conceptually similar to a query, it is succeeding at what it was designed to do. The problem is that document review is not only a semantic task. It is partly an exact-string retrieval task, and for that subset of the work, vector similarity is the wrong tool.

Atlan's analysis of semantic search behavior captures this precisely: "a query for 'Q4 2024 revenue ARR' needs lexical matching on those exact terms, not conceptual similarity, which is why production implementations require hybrid search." That framing applies directly to legal document review. A matter that turns on a specific account number, a specific contract identifier, a specific filing code, is a matter where conceptual similarity will retrieve the wrong set and exact-string matching will retrieve the right one.

The recall comparison between semantic-only and hybrid search on identifier-dependent document types follows a consistent pattern:

Query Type Semantic-Only Recall Hybrid Recall Gap
Conceptual natural language (topic, theme, concept) High High Minimal
Named entity (well-known person, company, place) Moderate to High High Small
Internal code names and project nicknames Low High Large
Alphanumeric account and contract identifiers Very Low High Very Large
Legacy system codes and proprietary abbreviations Very Low High Very Large

This table describes a structural property of vector-based retrieval, not a judgment on AI search technology broadly. Semantic search is genuinely superior for conceptual queries. BEIR benchmark evaluations confirm that dense embedding models outperform keyword search on natural-language question-answering, while BM25-based keyword search outperforms embedding APIs on certain specialized corpora, including biomedical literature and identifier-heavy records. No single method wins across all document types. Applied AI Research (2025) found that combining them improved recall by 15 to 30 percent over pure semantic retrieval.

What makes this a defensibility problem is that the missed documents are not random. They are the documents whose relevance is encoded in exact identifiers: the specific account, the specific contract, the specific project code. As Relevant Discovery builds into its review process, every answer must trace back to the exact exhibit it came from - but that linkage only holds if the exhibit entered the reviewed set in the first place. The shadow of its absence is not discovered until opposing counsel introduces it at deposition, and by then the procedural consequences are already in motion.

Two overlapping search circles representing keyword and semantic retrieval passes, their union glowing white against a dark digital background. Documents marked with alphanumeric contract codes fall only within the union

Does Your Document Review Search Both Ways?

Relevant Discovery combines keyword and semantic search in a single defensible workflow, with a logged keyword pass, union-of-results merging, and every answer traced to its source exhibit. See how it works on your actual collection.

Questions this article answers

  • Can AI semantic search miss important documents in e-discovery document review?
  • What specific document types does pure semantic search most reliably fail to retrieve?
  • Why is hybrid keyword-plus-semantic search the defensibility standard rather than an optional enhancement?

Why Is Hybrid Search the Mandatory Standard, Not a Premium Option?

Hybrid search is the architectural pattern that runs a keyword retrieval pass and a semantic retrieval pass in parallel, merges the result sets by union, and reranks the combined pool by relevance signal.

The keyword pass, typically BM25, retrieves documents containing exact query terms. The semantic pass retrieves documents conceptually similar to the query. The union contains documents that either pass matched, as of .

The case for treating this as a standard rather than an enhancement comes from the structure of the gap it closes. When a matter turns on an exact identifier, the semantic pass will systematically under-retrieve because embedding vectors have no anchor for that exact string. The keyword pass retrieves precisely the documents the semantic pass missed. The union closes the gap. There is no version of a sound retrieval protocol that achieves adequate recall on identifier-heavy matters without this union.

Paolo Perrone, writing in The AI Engineer, identified Reciprocal Rank Fusion (RRF) as the standard fusion method for production hybrid search: "every serious 2026 search system runs semantic alongside BM25, fuses the rankings with Reciprocal Rank Fusion, and reranks the top 100 results with a cross-encoder." That is not a description of a premium feature. It is a description of how the retrieval problem is solved at production scale.

The defensibility dimension adds a second reason hybrid search is not optional in legal document review. A semantic-only protocol produces no keyword log. When opposing counsel challenges the search methodology, a keyword log, the list of exact terms searched and the documents they returned, is the primary response. A protocol that cannot produce this log cannot withstand a standard search-term challenge, regardless of its overall recall performance on conceptual queries.

The Vespa.ai research team identified a further structural limitation of pure vector approaches in legal search: querying "force majeure" AND "(pandemic OR epidemic)" fails with pure vector databases that cannot execute boolean logic. Legal keyword searches frequently depend on boolean operators, term proximity, and field-specific constraints that vector similarity cannot model. The hybrid protocol handles these natively through the keyword pass.

A documented hybrid protocol for defensible AI-assisted review requires four elements, in my judgment:

  1. A parallel keyword pass alongside the semantic pass, with all keyword terms logged and preserved
  2. Union-of-results merging so that documents retrieved by either pass enter the reviewed pool
  3. RRF-style reranking of the combined pool to prioritize the highest-signal documents for attorney review
  4. A recall-validation sample drawn from the non-retrieved set to estimate actual recall, not only precision

The shadow of unclosed gaps, documents that do not enter the reviewed pool because no hybrid pass retrieved them, is not visible in precision metrics. Precision measures what you reviewed. Recall measures what you missed. A semantic-only protocol optimized for precision can show strong performance numbers while leaving an identifiable category of relevant documents untouched, in that grey circle of the un-retrieved. That circle narrows only with the union. The darkness behind the result set is not an AI limitation. It is a protocol choice.

The grey circle of retrieved documents is not the collection. It is the portion of the collection that the retrieval protocol reached. In that distinction lies the practical risk of semantic-only search in legal document review: the gap is not random, not uniformly distributed, not noise. It clusters on exact identifiers, account numbers, internal codes, proprietary abbreviations, the categories that define the documents whose relevance is most precisely encoded.

From what I have seen in practice, the practitioners most exposed to identifier-gap risk are those who adopted AI-assisted review early and stopped at semantic search because it seemed sufficient. The gap only becomes visible when opposing counsel introduces the missing exhibit, or when a sampling exercise reveals a recall shortfall the aggregate metrics concealed. By then, the protocol choice that created the gap is already a matter of record.

Hybrid search closes this. The union of a keyword pass and a semantic pass retrieves what either method alone misses. The keyword log produced by that union is the documentary foundation of a defensible search methodology. In 2026, these are not advanced features. They are the architecture that every serious production system already runs. Predictive coding is not the only defensible AI review path, and semantic-only is not a defensible alternative to hybrid. The shadow behind the result set narrows only when both passes run. Run both.

Written by

Michael

Kansky

Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

Can AI semantic search miss important documents in e-discovery?

Yes. Semantic search retrieves documents by conceptual similarity to the query. Documents whose relevance depends on exact identifiers, such as internal project codes, account numbers, and contract identifiers, are often not retrieved because those strings have no conceptual anchor in the embedding model's training. Applied AI Research (2025) found hybrid retrieval outperforms pure semantic search by 15 to 30 percent on complex document sets.

What types of documents does semantic search most reliably miss?

The five categories most consistently under-retrieved by semantic-only search are: internal project and deal code names, legacy account and policy numbers, contract and order identifiers in proprietary formats, regulatory filing numbers and docket identifiers, and proprietary product model numbers and SKU codes. These are standard document types in commercial litigation, securities matters, and insurance disputes.

What is hybrid search in the context of document review?

Hybrid search runs a BM25 keyword retrieval pass and a semantic vector retrieval pass in parallel, merges the results by union, and reranks the combined pool using Reciprocal Rank Fusion. The keyword pass retrieves documents containing exact query terms. The semantic pass retrieves conceptually similar documents. The union captures documents that either method alone would miss.

Is a keyword log required for defensible AI-assisted review?

A keyword log documents the exact terms searched and the documents returned by the keyword pass. When opposing counsel challenges search methodology, the keyword log is the primary response. A semantic-only protocol produces no keyword log and cannot satisfy a standard search-term challenge, regardless of its conceptual recall performance.

How does Relevant Discovery address the semantic search gap?

Relevant Discovery's review workflow combines keyword and semantic search in a single documented protocol, with a logged keyword pass, union-of-results merging, and every answer traced back to the exact source exhibit. This architecture closes the identifier gap while producing the keyword documentation a defensible review requires.

Read next

A clean, professional visual concept-style hero image showing two pricing paths diverging at a crossroads. On the left path, Per-Matter, a single briefcase or legal folder icon with a $195 price tagE discovery

Per-Matter vs Annual Platform: What a Small Firm Should Pay

If you typically have fewer than three matters actively running in your e-discovery platform at the same time, per-matter pricing will almost always cost you less than an annual subscription.September 21, 202628 min read
Judge reviewing AI document review workflow documentation in a federal courtroom settingE discovery

What Judges Expect From AI Document Review in 2026

In 2026, judges expect a documented, reproducible AI review workflow supervised by counsel. As practitioners noted following the Schulte v. LinkedIn ruling (N.D. Cal. 2026), AI-assisted review is already settled law. Courts are no longer asking whether you used AI.September 20, 202627 min read
Private AI deployment for e-discovery: dedicated single-tenant infrastructure versus shared multi-tenant cloudE discovery

Why 91% of E-Discovery Buyers Now Want Private AI

Private AI deployment keeps your client's privileged documents inside a dedicated single-tenant environment where no shared inference infrastructure can touch them.September 18, 202625 min read

See it on your matter

Bring us a messy collection - mailboxes, scans, phones, recordings - and watch it become one searchable, defensible record.