Home › Insights › Why Searching Outlook Isn't a Defensible Document Review

Why Searching Outlook Isn't a Defensible Document Review

Laptop showing an email inbox search beside a stack of printed emails and an external hard drive on a small law office desk

Quick Answer

No. Searching Outlook or Gmail directly is not a defensible review: it skips deleted items, other custodians, de-duplication, threading and an audit log, and an export can alter metadata.

It is not the cheap option either. Every duplicate and broken thread a mailbox search leaves behind gets read by hand. The lower bill comes from a documented collection, de-duplicated, threaded and culled before review begins, with a record that answers the question opposing counsel tends to ask last.

Did this answer your question?

Key Points

  • In an August 2026 r/paralegal thread, one commenter advised: "For discovery, I would avoid normal Outlook search entirely," because it leaves no record of how a set was collected.
  • On a mailbox under a Purview hold, anything the custodian moves or deletes is kept in Recoverable Items , outside the ordinary folders, and is retrieved through Purview rather than Outlook search.
  • In a test described on the Emerging Litigation Podcast in 2024, generative AI found 96% of relevant documents against 54% for human reviewers, with precision of 60% against 91%.
Three things small litigation teams believe about mailbox search. Myth or fact?
Call each one, then see how other readers called it.
1 Outlook's own search is good enough for discovery if the search terms are chosen carefully.
2 Purview eDiscovery keeps a record of the searches you ran on a mailbox.
3 Once email sits in Purview, every attachment in it can be searched.
Laptop showing an email inbox search beside a stack of printed emails and an external hard drive on a small law office desk

A mailbox search finds what one person can see. Discovery asks for the rest.

Shrink the set before anyone reads it: de-duplicate, thread and cull a documented collection, then rank what remains, so attorneys spend paid hours only on documents likely to matter.

The difficulty is where the set comes from. The American Bar Association's Law Technology Today has described "permissions debt": access that piles up as attorneys join matter teams, SharePoint rights that outlive the matter, Teams channels left open after the project ends. A collection drawn from that sprawl arrives large and repetitive. Searched mailbox by mailbox in Outlook, it leaves every duplicate and every broken thread to be read by hand, at manual review rates.

My own view is that the cheapest document in any review is the one nobody has to open.

Automation lowers the bill without abolishing it. A classifier that attorneys trust enough can let some documents be produced without an attorney reading them, a point a data scientist made in 2024. Generative AI moves the cost forward instead. One practitioner put the senior attorney time spent writing review prompts at 40, 50, even 60 hours, all of it spent before the first document is coded. The same practitioner called the "maybe relevant" pile the place where cost bloats.

A note on the evidence: nothing gathered for this article counts what native search finds against what a review platform finds, and no measured case shows metadata changing on export. Those figures would settle the argument in a sentence. Until someone collects them, the case rests on how practitioners describe the work.

The description begins with the search box itself, and with the narrow thing it was built to find.

Searching Outlook for discovery documents is not a defensible review. It is a search of one place, and discovery rarely lives in one place. A 2026 article in the American Bar Association's Law Technology Today listed what Microsoft Copilot can reach for a single user. The list ran to email "going back years," Teams messages and meeting recordings, OneDrive files, and SharePoint libraries across every site that user can see. Outlook's search box was built for the first item on that list.

The same article cited Wolters Kluwer's 2026 Future Ready Lawyer Survey, which found that over 90% of legal professionals use at least one AI tool daily. Lawyers are not short of software. A second 2026 piece in the same publication said the firms getting real value from AI had made verification a documented workflow step, with a named reviewer and a checklist. A defensible review is one whose collection, culling and coding can be explained afterward, step by step, to a party who would prefer it had gone badly.

Set that Copilot list beside an Outlook search and ask how many of those places the search actually opened. The results look complete. Deleted items, other custodians, duplicates and threads simply never arrive, and an export can change the metadata it was supposed to carry. Even the newer AI review tools, which judge each document alone against one set of instructions, inherit whatever the collection quietly left behind.

Was the Mailbox Search Ever Going to Be Enough?

Search finds email. It does not preserve it: holds belong to a case, not a query, and one administrator found a search alone "not good enough due to the legal structure." Generative AI shares a blind spot, which one practitioner called "no cross document intelligence." Recall is certified by sampling, never assumed. Bring the collection to Relevant e-Discovery before the first document is coded.

Review My Collection

Can You Just Search Outlook for Discovery Documents?

No. Outlook search is built to retrieve a remembered email, not to support a review, and the manual sorting it leaves behind is priced like manual review: roughly $19K per gigabyte.

A defensible review is one whose scope, method and results can be explained, step by step, to opposing counsel or a judge. Before relying on any mailbox search, I would ask three questions:

  1. Can you show in writing what was searched, when, and by whom?
  2. Does the set cover every custodian and every held item, or only one person's mailbox?
  3. Were duplicates and threads handled, so the same message is coded once and coded the same way?

If any answer is no, the search found email. It did not review it.

A 2025 thread on r/ediscovery shows how quietly this goes wrong. The poster, an attorney with 20 years of practice, explained that the organization exported the entire result set as a PST and left the attorneys to sort it in Outlook, "which is obviously not a review tool." The IT person then said Purview could not run a single search for emails meeting either of two conditions. What was offered instead was two separate searches, with "potentially enormous overlap." One commenter asked the only question that mattered: "If that person is ever deposed as a records custodian what are they going to say?"

Nothing in that exchange was dramatic. A search was refused, then split in two, then handed to an email client. Combining 4 sources points to one pattern: the expensive part of email discovery is seldom finding the messages, it is proving later how they were found.

The common assumption is that a search which returns the right messages has done the work of review. It has not. It has produced a pile, and a pile keeps no memory of how it was made. A kitchen drawer behaves the same way. Finding the receipt you wanted says nothing about what else was in the drawer, or who emptied it last week.

Our platform starts from the opposite end. Originals are immutable and content-hashed, audit trails are append-only, chain of custody is documented, and a fail-closed privilege gate sits in front of production, so the process holds up if opposing counsel challenges it. Enterprise platforms supply that spine; newer AI tools supply cited answers and chronologies. Relevant e-Discovery was built to be both at once, which matters most to firms that have learned a small case is rarely a small review job.

The harder question, which the r/ediscovery thread never quite reached, is what a mailbox search leaves out even when it runs exactly as intended, and Microsoft's own tooling turns out to have a few answers of its own.

What Does a Native Mailbox Search Leave Out?

A native search leaves out held and deleted items, other custodians, and any trail back to source; a proper review links every answer to the exact exhibit it came from.

A native search is a query run inside the email client itself, Outlook or Gmail, against whatever that client happens to display. Three gaps open at once:

  • Moved and deleted items. On a mailbox under hold, anything the user moves or deletes is kept in Recoverable Items, outside the ordinary folders, and is retrieved through Purview.
  • Other custodians. A mailbox search covers one mailbox. Discovery requests rarely stop there.
  • Preservation. Finding an email does not preserve it.

An IT administrator posting to r/ediscovery in May 2025 ran into all three. The organization received such requests about 4-5 times a year, and requestors wanted "any and all mailboxes and sites held that contain the name, email or phrase." A search alone, the poster wrote, had been tried and was "not good enough due to the legal structure." Then the Purview screen allowed only 100 users to be selected for a hold, although Microsoft's documentation, as the poster read it, allowed 2,000 user mailboxes and 2,000 sites in a single hold.

Contrary to how the screen presented it, 100 was not a ceiling. A commenter explained that it was an auto-populated preview, and that more users could be added by searching for them by hand or through PowerShell. Another warned of "severe sanctions in U.S. Federal and State courts" when a hold goes wrong. A third thought Recoverable Items storage was capped at 110GB, enough for a long-tenured custodian to fill.

None of this is exotic. It is the ordinary weather of email discovery. A screen that misstates its own limits will not warn you about what it skipped.

One commenter even offered an untested PowerShell script for bulk holds. Improvised tooling usually begins this way, and home-built review tools tend to break just as quietly, in places nobody thought to check.

Our answer to the source problem is structural. Every output links back to the exact exhibit it came from, so an attorney can verify it with one click before filing, which is the direct answer to the sanctions risk seen in Mata v. Avianca. Where that work runs matters as much. We deploy inside the client's own AWS account, under the client's own keys, a trust posture no consumer AI tool and almost no small-firm tool can match.

The fair test is not a brochure. Relevant e-Discovery asks buyers to bring a messy collection and a hard question, then watch the system read, code and cite their own evidence.

And even a hold that works leaves a second problem in place: the same email, kept in several mailboxes, arrives in the review once for every mailbox that kept it.

Where does an Outlook review actually break down?

We followed what attorneys, paralegals and IT administrators say they do with mailbox search results. The search box turned out to be the smaller problem.

In August 2025, an attorney with 20 years of practice described an odd arrangement on a forum for eDiscovery practitioners. Their organization licenses Microsoft's eDiscovery tooling, Purview. Even so, the standing practice was to export every result set as a PST file and leave the attorneys to sort through it somewhere else, "like…Outlook, which is obviously not a review tool."

Others use the same workaround. A year earlier, an administrator at a managed service provider had laid out the routine in more detail: take the PST, load it into a shared "Discovery Mailbox," and grant reviewers permission to read it in Outlook. "While it's not the easiest process, it's totally doable," the administrator wrote, though a partnering law firm clearly had "more streamlined tools."

The mail gets found. The trouble starts after the export.

Stage of the workaroundWhat practitioners report
Search in PurviewKeeps a record of the searches run, per a paralegal forum commenter. Another says it does not highlight search-term hits, and a third that content that is not OCR'd or sits in a zip file stays unsearchable.
Two searches instead of one"Potentially enormous overlap" between the result sets, all of which would need review.
Export to PSTWorkable for a small one-off matter, one commenter says, if the original export stays untouched and every search parameter is written down.
Searching in Outlook itself"Fine for finding the email you vaguely remember, not great for defensible searching."

The commenter behind that last row, posting in a paralegal forum in August 2026, singled out one Purview feature. Purview gives you "a record of the searches you ran," the commenter wrote, and "that last part matters if someone later asks how the set was collected." Ordinary Outlook search leaves nothing comparable behind. "The main thing is not relying on whatever Outlook happens to show today."

Someone may well ask. A review guide published on JD Supra by an eDiscovery software vendor notes that under Federal Rule of Civil Procedure 30, the opposing party can seek discovery about how a party collected, processed, reviewed and produced its documents. A commenter in the attorney's thread pointed out that the organization's IT person had first called the needed search impossible, then asked:

"If that person is ever deposed as a records custodian what are they going to say?"

Commenter on an eDiscovery practitioner forum, August 2025

The record matters for privilege as well. Federal Rule of Evidence 502(b) protects an inadvertent disclosure only where the holder "took reasonable steps to prevent disclosure." A party can show those steps far more easily if it wrote them down.

A search does not preserve anything

A second gap sits upstream of the review. An administrator handling broad requests through Purview wrote in May 2025 that "a search is usually not good enough due to the legal structure." A commenter explained what a hold adds: the custodian keeps working normally, but "anything they move or delete is retained in recoverable items" and can be pulled back through Purview. Without the hold, a search sees only what is still there to find.

The same commenter warned of "severe sanctions in U.S. Federal and State courts" when holds go wrong. The rule itself is narrower than the warning. Rule 37(e) applies when information that should have been preserved is lost because a party "failed to take reasonable steps to preserve it" and cannot be restored or replaced. The harshest measures require a finding of intent to deprive the other side.

Overlap, and the limits of the better tool

Told to run two searches where one would do, the attorney worried that every email in the first set might also sit in the second. Another commenter noted that a custodian's signature block alone would pull messages into both. The thread's remedy: load both exports into a dedicated review platform and deduplicate across the whole collection. The vendor guide on JD Supra describes the companion controls: "propagating tagging to duplicate documents and batching families of documents such as threaded emails and attachments."

None of this makes Purview the finish line. Besides the limits in the table, one commenter cautioned that a query working in August 2025 might not work a year later, after a Microsoft update.

Read side by side, the accounts put the risk in the same place. The teams in these threads collected with tools that can log a search and reach held items. They gave those advantages up at the export, when the results landed in a mailbox that cannot reconcile duplicates or show how the set was built. The search was rarely the step they struggled to defend.

Before anyone opens a result

  • Write the search log first: terms, date ranges, custodians, exclusions, and an untouched original export.
  • Ask who ran the search and whether that person could explain the query under oath.
  • Ask whether a hold was on each custodian's mailbox before anyone searched it, and who confirmed that with counsel.
  • If you ran several searches or added a custodian, ask how duplicates are reconciled and whether coding carries to every copy and every message in the thread.
  • Ask what the tool could not read, scanned attachments and zip files in particular, and how those items will be reviewed.

How we checked this

We read four practitioner forum threads from 2024 to 2026, one vendor-written review guide, and the text of two federal rules. No figures in this section come from our own data. The forum posts are individual accounts from anonymous attorneys, paralegals and administrators. They show how some teams work and cannot tell us how common the practice is, and several claims about Purview's features are one commenter's experience, which we did not test. The review guide was written by a company that sells review software, and so do we: our platform keeps append-only audit trails, so we have a stake in arguing that the record matters. Still unknown: how often courts actually probe a mailbox-based review, and how Purview's higher-tier review features compare, since no source here had used them fully.

  1. Reddit, r/ediscovery, thread on search and review limitations, August 12, 2025.
  2. Reddit, r/msp, thread on internal review methods, June 19, 2024.
  3. Reddit, r/paralegal, thread on Outlook search for discovery, August 30, 2026.
  4. Reddit, r/ediscovery, thread on Purview holds, May 14, 2025.
  5. JD Supra, vendor guide to document review searches, April 25, 2022. Written by an eDiscovery software company.
  6. Legal Information Institute, Federal Rule of Civil Procedure 37, retrieved October 10, 2026.
  7. Legal Information Institute, Federal Rule of Evidence 502, retrieved October 10, 2026.

What Turns an Email Search Into a Defensible Review?

Move the review into a platform that de-duplicates, threads and logs every step, and keep processing single-tenant or inside your own AWS account, with no vendor retention or model training.

Three steps separate a pile of search hits from a review that can be explained:

  1. Collect every relevant custodian, held items included, into one matter instead of one mailbox at a time.
  2. De-duplicate and thread the set before anyone codes a document.
  3. Record the terms, dates, custodians and exclusions, then validate the results by sampling.

The vocabulary for this already exists. On the Emerging Litigation Podcast in 2024, data scientist Lenora Gray described de-duplication as replacing identical copies with one representative copy, culling as cutting a collection by date range, file type or keyword, and de-NISTing as removing system files. She also explained technology-assisted review (TAR): attorneys code a small set of documents, and an algorithm builds a statistical model, called a classifier, that predicts coding for the rest and attaches a confidence level to each prediction.

Certification is the part a court cares about. Gray explained in 2024 that TAR was usually certified to the other side by random sampling and statistical estimation, in a sentence such as "We recovered at least 80% of the relevant documents from this collection." In 2024, Gray said, e-discovery standards usually required around 80% recall.

The same interview described a head-to-head test, run with a single prompt and no matter-specific tuning, that scored contract attorneys and generative AI against senior-attorney coding. In that 2024 account, the AI found 96% of the relevant documents against 54% for the humans, while its precision was 60% against 91%.

Recall measures what you found. Precision measures what you paid to read. An Outlook search produces neither number.

One review-software vendor makes the procedural point bluntly in its own marketing: under FRCP Rule 30, the opposing party can seek discovery about a party's collection, processing, review and production workflows, and a review platform should propagate tags to duplicates and batch threaded emails with their attachments. The same guide points to the conference that Rule 26(f) requires, where the parties should negotiate search terms and settle preservation, production and privilege questions before the review begins.

Where the processing happens is the last question, and often the first one a client asks. Our processing runs single-tenant, or in your own AWS account under your own keys, so nothing you feed it trains anyone's model or waives privilege. We built it for the solo or small firm with one messy matter and no administrator, the segment Relativity and Everlaw price out, which is why per-matter and annual platform pricing deserve a careful comparison before anyone signs.

On a 2026 podcast, a forensic technology director whose firm sells AI review services asked whether the lawyers understood a method well enough to put it in an affidavit, and a defensible review is one that passes that test. It is a recorded process that collects, de-duplicates, threads and samples, and that someone can explain under oath.

Toward documented, multi-source collections reviewed in dedicated tools, where a mailbox search with no written record of terms, dates and custodians behind it grows steadily harder to defend.

In 2019 an Office 365 administrator wrote that every discovery request they received was answered through Content Search alone. By 2026, practitioners advising on Outlook mail were telling colleagues to avoid normal Outlook search for discovery, keep the original export untouched, and write down search terms, date ranges, custodians and exclusions. My reading is simple. The search has not changed much. The paperwork around it has.

The American Bar Association's Law Technology Today made a parallel point about Copilot: knowing what the tool can see, and deciding about that scope deliberately, puts a firm in a "defensible position under Model Rule 1.6." A collection works the same way. Deleted items, other custodians, duplicates, threads and the log either sit inside the record or sit outside it, and an export that alters metadata moves things the wrong way.

A firm that tries a review tool the way a 2026 Law Technology Today article suggested, as a 30-day pilot on one recurring, low-stakes workflow with an owner and a review step, learns its time saved and corrections required before a real deadline arrives. On a 2026 law-firm podcast, Therese Caparo named the three places generative AI now enters document review: surfacing categories of relevant documents early, first-level tagging, and second-level checks on coding consistency and gaps.

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

What Else Do People Ask About Searching Outlook for Discovery?

Most follow-up questions concern exports, Purview, cost and AI review, and each answer returns to one point: a defensible set needs a record of how it was built.

Does exporting a mailbox to PST change its metadata?

It can, and the evidence gathered for this article does not measure how often. A PST is Outlook's own data file, built for moving mail rather than proving where it came from. I would treat the first export as evidence, leave it alone, and run every search on a copy.

Is Microsoft Purview enough on its own?

For holding and collecting Microsoft 365 mail, Purview is the usual starting point. For review it is thinner. Administrators in a 2025 forum thread described real grief with its redesign and a hold screen that did not always behave. A hold is not a review.

What is technology-assisted review?

Technology-assisted review (TAR) is supervised machine learning applied to documents. On the Emerging Litigation Podcast in 2024, it was described this way: one or a few attorneys code a small set, and the software builds a statistical model, a classifier, that predicts coding for the rest and gives each prediction a confidence level. Documents can then be ranked so the likely relevant ones are read first.

How do you test a generative AI review before trusting it?

Start small. One practitioner on a 2026 law-firm podcast suggested a 100-document test set: about 30 documents known to be relevant, about 30 known not to be, and 40 drawn at random. The prompt then moves to 1,000 documents, and only after that to the full population.

What does AI-assisted review cost compared with manual review?

At Relevant e-Discovery, AI-assisted review runs at cents per document, against the dollars per document that manual review costs.

The difference is meant to be visible on the buyer's own matter, and the Relevant e-Discovery contact page is where that comparison starts.

Written by

Michael

Kansky

Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.

Connect on LinkedIn

Read next

A very long itemized paper paperwork unspooling across a small law firmEdiscovery

What AI Coding 100,000 Documents Costs in Your Own AWS Account

AI coding 100,000 documents in your own AWS account has no single price. It splits into a metered cloud bill, a vendor license and human validation hours.October 9, 202626 min read
Attorney comparing two printed document review reports side by side at a desk beside a laptop and boxes of case filesEdiscovery

4 Records That Make an Agentic AI Review Reproducible

An AI document review is reproducible for court when four records survive each run: the instructions with every revision, the model version, the exact retrieval set and a fixed random seed.October 9, 202628 min read
Photorealistic editorial photo of a small law office conference table in late afternoon light: a printed software services agreement lies open with a pen resting on a marked-up clause, beside a neat stack of looseEdiscovery

Your AI Review Vendor Can Swap Models Mid-Matter. Check Your Contract

Unless your AI review contract pins a model version per matter and requires advance notice of updates, the vendor may swap models mid-matter, and March's coding may not reproduce in June.October 7, 202629 min read

See it on your matter

Bring us a messy collection - mailboxes, scans, phones, recordings - and watch it become one searchable, defensible record.