Home › Insights › What Share of a Chat Collection Is Actually Relevant?

What Share of a Chat Collection Is Actually Relevant?

Laptop showing a long workplace chat thread beside a tall stack of printed pages with only a few flagged for review

Key Points

  • Microsoft Purview eDiscovery (Premium) cannot isolate a single 1:1 Teams chat at search, because Conversation ID, Thread ID and Conversation Type are not supported search properties.
  • According to Greg Buckles of eDiscovery Journal, the new Microsoft Purview eDiscovery portal collects chat in 12-hour chunks, so unrelated conversation segments ride along into review.
  • ComplexDiscovery projects review's share of eDiscovery task spend at 52 percent for 2030, down from a reconciled 62 percent in 2025, even as business chat keeps growing.
Three things litigators believe about chat collections. Myth or fact?
Call each one, then see how other readers called it.
1 A careful Purview search can pull only the one Teams conversation a matter concerns.
2 The duty to preserve chat can arise before any complaint is filed.
3 Purging a Teams search removes the messages from every user's view.
Laptop showing a long workplace chat thread beside a tall stack of printed pages with only a few flagged for review

A chat collection's size reflects how it was gathered; the flagged few are what the matter requires.

Quick Answer

Nobody has published a measured figure. A chat collection's responsive share refers to the fraction of gathered messages that a request for production actually calls for, and the only defensible number is one measured on your own matter. Collection size is set by tooling, not by the dispute. Microsoft Purview eDiscovery (Premium) cannot isolate a single 1:1 Teams chat at search, and according to Microsoft Learn, that same search also fixes what a purge deletes. I'd price review on the scoped set, never the raw count.

Did this answer your question?

I shall state the conclusion first, as the evidence obliges me to: the number of messages in a chat collection records how the collection was gathered, and it tells a litigator remarkably little about how much of it a request for production will require. A chat collection is the body of Teams, Slack or similar messages preserved and gathered for a matter. Its responsive share is the fraction that a request actually calls for.

That share has not been published for chat. None of the sources cited in this article measures it, and I will not pretend otherwise. What the sources do show is how the collected count is formed. According to Microsoft Learn, Teams stores chat in two separate copies, one kept for compliance and used by eDiscovery, the other kept for the user. The compliance copy is what a search reaches, what a hold keeps and what a review set inherits.

It is worthy of notice that nothing in that arrangement is sized to the matter. The search mechanics decide the count; the dispute does not. A per-message quote applied to that count therefore prices the mechanics.

This article will cover three matters:

  • why chat collections come back larger than the matter that occasioned them;
  • why per-message pricing on the raw set overstates the review burden;
  • how to measure the responsive share on your own matter, and how to defend a narrowed scope.

If you are weighing Relativity against other eDiscovery platforms, I'd recommend putting one question to every vendor before comparing features: is chat review priced on the collected set or on the scoped one? The answer reveals more about the eventual bill than any feature list. In summary, the scoped set is the KEY figure, and the collected count is only its starting point.

Microsoft Purview eDiscovery (Premium) cannot isolate one 1:1 Teams chat at the search stage, so the participants' other chats are collected too, and the review set swells unread.

Allow me to begin with a particular instance. On 2026-07-18, a Microsoft 365 E5 customer asked on Microsoft Q&A how to export only one 1:1 chat between two users; searches on Participants, Sender, Recipients, KeyQL and Message Kind still returned other 1:1 chats, group chats and meeting chats. A Microsoft moderator answered that Conversation ID, Thread ID and Conversation Type "aren't supported search properties." The supported path is to collect the participants' chat, commit it to a review set (the working copy that reviewers code), and rebuild the thread there. Even the ConversationType field reads "Group" for 1:1 and group chats alike.

According to Microsoft Learn, Teams keeps chat in two separate copies, one for compliance, used by eDiscovery, and one for user access. The search therefore works on the compliance record. Precision must come afterwards.

My position follows from this. The message count in a chat review set records what the tools could not exclude, not what the matter requires. A per-message price built on that count overstates the burden. As a 2025 CCSP audio course put it, "Search strategies transform vast datasets into manageable review pools." No published source yet measures how manageable a chat pool becomes. I shall not invent that figure; I shall show how to find it.

What share of a chat collection is actually relevant?

No published source, OpenText's EDRM guidance included, measures the responsive share of a chat collection; the defensible figure is the one measured on your matter, each coded message traced to its source.

An analysis of 7 sources shows that none publishes a measured responsiveness rate for chat. In February 2022, OpenText reported on the EDRM blog that over 600 billion chat messages had been sent among businesses alone in the prior year, and that Microsoft Teams usage had increased by more than 330 percent; neither figure carried a cited source. By a responsiveness rate I mean the share of collected items that reviewers code as answering a discovery request. It is worthy of notice that the same post described chat as "inherently designed to encourage long threads and large volumes of data," yet offered no figure for how much of that volume ever bears upon a case.

Because the number is absent, I would put every chat estimate through what I call the reach-versus-need test, a three-question lens that separates what a collection tool gathered from what the matter requires:

  1. Reach: how many messages did the search and collection return?
  2. Need: how many of those fall within the custodians, dates and conversations the requests actually name?
  3. Proof: of that narrower set, how many does a reviewer code responsive, with each call traced to its source message?

A common misconception is that a larger chat collection means a proportionally larger review. The reality is otherwise, and OpenText's own guidance concedes the point: it recommended breaking chat histories into blocks of time "to narrow the volume of chat that requires eyes-on review." The same post, it is curious to note, enumerated the particulars that swell a chat set, namely replies, reactions, edits, deletions, leave and join events, and attachments. In practice, a raw message count measures reach, not need, which is why a small case with chat data can still arrive as a large review job. The difference is KEY.

The economics follow from that distinction. In our experience, AI-assisted review runs cents per document, against the dollars per document and roughly $19K/GB of manual review, a difference the buyer can see on their own matter. Our outputs also link back to the exact exhibit they came from, so an attorney can verify any answer with one click before filing. I am inclined to think that this pairing, a lower unit cost and a visible trail to the source, is what allows a firm to test the need figure rather than accept the reach figure on faith.

What would a measured rate show once collected? For a particular matter, it would give the proportion of collected messages coded responsive, and it would let a buyer set that proportion beside the raw count before accepting any quote. The evidence on record does not yet supply that proportion. I shall not supply one in its place.

The takeaway is plain: price the need, not the reach.

Summary: The responsive share of a chat collection is not yet a published number, and the only figure worth trusting is the one your own matter produces, traced message by message to its source.

What will matter most for chat review over the next 12 to 24 months?

Over the next 12 to 24 months, I expect chat budgets to shift from paying per message reviewed toward paying to scope, isolate and cull conversations before any reviewer reads them.

I hold this view because three observations, drawn from quite different quarters, point the same way. None of them measures the responsive share of chat directly. That figure has not yet been published by anyone, and each signal below should be read as circumstantial rather than conclusive.

PredictionWeak signalWhy it mattersSource
Teams collections will keep outgrowing the conversation a matter concerns, and cost negotiation will turn toward isolating conversations. According to a Microsoft moderator on Microsoft Q&A, Conversation ID, Thread ID and Conversation Type "aren't supported search properties" in Purview eDiscovery (Premium). When a conversation cannot be isolated at search, the message count reflects the tool rather than the matter. A per-message quote then prices chats nobody requested. Microsoft Q&A, July 2026
Chat holds will be set earlier and drawn wider, well before a complaint is served. The duty to preserve begins when a reasonable person would expect litigation, and cloud holds may involve suspending deletion policies for chat histories. An early, broad hold enlarges the preserved set long before anyone knows what is responsive, so the distance between what is kept and what is needed will widen. CCSP audio course, Episode 91, September 2025
Chat search definitions will draw scrutiny from both sides of a matter. Microsoft's guidance ties a purge to every item a search returns, while Teams keeps separate compliance and user copies of chat. An over-broad search inflates what is held and reviewed; an under-broad one risks spoliation or an incomplete purge. A documented, conversation-level scope serves as cost control and risk control at once. Microsoft Learn, June 2026

It is worthy of notice that all three signals concern the stage before review. The collection is fixed by search properties, the hold by the date on which litigation became foreseeable, and the purge by the same search that set the hold. Review inherits whatever those earlier choices produce. For that reason I am inclined to think the budget conversation will migrate upstream, toward whoever can show, in writing, why a given conversation was kept or set aside.

What would change this forecast: if Microsoft added Conversation ID, Thread ID or Conversation Type as supported search properties, chat could be isolated at the search stage. Over-collection would then shrink at the source. I would expect the argument to move from scope toward review quality, and the case for culling would weaken accordingly.

What most buyers miss: more business chat does not mean a matching rise in review bills. ComplexDiscovery projects review's share of eDiscovery task spend at 52 percent for 2030. Volume grows at collection. Need grows only as fast as the requests for production, and a buyer who budgets on the first will be paying for reach.

Why do chat collections come back larger than the matter?

Chat is collected in coarse slices, whole channels and 12-hour chunks, so non-relevant, even privileged, talk rides along; that surplus should be processed where it never leaves your control.

According to Greg Buckles, editor of eDiscovery Journal, writing on December 26, 2024, the new Microsoft Purview eDiscovery portal collects chat in 12-hour chunks. In Microsoft 365 Teams, 1-1 chats are stored in participant mailboxes, while channel messages are owned by the Team site group and stored in that group mailbox. Buckles thinks of channel chats as "recordings of conversations, reactions and item exchanges that take place in a virtual 'room' over time," in which participants come and go across "multiple interwoven topics that can cover months." It appears, upon such a description, that the unit of collection is seldom the unit of relevance.

Three particulars, as Buckles records them, carry a collection past the matter's need:

  1. Coarse slices. Time-clustered hits fit the 12-hour chunks, but "occasional mentions scattered over days may require entire channel chat collection by broad time bands."
  2. Uncertain sources. Custodians "frequently forget the full names of channels," so interviews alone cannot fix where a conversation lived.
  3. Versioned attachments. Cloud attachments can be retrieved as the latest version only, the last 10, the last 100, or all versions, and he warns that version retrieval "can make a huge impact on matter volume."

Buckles adds that broad collection "may pose redaction/disclosure issues by collecting non-relevant conversation segments." In practice, over-collection is a privacy cost as well as a review cost.

Here, however, the simple answer breaks down. A small responsive share does not make chat a minor source. Buckles writes: "All of the key determinative evidence I have extracted from bloated holds recently seems to come from cloud instant messaging sources." He also recalls the Enron era, when his team kept a "5000/1 rule" on email: for every 5,000 emails reviewed, one violation was referred to HR. It is worthy of notice that this was never a responsiveness rate; it counted referrals, a far narrower thing. What this means is that rarity and importance travel together. The decisive message may be one among thousands.

The tension, then, is this: the surplus must be culled, yet culled with care, and it must be held somewhere safe while that work proceeds. In our experience, privileged evidence never leaves the client's control. Processing runs single-tenant, or inside the client's own AWS account under their own keys, with no vendor retention and no model training, so nothing fed to it trains anyone's model or waives privilege. I am inclined to think that in-account deployment of this kind is a trust posture no consumer AI tool, and almost no small-firm tool, can match.

TIP: I'd recommend testing any platform on the hard case rather than the easy one. Bring a messy chat collection and a hard question, and watch the system read, code and cite your own evidence instead of reading a case study about someone else's.

Summary: Chat collections outgrow the matter because the tools gather by chunk, channel and version rather than by conversation, and the few messages that matter are rare but decisive, so the surplus must be culled carefully and processed under your own control.

Does narrowing a chat search create defensibility risk?

Narrowing carries risk when it is careless or undocumented; hash-verified originals, an append-only audit trail and a fail-closed privilege gate let a narrowed chat set withstand challenge.

It is remarkable how evenly the danger is divided. According to Episode 91 of the Certified CCSP Audio Course, published in September 2025, "Poorly designed searches risk missing critical documents or overproducing irrelevant material, increasing cost." The same episode notes that in cloud environments a legal hold may involve suspending deletion policies for mailboxes, chat histories, or object storage, and that "preserving a chat without timestamps or participants renders it nearly useless in court." A narrow search is therefore no remedy if it strips away the context that makes a message intelligible.

Microsoft's own documentation shows how much weight a single chat search definition carries. According to Microsoft Learn's guidance on finding and deleting Teams chat messages, dated June 2026, "The deletion process deletes all items returned by the search," and once the purge command runs, "you can't undo it and email messages and chats can't be restored." What this means is plain. One search definition governs what is held, what is reviewed and what may be destroyed.

For that reason I'd recommend keeping what I call a scope record, a short written account of every narrowing decision, built in three steps:

  1. Name the locations. Microsoft Learn places 1:1 and group chats in participants' mailboxes, standard and shared channels in the parent team's mailbox, and private channels in a dedicated mailbox for each private channel.
  2. Record the terms. Microsoft advises including "a date range or several keywords to narrow the scope of the search to items relevant to your investigation," and recommends the Type condition with the "Instant messages" option to identify Teams conversations most comprehensively.
  3. Report before acting. A report-only export lets you examine detailed metadata before deletion, and I would keep that report in the matter file.

The CCSP episode adds that courts respect documented egress and scoping plans as "evidence of reasonableness." In practice, a written scope is a cost control and a risk control at once.

The record needs a sound vessel. In our experience, the process holds up if opposing counsel challenges it when originals are immutable and content-hashed, audit trails are append-only, chain of custody is documented, and a fail-closed privilege gate stands before production. A fail-closed gate, in plain terms, is one that defaults to blocking production rather than allowing it.

Such discipline was long confined to enterprise platforms. Relevant e-Discovery was built to bring enterprise-grade defensibility to small-matter economics, operable and affordable for a solo or small firm on a single messy matter, the segment Relativity and Everlaw price out and cannot serve without an administrator. The takeaway is that careful narrowing no longer requires an enterprise budget.

Summary: Narrowing a chat search is not itself the danger; an undocumented or context-stripping narrowing is. Name the locations, record the terms, report before acting, and keep hashed originals behind an audit trail, and the narrowed set will bear scrutiny.

The next 12-24 months, scored

Where the cost of reviewing chat heads next

Forecasts on how chat volume, Microsoft Purview search limits and shifting eDiscovery spend will change what reviewing chat really costs.

6 sources analyzed5 web sources1 podcast
A

What changes for chat review budgets

Use these forecasts to test chat review quotes, collection scopes and preservation plans over the next two years.

75/100
Medium confidence 12-24 months

Chat search definitions will face more scrutiny from both sides of a matter. The same search that sets a legal hold on chat histories also determines exactly which Teams messages a purge permanently destroys.

62/100
Medium confidence 12-24 months

Over the next 12-24 months, Microsoft Purview eDiscovery (Premium) users will keep collecting far more Teams chat than they need to reach a single conversation. Cost negotiations will shift toward isolating conversations after collection rather than haggling over per-message review rates.

Weak signals watched: On 2026-07-18 a Microsoft 365 E5 customer reported being unable to export a single 1:1 Teams chat. A Microsoft moderator confirmed that Conversation ID, Thread ID and Conversation Type aren't supported search properties. Separately, 1-1 chats are stored in participant mailboxes. Review fell from 73 percent of eDiscovery task spend in RAND's 2012 baseline to a reconciled 62 percent in 2025. Over the same period, over 600 billion business chat messages were reported sent in a single year. Microsoft's guidance states that a purge deletes every item the search returns and that deletions are permanent. Cloud legal holds may require suspending deletion policies for chat histories, and chat must be captured during its short lifespan.

B

Sources behind the chat cost outlook

Each public source below is listed with the specific line it contributes to the chat discovery forecasts.

Source What it states Forecasts it backs
Managing Emerging Data in eDiscovery - ComplexDiscovery [Web source] Review's share of total eDiscovery task spend has fallen from 73 percent in RAND Corporation's 2012 baseline to a reconciled 62 percent in 2025. “Collection, over the same span, has expanded over threefold, from 8 percent to a projected 25 percent, a 17-percentage-point gain.” Chat growth decouples from review spend
How to handle chat data in eDiscovery and investigations - EDRM [Web source] "In the past year, over 600 billion chat messages were sent globally among businesses alone." (no underlying source cited). “Chat data is ephemeral and needs to be captured during its short lifespan.”
Chat is described as ephemeral ("needs to be captured during its short lifespan"), structurally distinct, and designed to produce "long threads and large volumes of data.".
Chat growth decouples from review spend
Chat search scope draws defensibility scrutiny
2025 - The Year of Chat - eDiscovery Journal [Web source] In the Enron discovery era, Buckles' team used a "5000/1 rule" on email: for every 5,000 emails reviewed, one violation was referred to HR. “Email is dead, long live chat!”
In M365 Teams, channel messages are owned by the M365 Team site group and stored in that group mailbox. 1-1 chats are stored in participant mailboxes.
Chat growth decouples from review spend
Chat over-collection outlasts search fixes
Find and delete Microsoft Teams chat messages in eDiscovery [Web source] The purge deletes every item the search returns, so search scope directly determines what is destroyed. Deletions are permanent and unrecoverable. “Verify Teams messages carefully before purging and avoid using cmdlets to purge Teams messages (the user copy isn't deleted).” Chat search scope draws defensibility scrutiny
Episode 91 - E-Discovery: Preservation, Collection and Production in Cloud [Podcast] In cloud environments, legal holds may involve suspending deletion policies for mailboxes, chat histories, or object storage. “For instance, preserving a chat without timestamps or participants renders it nearly useless in court.” Chat search scope draws defensibility scrutiny
How can I export only a specific 1:1 Microsoft Teams chat using [Web source] Microsoft External Staff moderator Srikanth Chavithina stated that Conversation ID, Thread ID, and Conversation Type "aren't supported search properties" in Purview eDiscovery (Premium). “I also noticed metadata fields such as Conversation ID, Thread ID, and Conversation Type, but I can't find a way to use them to isolate a single conversation.” Chat over-collection outlasts search fixes
The sources behind the forecasts above: what each one states, and which forecasts lean on it.
C

What would shift the chat cost outlook

These changes in Microsoft tooling, preservation practice or review economics would weaken or reverse the forecasts above.

A note on uncertainty

A score measures how much current evidence backs a call, and that evidence keeps moving. The top forecast here sits at 75/100, while the minority view at 75/100 shows where the sources still disagree.

  • Chat growth decouples from review spend. The moment regulators or buyers head the other way, that call is the exposed one.
  • Chat growth decouples from review spend. Should the evidence swing against the mainstream view, that forecast outlasts the rest.
Methodology Each forecast is scored 0-100 from the public sources shown for it: how many there are and how authoritative they are.
Attorney initialing a written chat search scope beside a sealed evidence drive with a chain-of-custody tag
A narrowed chat scope holds up when it is written down, signed and tied to verified originals.

Is your chat review priced on reach or on need?

A per-message quote multiplies a raw chat count by a review rate, yet that count reflects how the tool gathered, not what your requests require.

ComplexDiscovery reports review's share of eDiscovery task spend falling from 73 percent in RAND's 2012 baseline to a reconciled 62 percent in 2025. Scoping is where the money now goes. Bring a chat collection and begin where Greg Buckles and OpenText's EDRM guidance both begin, by narrowing what needs eyes-on review through a documented search, the same search Microsoft warns decides what a purge destroys.

Scope my chat collection

How can I reduce the cost of document review in litigation?

Reduce the set before you price it: document a conversation-level scope, measure the responsive share on a sample, and pay for review of the need rather than the reach.

For chat, I expect the price of review to follow the price of scoping. So long as Purview's search cannot isolate one conversation by its Conversation ID, the review set will record the tool's reach, and a per-message quote will charge for it. That is the whole of my argument.

According to Microsoft Learn, a report exported before a purge "helps refine your search scope and ensures a more targeted and accurate purge." I would hold chat review to the same discipline. Report first, then read.

I confess the trade-off is real. A 2025 CCSP course observed of custodians that "Missing someone risks leaving gaps in evidence, while over-inclusion inflates costs and complexity." No scope escapes that tension; a written one merely shows where the line was drawn.

In summary, the measured responsive rate for chat is still missing from the published record. Put your next collection through the reach-versus-need test, write the result into your scope record, and let the quote follow the number.

Written by

Michael

Kansky

Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

What else do litigators ask about chat relevance?

Most questions about chat review turn on three matters: when the duty to preserve begins, how narrowly a collection can be drawn, and how the review ought to be priced.

What percentage of collected chat messages are relevant?

No published source yet measures it, and I would distrust any quote that assumes a figure. A responsive message is one that a request for production actually calls for. The defensible number is the one you measure on a coded sample from your own matter.

When does the duty to preserve chat begin?

It begins when a reasonable person would expect litigation, not when a complaint is filed. A demand letter, a regulatory inquiry or an internal complaint may each suffice. In cloud systems, a legal hold (an instruction that keeps data from deletion) may require suspending deletion policies for chat histories.

Can I collect only the one Teams conversation my matter concerns?

Not at the search stage in Purview eDiscovery (Premium). Conversation ID, Thread ID and Conversation Type are not supported search properties there. I'd recommend collecting what the search returns, isolating the conversation during processing, and writing down exactly how it was done.

Does a Purview purge remove Teams chat from users' view?

Not necessarily. According to Microsoft Learn, Teams stores chat in two copies, one for compliance and one for users, and a cmdlet-based purge removes only the compliance copy. The user copy stays visible. For that reason I would never treat a purge as a culling tool.

How do I manage a large chat collection under deadline pressure?

Reduce before you review. In my view, the order that serves best is a written conversation-level scope, then a coded sample to measure the responsive share, then review of the narrowed set alone. Each step shrinks what the deadline must cover.

Should chat review be priced per message?

Only on the narrowed set. A per-message rate applied to a raw collection prices how the tool gathered, not what the requests require. I am inclined to think the fairer basis is the measured responsive share of your own matter.

Read next

Laptop running a homemade AI review workflow next to boxes of litigation documents and a sealed evidence boxEdiscovery

Building Your Own AI Review Agent: Where It Breaks

Yes, you can build one in an afternoon. A do-it-yourself AI review agent refers to a trigger, a container job and a model call, and it skips what makes review defensible.September 27, 202632 min read
Law firm server room showing AI document indexing pipeline running before privilege review screen is completeEdiscovery

Does Privileged ESI Get Embedded Before the Screen?

Yes. In most RAG e-discovery platforms, privileged ESI refers to attorney-client communications and work product that gets embedded into the vector index automatically at collection, before any privilege review queue runs.September 25, 202625 min read
A legal professional reviewing threaded Slack conversation data on a monitor, with JSON code visible alongside a structured conversation view, showing the transformation from raw data to readable threads in an eDiscovery contextEdiscovery

How to Produce Slack Messages Without Breaking Threads

To produce Slack messages defensibly in discovery: (1) Export the full workspace using the standard tool for public channels, or Slack's Discovery API (Enterprise Grid only) for private channels and DMs.September 14, 202630 min read

See it on your matter

Bring us a messy collection - mailboxes, scans, phones, recordings - and watch it become one searchable, defensible record.