
Quick Answer
AI coding 100,000 documents in your own AWS account has no single price. It splits into a metered cloud bill, a vendor license and human validation hours.
The meter is the cheap, public part: Amazon's own worked example prices plain OCR at $150 for 100,000 pages, before the model's per-token coding pass. The rest is harder to see. One account of discovery's digital shift argues that digitization could have lowered costs, and litigators instead used the document surge to drive up opponents' costs. Relevant e-Discovery's answer on fees is cents per document against dollars for manual review, shown on your own matter.
Key Points
- Amazon's price page lists Textract's plain OCR at $0.0015 a page for the first million pages, or $150 for 100,000 pages in Amazon's own worked example.
- ComplexDiscovery's Winter 2026 eDiscovery Pricing Survey, based on 53 responses, found per-document and hybrid GenAI review pricing tied at 28.3% each , with only 5.7% naming per-token pricing.
- Human validation is billed at human rates: a 2025 eDiscovery review update on JD Supra reported most remote review rates fell between $25 and $40 per hour .
Only the cloud's page meter comes with a public price list. The other lines on an AI review bill do not.
Amazon will tell you what a page costs to four decimal places. It publishes the rate, meters every page and sends you the arithmetic, which is, needless to say, more candor than most review quotes manage.
I think that asymmetry is the whole story of AI coding in your own AWS account. The cloud meter is public and checkable. The vendor license and the validation hours are not, and the index hosting bills for as long as the matter stays open. I would test those on your own documents: Relevant e-Discovery's demonstration asks a buyer to bring a messy collection and a hard question, then watch it read, code and cite their own evidence.
Discovery has run this experiment before. One account of the shift to digital files argues that digitization could have lowered discovery costs by making review easier, and that litigators instead used the surge in digital documents to drive up their opponents' costs. Total litigation costs stayed high.
The AI version is already visible:
- A 2025 industry analysis projected that review would fall from 64% to 52% of total eDiscovery spending by 2029.
- A 2024 survey of senior legal leaders across eight countries found that few firms using AI passed the savings to clients, while a larger group charged premium rates.
- In one early 2026 illustration of a review that got dramatically cheaper, clients rarely saw the whole $170,000 saving.
The money is not vanishing. It is changing lines.
Relevant e-Discovery runs single-tenant or inside the firm's own AWS account under the firm's own keys, so the cloud portion lands on an invoice the firm can read for itself. What follows prices one matter the way the bills really arrive: by the page, by the document, by the hour and by the gigabyte per month. Start with the page.
Coding 100,000 documents with AI inside your own AWS account produces three bills, not one: a metered cloud charge, a vendor license and human validation hours.
AI coding means a model reads each document and tags it for relevance, issue and privilege. Own-account means that work runs under your keys, so Amazon bills you directly: Amazon Bedrock for the model's tokens, plus storage, the search index and compute.
I would first ask whoever administers your AWS account whether this matter's charges can be separated from the firm's other cloud spending, because a bill you cannot read by matter cannot be checked or passed on. Own-account processing also buys something no price list shows: Relevant e-Discovery states that it runs with no vendor retention and no model training, so nothing a firm feeds it trains anyone's model or waives privilege.
Consider how people rate cloud software they plainly adore. On Gartner Peer Insights, Box holds a 4.4 overall rating from 573 ratings, with 45% at five stars, 44% at four stars and 0% at one star. Nobody, apparently, hates the product. One complaint that does surface concerns the licensing.
Review pricing has the same habit. The first question for any quote is therefore its unit: per document, per gigabyte, per token, a flat monthly subscription or a hybrid of them. Technology-assisted review, the previous generation, earned its savings through tiered staffing, so ask any vendor quoting a saving which baseline it was measured against. Relevant e-Discovery states its own fee position plainly: cents per document for AI-assisted review, dollars per document for manual review, a gap the buyer can see on their own matter. So begin with the one line Amazon itemizes for you.
What does your own cloud account charge for the AI first pass?
Your own account bills extraction per page and inference per token, under your own keys, and the extraction route you choose moves that bill more than the model does.
Before pricing a single token, settle four things:
- Count pages, not documents, because extraction services bill by the page.
- Pick the extraction route first: plain OCR, layout conversion, or page images sent straight to a vision model.
- Store every extraction output, so a rerun never pays for the same page twice.
- Price inference from the provider's current rate page on the day you budget, because no published Amazon Bedrock rate is quoted in this analysis.
The per-page rule is not a technicality. In 2025, a developer on Google Cloud's community forum woke up to a budget warning after a PDF tool was invoked 26 times in a single day. Commenters explained, with the patience of people who had also once woken up to a budget warning, that every Document AI service bills per page, not per request. The same 2025 thread priced Google's Enterprise Document OCR at $1.50 per 1,000 pages for the first 5 million pages, with async processing at 60 cents per 1,000 pages. One commenter summarized the learning curve with admirable economy: 80 bucks spent before figuring it out, and $1.50 for the rest of the job afterward.
Volume can also be bought ahead: in a December 2025 r/AZURE thread, one team said it handled roughly 2.6M pages of burst ingestion in about 12 hours and bought a pre-paid quantity for a first-month discount to absorb it. Teams running Azure pipelines report the Read OCR model at $1.5 per 1,000 pages, the Layout model at $10 per 1,000, and page images sent directly to GPT-4o vision at roughly $0.05 to $0.07 per page. One engineer's cost modeling found that letting the language model do everything was an order of magnitude more expensive than converting to Markdown first and calling the model second. Another reported processing an organization's 1 million files for about $1,000 instead of $20,000.
In practice, your page count sets the floor. Reruns are what quietly raise it.
The common assumption is that the model is the expensive part. On extraction, the reverse is closer to the truth: the cheap meter is OCR, and the costly habit is mailing photographs of documents to a model that then has to squint at them, page by page, at a premium.
There is an honest trade-off buried here. Async jobs and self-hosted open-source OCR are cheaper and also slower, which matters very little for a one-time matter ingestion and enormously for anything a partner wants tomorrow.
Where the meter runs is the other half of the question. Our in-account AWS deployment runs under the client's own keys, so extraction and inference charges land on the firm's own invoice instead of disappearing inside someone else's price. For a fuller account of what that changes, see who controls what in a managed service versus your own AWS account.
Managed platforms fold that meter into a license, and buyers notice. A reviewer on Gartner Peer Insights who rated Box a full 5.0 still called its account licensing "confusing and quite expensive for the product provided." The grievance was the licensing, not the file sharing it paid for.
Set side by side, 4 sources show one pattern: per-unit meters are cheap, and bills grow through volume, repetition and licensing terms. Which brings the ledger to the line most vendors would rather quote as one tidy number per document.
How wide is the gap between the cloud meter and the review bill?
A developer woke to a budget warning after a document parser had run just 26 times in a day. The bill stood at $38, and the reason matters for anyone pricing AI review.
In a May 2025 thread on r/googlecloud, one commenter explained: "All Document AI services charge per page not request." At the Form Parser rate of $30 per 1,000 pages, $38 came to roughly 1,266 pages, or about 48 per request. Extraction in your own AWS account is metered the same way.
Amazon's price page shows how small that line can be: Textract's plain OCR costs $0.0015 a page for the first million pages, or $150 for 100,000 pages in Amazon's own worked example. Google lists its enterprise OCR processor at $1.50 per 1,000 pages. Richer routes cost more. An engineer comparing Azure options in December 2025 reported about $10 per 1,000 pages for prebuilt models, $30 for custom extraction, and $0.05 to 0.07 a page to send page images straight to a vision model. Another commenter in the thread found the all-LLM route "an order of magnitude more expensive" than converting documents to Markdown first.
Buyers are quoted in other units. In ComplexDiscovery's Winter 2026 eDiscovery Pricing Survey, based on 53 responses, only 5.7% named per-token pricing as their primary model for GenAI review. Per-document and hybrid models tied at 28.3% each. The survey's analyst concludes that "providers are currently absorbing token cost variability and presenting buyers with higher-order pricing units." The analyst sees $0.11 to $0.50 a document emerging as the competitive range for GenAI review, set against $0.50 to $1.00 for human review, and traces the price spread to "task complexity, model selection, and quality control overhead."
Our software runs inside a firm's own AWS account under its own keys, so we have a stake in seeing the meter itemized. The bigger question sits further down the bill.
"The most important near-term challenge for the market is not the headline per-document or per-GB rate, but the hidden cost variables: exception document handling, quality control overhead, model retraining requirements, and the total cost of ownership of integrating GenAI review into existing workflows."
ComplexDiscovery analyst, Winter 2026 eDiscovery Pricing Survey analysis, 2026
The 2025 edition of the survey found about half of respondents unsure how exception documents are priced in GenAI review. In a December 2024 thread on r/ediscovery, practitioners described those costs. One listed the "additional cost of set up" and "QC stages with senior reviewers." Another said AI output gets samples and elusion tests at roughly "triple" the usual level, especially with privileged information. A third found the "reoccurring charges for rerunning models" too costly for their firm. Individual accounts cannot show how common such costs are, but they match the analyst's list.
Human rates drive those costs. In the 2026 survey, 41.5% of respondents put remote review attorneys at $25 to $40 an hour, and 35.8% put them above $40. RAND's 2012 study of 57 large cases put the fastest human pace at about 100 documents an hour, assuming simple documents and motivated, experienced reviewers. Validation relies on the same people at the same pace, so its cost grows with how much the protocol asks them to check.
Hosting runs on a clock of its own. Basic hosting has commoditized: 54.7% of 2026 respondents reported under $10 per GB a month, while 11.3% reported analytics hosting above $25. The charge runs for as long as the matter stays live, however quickly the coding pass finishes.
Each line uses a different unit: extraction is billed by the page, the vendor quote by the document, validation by the hour and hosting by the gigabyte per month. Only the page meter has a public price list you can check against your own cloud bill. The lines that decide the total are measured in reviewer hours and in the months a matter stays open, and a blended per-document price hides exactly those lines.
Five checks before you sign
- Count pages before you price extraction, and ask which route runs first: plain OCR, a layout model, or page images sent straight to a model.
- Ask any vendor, us included, whether the per-document fee covers model inference only, or QC, exception handling and reporting as well.
- Get exception handling in writing: what sends a document out of the AI workflow, how that document is billed, and who checks it.
- Ask how much sampling and elusion testing the validation protocol calls for, especially on privilege, and price those hours at your reviewers' actual rates.
- Multiply the hosting rate per GB per month by the months you expect the matter to stay open, and ask whether analytics hosting is priced separately.
How we checked this
We checked Amazon's and Google's published extraction prices on October 8, 2026. We also read ComplexDiscovery's analyses of its 2025 and 2026 eDiscovery pricing surveys, a 2012 RAND study of discovery costs, and three practitioner threads on Reddit. This section uses no figures of our own. The Azure rates come from a practitioner's research, not from Microsoft's price list. Some Reddit commenters disclosed that they work for document-extraction vendors. The 2026 survey had 53 responses, and the RAND data is more than a decade old. The chart assumes one page per document, so real extraction costs will be higher, and it leaves out the model's per-token coding charges. We sell AI review software that runs in a firm's own AWS account, so we have a commercial stake in this question. Still unknown: how many validation hours a typical AI coding pass needs, and how often documents fall out of the AI workflow as exceptions.
- Amazon Web Services, Amazon Textract pricing page, checked October 8, 2026.
- Google Cloud, Document AI pricing page, checked October 8, 2026.
- r/googlecloud, thread on Document AI per-page charges, May 2025.
- r/AZURE, thread on AI document extraction options, December 2025.
- ComplexDiscovery, analysis of the Winter 2026 eDiscovery Pricing Survey, March 2026.
- ComplexDiscovery via JD Supra, 2025 eDiscovery review update, June 2025.
- r/ediscovery, thread asking whether AI review is too expensive, December 2024.
- RAND Corporation, research brief on producing electronic documents, April 2012.
Why does the vendor license hide inside a per-document price?
AI-assisted review runs cents per document against the dollars per document of manual review, yet most vendor quotes never show which cents are cloud meter and which are license.
The ComplexDiscovery Winter 2026 eDiscovery Pricing Survey, built on 53 responses, found that hybrid and per-document models tie at 28.3% each as the primary way GenAI review is sold. Per-GB pricing trails at 11.3%. Per-token pricing and flat monthly subscriptions each sit at 5.7%, with outcome-based deals at 3.8%. The survey's analyst read the split plainly: providers are "absorbing token cost variability and presenting buyers with higher-order pricing units."
Start with the model: which one does the quoted price assume, and what does the price become if the vendor changes it? Ask, too, whether the price assumes one platform or two: Relevant e-Discovery describes enterprise tools as supplying Bates numbering, privilege and chain of custody, newer AI entrants as supplying cited answers and chronologies, and its own product as the combination.
Then ask for the same quote with the model's usage billed to your own AWS account, and see how much of the per-document price remains. Somewhere inside it sit token consumption, GPU infrastructure and model licensing, the cost drivers the survey analysis names, baked together at an even temperature and served by the slice. The labels on the newer pricing units are no clearer: 64.2% of respondents said per-GB GenAI pricing either does not apply to them or is unknown.
The same analysis names $0.11 to $0.50 per document as the emerging competitive zone for GenAI review, and it tells buyers to find out what that fee actually covers. Before signing any per-document quote, ask three questions:
- Does the fee cover model inference alone, or quality control as well?
- How are exception documents, the ones that fall out of the AI workflow, billed?
- Is reporting included, or priced separately?
Contract terms move the number as much as the rate does, which is why it pays to read how an AI review vendor can swap models mid-matter before comparing quotes side by side.
Our own position on fees is that the economics should be demonstrable on the buyer's own matter, not asserted in a brochure, against a manual review benchmark of roughly $19K/GB. At Relevant e-Discovery we built the platform to be operable and affordable for a solo or small firm on a single messy matter, the segment the large enterprise platforms price out.
Who sees that difference is another question entirely. Axiom's 2024 study of more than 600 senior legal leaders, summarized by Amy Swaner in early 2026, found that 79% of law firms use AI to boost efficiency, while only 6% pass the savings to clients and 34% charge premium rates for AI-enhanced work. Swaner's own illustration is a document review that once cost $200,000 and now runs $30,000, with clients rarely seeing the whole saving.
A blended price protects whoever sets it. Itemized lines protect whoever pays it. And none of the three questions above prices the people who will check the machine's work.
How much human validation does AI coding still need?
Validation is the line billed at human rates, and it does not shrink with token prices; hashed originals, append-only audit trails and a fail-closed privilege gate are what make it defensible.
The costs that never make it into a per-document quote are well catalogued. In a 2024 r/ediscovery thread titled "Is AI too expensive?", practitioners listed the "additional cost of set up," "QC stages with senior reviewers," and recurring charges for rerunning models. One commenter said quality control on AI output meant samples and elusion tests at roughly "triple" the level used for existing technology, especially for privileged information. An elusion test, for the blessedly uninitiated, is a statistical sample drawn from the pile the system set aside, to find out how much that mattered was left lying in it.
Another commenter summarized the mood with the bluntness of someone typing at the end of a very long day: "no one trusts AI." A third put the economics in one line. AI review "still requires QC...so someone billing hourly."
That is the line I would budget first. Ask the vendor how large a sample its own protocol draws and who is expected to read it, then put that person's hourly rate beside the number.
The same 2024 thread shows how fragile the savings arithmetic is. In 2024, its original poster priced AI first-pass review at $0.20 per document and called it 10 times cheaper than manual review, calculated at $60/hour and 30 documents per hour. A practitioner replied in that 2024 thread that 30 documents per hour was "pretty low" for first-pass review. Raise that baseline, and the gap narrows before a single validation hour is counted.
Technology-assisted review went through the same reckoning a decade earlier. Henry M. Sneath of Houston Harbaugh writes that TAR cuts cost and time against manual linear review by some estimates 40 to 60%, and that the savings come from a tiered model: lower-rate first-level reviewers handle large batches, while higher-rate second-level reviewers spot-check smaller ones. Validation uses statistical sampling to measure recall, the percentage of relevant documents found, and precision, the percentage of retrieved documents that are actually relevant.
For a 100,000-document matter, the validation line has three parts to price:
- Sampling and elusion testing on the documents the AI set aside.
- Second-level attorney review of privilege calls, which even vendors selling AI review say should stay under attorney oversight.
- Written documentation of the process, because a challenged review will likely need affidavit support from counsel or the vendor.
None of the sources in hand measures how many validation hours a 100,000-document matter takes. On fees, Relevant e-Discovery's own answer is that the economics are demonstrable, not asserted: AI-assisted review runs cents per document against the dollars per document of manual review, a difference the buyer can see on their own matter. Ask Relevant e-Discovery to show it on your collection, then price the validation hours in your own protocol at your own reviewers' rates, because they depend on the sample design and the share of privileged material. For a closer look at what those reviewers should be checking, see where AI relevance scores come from and what they hide.
Documentation is where this spend can be contained. Our platform keeps immutable originals with content hashing, append-only audit trails, documented chain of custody and a fail-closed privilege gate on production, so the record a challenge would demand already exists when opposing counsel asks for it.
The machine's work is billed in cents. How much of it gets checked can be argued over before a single hour is billed: in the 2025 In re Insulin Pricing Litigation ruling Sneath describes, plaintiffs wanted pre-set criteria for when review could or must pause for validation, while defendants wanted to pick the stopping point themselves.
Which line on the bill decides what own-account review costs?
Not the tokens. On fees, Relevant e-Discovery's answer is that AI-assisted review runs cents per document against the dollars per document of manual review, a difference the buyer can see on their own matter.
Ask Relevant e-Discovery to show that difference on your own collection. Which of the other lines grows fastest is then a property of the matter: the longer a case stays open, the more monthly hosting it accrues, and the more privileged material it holds, the more attorney checking it needs.
The old numbers explain why. In 2012, RAND examined 57 large-volume cases at eight very large companies and found that review took a median 73 percent of the cost of producing electronic documents, with processing at 19 percent. People were the bill. Machines were the side dish.
I expect AI coding to keep moving that weight, not erase it. The machine reads the first pass for cents, apparently without complaint, and the hours migrate to sampling, spot checks and privilege review. Meanwhile, what the client pays is a separate decision: Amy Swaner describes firms holding their rates, adding an AI premium, or landing somewhere in between.
So ask for three lines on any quote before you sign:
- The cloud meter, billed by Amazon to your account.
- The vendor license, stated on its own.
- The validation hours, with who performs them and at what rate.
Relevant e-Discovery's answer on fees is that the economics should be demonstrated, not asserted: cents per document for AI-assisted review against dollars for manual review, shown on your own matter. Bring the collection. Then decide which of the three lines your own client will see.
Summarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently Asked Questions
What else do people ask about AI review costs in their own cloud?
Most follow-up questions come down to the same split: what the cloud meters, what the vendor licenses and what people still have to check by hand.
What does AI document review cost in my own AWS account?
It costs three things, billed by three parties. Amazon meters extraction, tokens, storage and the index. The vendor charges a license, and people bill hours for validation. None of the sources behind this article publishes a worked total for a 100,000-document matter, which is exactly why each line should be priced on its own.
What does a per-document AI review fee actually include?
A per-document fee is one price charged for each document the AI codes. It can blend cloud processing with the vendor's license, and the quote rarely says in what proportion. Ask whether reruns, storage and index hosting sit inside that fee or beside it.
How much human QC does AI document review still need?
Enough to defend the result, and all of it is billed at human rates. A 2025 eDiscovery review update on JD Supra reported that most remote review rates fell between $25 and $40 per hour, with some above $40. Elusion testing, which means sampling the documents the AI set aside to estimate what it missed, sits on that line. No token discount touches it.
How can I reduce the cost of document review in litigation?
Separate the bill before you try to shrink it. I would price the cloud meter, the license and the validation hours as distinct lines, then pick the extraction route on purpose, because it moves the cloud line more than the model does. Compare the result against technology-assisted review, not only against linear manual review.
How is AI coding different from technology-assisted review?
Technology-assisted review (TAR) is software that learns from reviewers' coding decisions to rank or classify the rest of a collection. Its savings came largely from tiered staffing. AI coding hands more of the first pass to a model, and yet the checking stays human.
What does Relevant e-Discovery charge?
Relevant e-Discovery's answer is that the economics are demonstrable, not asserted: AI-assisted review runs cents per document against the dollars per document of manual review, a difference the buyer can see on their own matter. For a figure on your collection, contact Relevant e-Discovery and bring the messy one.
Written by
Michael
Kansky
Michael Kansky is a serial software entrepreneur who has spent more than two decades building and bootstrapping profitable SaaS and services companies.
Connect on LinkedIn

