Why "summarize this" gives you a smooth summary with holes in it
I have pasted hundreds of long documents into Claude, and for months I asked for what everyone asks for: "summarize this." The output always read well. It was also, more often than I liked, missing the one detail I would later need. A termination notice period. A footnote that reversed the headline finding. The one customer who mentioned a bug that turned out to be real.
The problem is not that Claude is bad at reading. The problem is what the word "summarize" asks for. A summary is a compression, and compression means choosing what to drop. With no criteria for what matters, Claude falls back on what a summary usually looks like: main themes, overall tone, the conclusion. Numbers, dates, names, exceptions, and obligations get smoothed over because they do not fit a paragraph that flows.
So the fix is not a cleverer summary prompt. The fix is to stop asking for the summary first.
Extract first, summarize second
The pattern that changed my results is a two-step order. Step one is a structured extraction: Claude pulls every concrete item in the document into fixed categories before writing any prose. Step two is the summary, written from that extraction and the document, with instructions to reference the extracted items rather than paraphrase around them.
Extraction is lossless and summarization is lossy. When the lossless pass comes first, the lossy pass has a checklist it cannot quietly skip. Here is the prompt I use for almost any document under about 40 pages.
I am going to give you a document. Do not summarize it yet.
First, read the whole thing and produce a structured extraction with these sections. Put "None found" under any section that is genuinely empty. Do not merge items. Do not round numbers.
1. Facts and claims: every factual statement the document makes, one per line, with the page or section it came from.
2. Numbers: every figure, percentage, currency amount, quantity, and threshold, with its unit and what it refers to.
3. Dates and deadlines: every date, time period, and deadline, and what happens on it.
4. People and entities: every person, company, team, or role named, and what the document says they do or owe.
5. Obligations and commitments: anything anyone must do, may do, or must not do. Quote the exact wording for each one.
6. Exceptions and conditions: every "unless", "except", "provided that", "subject to", and similar carve-out.
7. Open questions: anything the document leaves undecided, says it will address later, or contradicts elsewhere in the same document.
After the extraction, write a summary of no more than [300] words. Every sentence in the summary that contains a number, date, or obligation must match an item in the extraction above.
Document: [paste the document]
The output is long. That is the point. You are trading a tidy five-paragraph summary for a checklist plus a summary, and the checklist is where the value is.
When the document is longer than the context window
Before chunking anything, check whether you need to. Current Claude models handle far more than most people assume. According to Anthropic's context window documentation, Claude Opus 5, Claude Sonnet 5, and the Fable models all have a 1M-token context window, while Claude Haiku 4.5 has 200K tokens. The models overview puts 1M tokens at roughly 555,000 words on the current tokenizer, and 200K tokens at roughly 150,000 words. A 300-page contract is nowhere near either limit. I wrote up how tokens map to real documents in Claude's context window explained if you want the fuller picture.
File uploads are a separate limit from the context window. Anthropic's support page on uploads says claude.ai accepts up to 500MB per file and 20 files per chat, and PDFs are capped at 1,000 pages. One detail there matters more than the caps: for PDFs of 100 pages or fewer, Claude analyzes both the text and the visual elements, and for PDFs from 101 to 1,000 pages it processes text only. If your report's charts carry the findings, split it so the chart-heavy sections land under 100 pages, or those numbers will not be read.
There is one more reason to chunk even when everything technically fits. The same context window page states that as token count grows, accuracy and recall degrade, which Anthropic calls context rot. In practice that shows up as a summary that is sharp on the first third of a very long document and vague on the last. So for anything past a few hundred pages, or anything where I need every clause, I chunk it and keep a running ledger.
The ledger is the piece most chunking advice leaves out. Summarize each chunk independently and then summarize the summaries, and you compress twice and lose twice. Instead, carry one growing extraction forward and have Claude add to it, so the final ledger is one lossless document you summarize once.
We are working through a long document in parts. This is part [3] of [8].
Below is the running ledger from the earlier parts. Do not rewrite or shorten it. Your job is to:
1. Read the new part.
2. Append new items to each section of the ledger (facts, numbers, dates, entities, obligations, exceptions, open questions), tagging each new item with "Part [3]".
3. If anything in this part contradicts, updates, or resolves an item already in the ledger, do not delete the old item. Add a line directly under it that starts with "UPDATE (Part [3]):" and explain the change.
4. At the end, list anything in this part that refers to a section you have not seen yet, so I can watch for it.
Return the complete updated ledger, not just the additions.
Running ledger:
[paste the ledger from the previous step]
New part:
[paste part 3]
When the last part is done, I paste the final ledger into a fresh conversation with the first prompt's summary instruction, and I get a summary written from a complete inventory instead of from a memory of eight chunks. It takes longer. It also ends the last-chapter surprise.
Prompts for specific document types
The generic extraction works, but each document type has its own places where details hide. These are the versions I keep saved.
Contracts
In a contract, the danger is not the clause you read. It is the clause that modifies the clause you read, three pages later.
Extract the following from this contract before writing anything else. Quote the exact language for every item and give the clause number.
1. Parties, and which defined term refers to which party.
2. Term: start date, end date, renewal mechanics, and whether renewal is automatic.
3. Termination: every way either party can end this, the notice required, and what survives termination.
4. Money: every fee, payment deadline, late penalty, price change mechanism, and cap.
5. Obligations of [my company name]: everything we must do, with any deadline.
6. Obligations of the other party: everything they must do, with any deadline.
7. Liability and indemnity: caps, exclusions, and who indemnifies whom for what.
8. Anything that is defined in one section and modified in another. List both locations.
9. Anything I would need a lawyer to interpret. Say why.
Then give me a one-paragraph plain English summary of the deal, and a separate list of the three clauses you think are least favorable to [my company name].
Contract: [paste the contract]
This prompt makes me a better-prepared reader before I talk to a lawyer. It does not replace the lawyer.
Meeting transcripts
Transcripts are long, repetitive, and full of decisions made in passing. I go deeper in how to use Claude for meeting notes, but the summarization-specific version is this.
This is a raw meeting transcript. Before summarizing, extract:
1. Decisions: anything the group agreed to. Quote the line where it was agreed and who said it. If it was proposed but not clearly agreed, put it under "Proposed, not confirmed" instead.
2. Action items: task, owner, deadline. If any of the three is missing, write "unassigned" or "no date" rather than guessing.
3. Numbers mentioned: budgets, headcounts, dates, metrics, with who said them.
4. Disagreements: any point where two people held different positions, and whether it was resolved.
5. Things deferred: anything someone said they would "come back to" or "take offline".
Then write a summary of no more than 200 words for someone who was not in the room. Do not include anything in the summary that is not in the extraction.
Transcript: [paste the transcript]
The "proposed, not confirmed" bucket is the one that earns its place. Most bad meeting summaries come from treating a suggestion as a decision.
Research reports
With reports, the headline finding is rarely the problem. The caveats live in the methodology and the footnotes.
Read this report and extract, before any summary:
1. The main findings as the authors state them, with the page number.
2. For each finding, the sample size, time period, and population it is based on.
3. Every limitation, caveat, or "further research is needed" the authors themselves state, and which finding it applies to.
4. Every number in the executive summary, and whether it matches the number in the body or tables. Flag any mismatch.
5. Who funded or commissioned the report, if stated.
6. Claims that are cited to another source rather than the report's own data.
Then write a summary that leads with the findings but attaches the relevant limitation to each one in the same sentence.
Report: [paste the report]
Item 4 catches more than you would expect. Executive summaries get edited late and separately from the body, and the numbers drift.
Customer feedback exports
A CSV of 800 support tickets is a different kind of long document. The failure mode is the summary saying "customers mostly want X" when four of them reported something that costs you money.
This is an export of customer feedback. Do not summarize by theme yet.
First:
1. Count the total responses and state the number.
2. Pull out every response that mentions a bug, an error, a refund, a cancellation, a competitor by name, or a legal or safety concern. Quote each one in full with its row ID. Do not group these.
3. List every specific feature request, one per line, with a count of how many responses asked for it. Quote one example each.
4. List every number a customer stated (a price, a time, a quantity).
Then group the remaining responses into themes, give a count per theme, and write a summary. In the summary, the individually quoted items from step 2 must appear by row ID, not be folded into a theme.
Export: [paste the export, or attach the file]
The verification pass
Whatever prompt you used, the summary is a claim about the document, and claims should be checked. I run this as a second message in the same conversation.
Go back through the summary you just wrote. For every number, date, name, deadline, and obligation in it, produce a table with three columns: the item as it appears in your summary, the exact quote from the source document that supports it, and the page or section where that quote lives.
If you cannot find a supporting quote, write "NOT FOUND IN SOURCE" in the second column. Do not rephrase the source to make it match. Do not remove the item from the summary. I want to see the gaps.
When a row comes back NOT FOUND, either Claude inferred something reasonable that the document does not actually say, or it merged two nearby facts into one. Both are worth knowing before that summary goes to a client. I cover the broader habit in how to fact-check Claude responses; this table is the fastest version of it for documents.
If you find yourself reaching for these prompts every week, the extraction and verification patterns above are two of the ones in the Claude Cheat Sheet, 120 prompt patterns with worked examples and notes on when each one does nothing, $19 once, two free samples on the page.
What never to trust a summary for
A good summary tells you where to look. It does not replace looking. I read the source every time in these cases, however clean the extraction came back.
Anything you are about to sign. If a clause is going to bind your company, read the clause itself, and if the amount is material, pay a lawyer to read it too.
Anything with a legal, medical, tax, or regulatory consequence. A summary that gets the general shape right and one threshold wrong is worse than no summary, because it feels finished.
Anything where the exact wording matters. Warranties, disclaimers, guarantees, "business days" versus "calendar days." Claude will usually quote these correctly when asked, but you are relying on the quote, so confirm the quote.
Anything you will be quoted on. Before you tell your board "the report found a 12 percent decline," open the report to the page the verification table pointed you to and look at the number yourself.
Anything scanned, image-heavy, or over 100 pages as a PDF. Visual elements are not analyzed past 100 pages, and a scanned page with poor text extraction can silently drop content. If the summary of a long scanned PDF feels thin in one section, that is usually why.
None of this makes the two-step method less useful. It makes it useful for the right thing: getting you to the parts of a long document that deserve your full attention, with a complete inventory of what is in there.
Where to go from here
If you want the fuller set of habits behind this, start with the free 75-page Claude guide, which covers how to structure a request so Claude does the lossless work before the lossy work. When you want the saved versions of prompts like the ones above, the Claude Cheat Sheet is the next step, with two free samples on the page so you can see the format first.