Introduction
Legal research often involves a large number of multi-page documents. A case file can grow to thousands of pages over time and may contain complaints, motions, exhibits, court orders, and other material. Importantly, these documents describe different stages of the same matter, often using different language for the same people, events, and legal issues.
A retrieval system that provides useful answers may therefore depend on evidence spread across the case file. For example, a question could require connecting an allegation in a complaint with a later court ruling and the remedy recorded in a settlement.
This is a multi-hop question: the answer depends on several pieces of evidence that appear in different passages and documents.
This makes retrieval more difficult as the case file grows. The system has to identify which documents are likely to contain relevant information and then recover the right passages from them.
In this article, I explore whether document maps can help with this by allowing a RAG system to find the right documents in a large case file before surfacing specific passages.
The benchmark in this article uses legal case files from Multi-LexSum, a dataset based on civil-rights litigation records. I compare six document-map strategies with a standard retrieval baseline and evaluate how well each approach recovers the evidence needed to answer the questions. See the full benchmark in the repository.
The central question is:
Can document maps help a retrieval system route complex questions to the right evidence?
The Retrieval Problem
To understand why document maps might help, it is useful to start with how a standard RAG system retrieves information.
A basic RAG pipeline can split documents into passages, or chunks, embed those passages, and use dense retrieval to find the chunks most similar to a question:
In dense retrieval, each chunk is scored independently against the question. This can become a problem for multi-hop questions, where the answer depends on evidence spread across the corpus.
Long documents spread related information across many passages. Multi-document case files make this even harder because different documents may describe the same people, events, or issues at different stages of the case.
Consider this example. A lawyer may want to know how a challenged city policy changed over the course of a case. The complaint describes the policy and why it was challenged, a later court order explains how the judge ruled, and a settlement records the changes the city agreed to make. Answering the question requires evidence from all three documents.
This creates a problem for passage-level retrieval:
- The answer may require a combination of evidence from several different documents.
- Searching all passages together can return several strong matches from one document while missing another document that is essential to the answer.
This experiment focuses on the document-level retrieval step. Document maps are constructed during ingestion and used to create a subset of relevant documents to search for each question.
After that, every system uses the same passage retriever. This lets us test whether better document selection alone helps the system retrieve more of the evidence needed to answer the question accurately and completely.
What Is a Document Map?
A document map is a structured representation of what a document contains. It is created when the document is ingested and gives the retrieval system a compact view of its contents.
Each map contains:
- an overview of the document;
- entries for important facts, events, claims, decisions, and other topics;
- references from each entry back to the original source passages.
For example, a map entry might describe a specific claim in a complaint, summarize what it is about, and record which passages contain the supporting text.
kind: claim
label: [CLAIM OR ISSUE]
summary: [WHAT THE DOCUMENT SAYS ABOUT IT]
source_references:
- document_id: [DOCUMENT_ID]
segment_ids: [SEGMENT_ID]
The map therefore gives the system a higher-level view of the document while keeping every entry connected to the original evidence.
How Document Routing Works
Once the maps have been created, they can be used before passage retrieval.
For each question, a router reads the document maps for the case and selects the documents that are most likely to contain the required evidence. The passage retriever then searches only within those selected documents.
In the benchmark, the router can select up to three documents. The same passage retriever is then used to find the most relevant passages inside that subset.
The baseline skips this routing step and searches all passages in the case directly:
This gives us a controlled comparison. The passage retrieval and answer generation stages remain the same; the main difference is whether the system first uses a document-level representation to narrow the search.
Six Ways to Build a Document Map
I explored six techniques for constructing a document map during ingestion.
Each strategy starts from the same source passages and produces a document map in the same format. The main design choice is how much source material the model sees at each step, and how information from those steps is combined into the final map.
Several of these approaches have close equivalents in long-document summarization. Create-and-refine and recursive tree summarization, for example, are established ways of processing text that is too large to handle comfortably in a single call (LlamaIndex response synthesizers). The idea is that a document map acts as a compact preview of the document, helping the retrieval system decide where to search before retrieving specific passages. I selected six construction strategies to compare.
1. Stuffing
The simplest approach is to give the model the entire document in one call and ask it to create the map.
The model sees the complete document while constructing the map, so information from distant parts of the source can be considered together.
The main constraint is document size. The source text, prompt, and generated map all need to fit within the available context budget. In my implementation, documents that exceed the configured input limit are marked as context_overflow and skipped.
A second constraint is compression. The full document still has to be reduced to a much smaller representation, so some details may be omitted from the final map.
2. Map-reduce
Map-reduce builds the document map in stages.
First, each source passage is mapped independently:
The partial maps are then recursively combined until one document map remains:
In my implementation, each reduction step combines up to eight partial maps.
This keeps the amount of source text processed by each initial call bounded, which makes the method suitable for documents that cannot be handled in one pass.
A potential trade-off is information loss propagation. If a detail is omitted from an early partial map, that omission can propagate through later reduction stages because the reducer only sees the partial maps.
This pattern is closely related to recursive summarization approaches such as LlamaIndex’s tree_summarize (LlamaIndex documentation).
3. Refine
Refine also builds the map in stages, but it keeps one evolving map instead of creating separate partial maps.
The first passage creates the initial version:
Each following passage is then used to revise that map:
This allows later passages to add new information or update entries created earlier in the document.
Like map-reduce, refine can lose information as the document is compressed. Here, the risk comes from repeatedly rewriting the same map. If a detail disappears in one revision, the next step only sees the revised map, so the omission carries forward.
Refine is also sequential: each step depends on the map produced by the previous one.
LlamaIndex uses the same general idea in its refine response synthesizer, where an initial answer is updated as new chunks are processed (LlamaIndex documentation). Here, the object being updated is the document map.
4. Hierarchical mapping
Hierarchical mapping builds the document map through a tree of increasingly broad representations.
The document is first divided into small groups of neighboring passages:
The leaf maps are then combined recursively:
In my benchmark, each leaf starts from two neighboring passages, and the hierarchy uses up to three reduction levels.
The method is closely related to map-reduce. The main difference is when the first compression happens: map-reduce creates a map from each passage independently, while hierarchical mapping lets a small group of neighboring passages contribute to the same leaf map.
This gives the model more local context before information is compressed. A potential trade-off is that the map is still compressed repeatedly as it moves up the hierarchy, so information lost at one level can affect the levels above it.
RAPTOR uses a related idea, recursively building summaries at different levels of abstraction for long-document retrieval (RAPTOR paper). In this experiment, the hierarchy is used only to construct the final document map; retrieval later uses that final map rather than traversing the tree.
5. Outline-then-fill
Outline-then-fill starts by creating a high-level structure for the document.
The outline is then expanded into more detailed map entries:
The main idea is to decide how the document should be organized before filling in the details.
A potential benefit compared with the earlier map-assembling methods is consistency: related information can be grouped under the same topic instead of emerging independently across separate partial maps.
The initial outline is generated from the full document, so this method still depends on the document fitting within that request.
In my implementation, source passages are assigned to outline nodes using a simple round-robin rule: passages are dealt out to the outline nodes in turn, by position, regardless of content. This keeps passage assignment simple, but it does not guarantee that each outline node receives the most relevant passages. A stronger version would match passages to outline sections semantically.
Outline-first generation has also been explored in long-form generation, where a high-level plan is created before individual sections are expanded (Summarize, Outline, and Elaborate).
6. Agentic mapping
The final strategy, agentic mapping, builds the document map through an iterative inspect-and-revise loop.
The system starts with an empty map. After each revision, it checks which source passages are already referenced by the map and selects another uncovered passage.
In my benchmark, the process aims for full passage coverage and allows up to eight steps. If the limit is reached first, the document is recorded as only partially indexed.
The current selection policy is simple: it chooses the first uncovered passage in document order. The model revises the map, while the surrounding loop controls passage selection and stopping.
A limitation is that coverage only tracks references. A passage can be referenced by the map even if some important information from it has been omitted.
The broader idea is related to agentic RAG systems that iteratively inspect retrieved information and decide whether more evidence is needed (Microsoft Agentic RAG guide). Here, that loop is used during document-map construction rather than at query time.
Method
The benchmark uses 20 legal case files from Multi-LexSum and 60 hand-curated questions. The questions are split evenly across three levels of difficulty:
- 20 single-document questions;
- 20 multi-document questions;
- 20 chained multi-document questions.
In this benchmark, single-document questions can be answered from one document alone (single-hop). For multi-document questions, the answer depends on evidence from two or more documents in the case file (multi-hop). The chained questions go a step further. The steps build on each other, so the question follows the case through several stages, such as an allegation in the complaint, the court order that addressed it, and how that order was later tested. Answering one requires three or four pieces of evidence, in order (chained multi-hop).
Each question includes the expected answer, the documents required to answer it, the relevant passages, and supporting evidence quotes.
All seven systems answer the same 60 questions, producing 420 answer runs in total. The six document-map systems first select up to three documents and then use the same passage retriever inside those documents. The flat baseline skips document routing altogether and searches all passages in the case directly. For the flat baseline, document-level metrics use the first three distinct documents appearing in its eight retrieved passages, in retrieval order.
The shared query flow has three stages. Below is a simplified excerpt:
document_maps = _load_document_maps(index_artifact)
# 1. Route: the model reads the maps and picks the documents to search.
routing = await route_documents_from_maps(
query=item.query,
document_maps=document_maps,
services=services,
top_k=pipeline.selected_documents, # 3 in the benchmark
)
# 2. Retrieve: the same dense retriever, restricted to those documents.
retrieved = await services.retrieval_backend.search(
_retrieval_index_id(index_artifact),
item.query,
document_ids=routing.decision.selected_document_ids,
top_k=pipeline.retrieved_segments, # 8 in the benchmark
)
# 3. Answer: the shared answer model uses the retrieved passages.
answer = await answer_from_retrieved_evidence(
item=item,
retrieved_items=retrieved,
services=services,
)
For the results below, I focus on four metrics, two for each retrieval stage.
The first two measure document routing:
| Metric | What it measures |
|---|---|
| Document recall @3 | The fraction of required documents found among up to three selected documents |
| Required-document coverage | Whether every document required to answer the question was selected |
The second two measure passage retrieval inside the selected documents:
| Metric | What it measures |
|---|---|
| Context recall @4 | The fraction of required passages found in the first four retrieved passages |
| Evidence-quote recall @4 | The fraction of labeled evidence quotes found in those first four passages, after normalizing case and whitespace |
I also report context precision @4 as a supporting metric. It measures how well the first four results rank relevant passages ahead of irrelevant ones.
The answering pipeline can use up to eight retrieved passages, so I also inspect retrieval at the full eight-passage budget where useful. Beyond these, the benchmark records metrics for map validity, citations, answer quality, runtime, and model usage. The full set of metrics is available in the repository.
Results
Document maps improved document routing
The clearest improvement appeared at the document-selection stage.
With flat retrieval, 78.9% of the documents required to answer a question appeared among the first three document choices. The agentic map increased this to 93.1%.
The gain was larger on the stricter measure of finding every required document. Flat retrieval found all required documents for 58.3% of questions. Agentic routing increased this to 86.7%.
| Document recall @3 | All required documents found | |
|---|---|---|
| Flat retrieval | 78.9% | 58.3% |
| Agentic map | 93.1% | 86.7% |
| Change | +14.2 points | +28.4 points |
The same pattern appeared across most map-construction strategies.
| Strategy | Document recall @3 | All required documents found |
|---|---|---|
| Flat retrieval | 78.9% | 58.3% |
| Stuffing | 70.6% | 50.0% |
| Map-reduce | 92.2% | 81.7% |
| Refine | 90.0% | 76.7% |
| Hierarchical map | 89.4% | 75.0% |
| Outline-then-fill | 88.9% | 71.7% |
| Agentic map | 93.1% | 86.7% |
Five of the six document-map strategies improved both routing metrics over the flat baseline. Agentic mapping produced the strongest document-routing results, followed by map-reduce.
Stuffing was the exception. It created maps for only 74.4% of the source documents because some documents exceeded its configured input limit.
Passage retrieval improved much less
After document selection, the system still has to find the exact passages containing the evidence.
Here, the difference between flat retrieval and agentic routing was much smaller.
| Metric | Flat retrieval | Agentic map | Change |
|---|---|---|---|
| Context recall @4 | 53.9% | 57.8% | +3.9 points |
| Evidence-quote recall @4 | 58.2% | 60.7% | +2.5 points |
| Context precision @4 | 54.6% | 56.7% | +2.1 points |
Passage retrieval showed a small positive tendency, but the gains were modest. The routing improvement was therefore much larger than the improvement in the passages ultimately retrieved.
This pattern is also visible across the six map strategies. Their document-routing scores vary substantially, while their passage-level retrieval scores remain relatively close together.
At the full eight-passage context budget used by the answering pipeline, the gap increases somewhat: both context recall and evidence-quote recall improve by about 5.4 percentage points overall.
The difference was largest on the harder questions
The routing advantage became especially clear for the 20 chained multi-document questions.
| Metric | Flat retrieval | Agentic map |
|---|---|---|
| Document recall @3 | 61.7% | 86.7% |
| All required documents found | 15.0% | 75.0% |
Flat retrieval found every required document for only 3 of the 20 chained questions. Agentic routing did so for 15 of 20.
At the passage level, the top-four results were again mixed. Context recall @4 fell from 41.7% to 38.3%, and evidence-quote recall @4 from 47.5% to 43.3%. At the full eight-passage budget, however, the direction reversed: context recall increased from 55.8% to 65.4%, and evidence-quote recall from 64.6% to 70.8%.
The chained questions show the same split as the overall results: large routing gains, small or mixed passage-level gains.
The experiment separates two retrieval problems:
The document maps substantially improved the first step. The second step, however, was not automatically improved, because every map-based system ultimately relies on the same passage retriever once the documents have been selected. Even without changing the passage retriever, passage-level metrics improved slightly overall when it searched a more relevant subset of documents.
For this benchmark, the main finding is:
Document maps were effective document routers, especially for questions that required evidence from several sources.
Passage-level gains were present, but smaller. The improvement in passage retrieval depended on the retrieval budget, that is, how many retrieved passages the answer model receives. For the chained questions, for example, recall fell at four passages but improved at eight.
Latency and model usage
The routing gains come with additional work. Every map-based system adds a routing step at query time and requires document maps to be created during ingestion.
| Indexing tokens | Query tokens | Total tokens | Avg. query latency | |
|---|---|---|---|---|
| Flat retrieval | 0.99M | 0.88M | 1.87M | 16.6 s |
| Stuffing | 0.41M | 1.84M | 2.25M | 34.4 s |
| Map-reduce | 2.65M | 2.26M | 4.91M | 34.2 s |
| Refine | 2.97M | 1.94M | 4.92M | 32.7 s |
| Hierarchical map | 2.60M | 1.17M | 3.77M | 30.3 s |
| Outline-then-fill | 3.76M | 1.22M | 4.98M | 30.3 s |
| Agentic map | 2.07M | 1.16M | 3.23M | 30.8 s |
The flat baseline used fewer tokens and was faster. The agentic-map pipeline, which produced the strongest routing results, used 3.23M tokens across the benchmark and averaged 30.8 seconds per query, compared with 1.87M tokens and 16.6 seconds for flat retrieval.
Part of this cost is a one-off, paid during ingestion. Once a document map has been created, it can be reused across later questions about the same case file. Stuffing’s lower indexing usage should also be read together with its lower map-completion rate.
Conclusion
The benchmark suggests that document maps can improve retrieval for long, multi-document case files by giving the system a stronger document-selection step. The gains were clearest on questions that required evidence from several documents.
The results also identify the next part of the pipeline to improve. Better document selection increased the quality of the search space, but it did not consistently produce equally strong gains in the passages ranked inside that search space. Passage-level retrieval quality within the selected documents remains the main next target.
The additional latency and model usage may be a reasonable trade-off in high-stakes applications such as legal research.
Missing a document that contains required evidence can produce an incomplete answer even when the retrieved passages appear relevant.
This benchmark covers 20 legal cases and 60 hand-curated questions, with plain dense retrieval as the reference point. It is designed to add and compare more retrieval architectures, with larger collections and other document-routing approaches as the natural next comparisons. For real document collections, the right design depends on the domain, the dataset, and the retrieval failures that matter. Benchmarking and evaluation are crucial when taking retrieval systems into production.
A hybrid retrieval pipeline is one way to build on these results. Document routing can narrow the initial search, while passage retrieval inside the selected documents can be strengthened with techniques such as larger candidate sets and reranking, hybrid dense–lexical retrieval, or iterative retrieval for incomplete evidence.
A broader retrieval pipeline could look like this:
Document routing narrows the initial search, while later retrieval stages can recover additional evidence when the first pass is incomplete.
No retrieval architecture works best everywhere. The goal is to identify the right combination for the collection and task being evaluated.
Inspiration List
Show listHide list
- Towards Reliable Retrieval in RAG Systems for Large Legal Datasets — the closest legal paper to the problem in this article. It identifies document-level retrieval mismatch: retrieving semantically relevant passages from the wrong source document, and proposes adding document-level summaries to chunks to improve retrieval. ACL paper
- RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval — builds hierarchical summaries at multiple levels of abstraction and retrieves over that structure. Directly related to the hierarchical mapping strategy. RAPTOR paper
- TreeRAG: Unleashing the Power of Hierarchical Storage for Enhanced Knowledge Retrieval in Long Documents — targets long-document retrieval with tree-structured representations and bidirectional traversal. Notably, it evaluates on a law subset as well as finance and medicine. TreeRAG paper
- LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs — investigates larger retrieval units and the limitations of very fine-grained chunking. Background on why passage granularity matters. LongRAG paper
- Walking Down the Memory Maze: Beyond Context Limit through Interactive Reading (MemWalker) — creates a hierarchy of summaries over long text and navigates it iteratively to answer questions. Closely related to the idea of creating a navigable representation before accessing source evidence. MemWalker paper
- LlamaIndex Response Synthesizers — reference implementations of the sequential
refineand recursivetree_summarizepatterns, which informed the refine, map-reduce and hierarchical strategies here. LlamaIndex documentation - A2JRAG: Process-Aware Retrieval-Augmented Generation for Public Legal Information Systems — conditions retrieval on the user’s procedural stage so that legally correct information from the wrong stage is not treated as relevant. It offers a related example of using structured procedural context to guide legal retrieval. A2JRAG paper
- LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning — legal-specific RAG using a hierarchical representation of cases, statutes, interpretations, facts and rules, with separate retrieval, verification and reasoning stages. Broader in scope than document maps, but relevant to moving beyond flat legal-document retrieval. LegalGraphRAG paper
- Legal RAG Bench: an end-to-end benchmark for legal RAG — separates retrieval failures from generation and reasoning failures, and finds retrieval to be a major driver of end-to-end legal RAG performance — the same split this benchmark relies on. Legal RAG Bench
- Multi-LexSum: Real-world Summaries of Civil Rights Lawsuits at Multiple Granularities — the multi-document legal dataset used in this benchmark. Multi-LexSum paper