GraphRAG: does building a knowledge graph fix retrieval?
Microsoft's GraphRAG extracts an entity graph from a corpus, clusters it into communities, summarises each community, and answers global questions by map-reduce over the summaries. A method deep dive on what it wins, what the index costs, and why the evaluation leaves the central question open.
The question ordinary RAG cannot answer
Vector retrieval answers questions whose answer lives in a few chunks. It has no good response to a question like what are the main themes across this corpus, because the answer is spread over everything and no chunk is relevant by itself. Darren Edge and colleagues at Microsoft call these global sensemaking questions, and their April paper proposes to answer them by building a structure over the corpus in advance and querying the structure rather than the text. The open source implementation is now on GitHub, which is why the method is being tried widely.
The pipeline has more steps than a summary usually admits. Documents are chunked, and the chunk size matters. At 600 tokens the extraction step finds about twice as many entity references as at 2,400 tokens. A language model then reads each chunk and emits entities with names, types and descriptions, plus relationships between them, with optional gleaning rounds where the model is asked whether it missed anything and forced to answer yes or no by logit bias. Duplicate references are merged into one description per element. The Leiden algorithm partitions the resulting graph into a hierarchy of communities, and the model writes a summary for each community at each level. At query time the summaries at a chosen level are shuffled into batches, each batch produces a partial answer with a helpfulness score from zero to 100, and the top-scoring partials are combined into the final answer.
What the evaluation shows
Two corpora were used, a set of podcast transcripts of about one million tokens in 1,669 chunks, and a news collection from the MultiHop-RAG benchmark of about 1.7 million tokens in 3,197 chunks. The podcast graph has 8,564 nodes and 20,691 edges and the news graph 15,754 nodes and 19,520 edges. Questions were generated by prompting a model to imagine five user personas with five tasks each and five questions per task, giving 125 questions per corpus. Answers were judged head to head by GPT-4-turbo on comprehensiveness, diversity, empowerment and directness.
Against naive vector RAG, the graph conditions win comprehensiveness 72 to 83 percent of the time on podcasts and 72 to 80 percent on news, and diversity 75 to 82 and 62 to 71 percent. Naive RAG wins directness, which the authors expected, since short answers are more direct. The more informative comparison is against a baseline that skips the graph entirely and map-reduces over the raw source text chunks. There the intermediate community levels win, but narrowly. C2 on podcasts takes 57 percent on both comprehensiveness and diversity, and C3 on news takes 64 and 60 percent. Empowerment is mixed across the board.
What the index buys and what it costs
The token cost table is the strongest argument in the paper. Answering a question from the root-level communities uses 26,657 tokens on podcasts and 39,770 on news, about 2.6 and 2.3 percent of the tokens needed to map-reduce over the source text. The intermediate levels use 22 to 57 percent. The leaf level uses 67 to 74 percent. So the graph lets you trade answer quality for query cost along a dial, and the root level gives you a cheap answer that is still competitive.
The cost that table does not show is building the index. Every chunk goes through the extraction model, plus gleaning passes, plus a summary per element, plus a summary per community per level. The paper does not report an indexing cost figure, and the GitHub README opens with a warning that indexing can be an expensive operation and that users should read the documentation and start small. For a corpus you query once, the index will never pay for itself. For a corpus you query thousands of times with global questions, the arithmetic changes. The README's own advice to start small is the right advice, and it is easy to skip.
Why the central question is still open
The paper's own limitations section is candid. The evaluation covers one class of question on two corpora of about a million tokens, the questions and the judgments were both produced by language models with no end-user validation, and there is no hallucination analysis. The gap that bothers us most is the one the authors name last. They do not know whether the graph index beats the source-text map-reduce by enough to justify its cost, because a 57 percent win rate on a judge whose own reliability is unmeasured is a small effect sitting on an uncertain instrument.
We have not seen a controlled replication on a different corpus that reports both the win rate against a source-text baseline and the indexing cost, and until one exists the honest summary is that the method wins the corpus-overview question at a query cost it can dial down, at a build cost nobody has published.
What we would test
The experiment we want is the one the paper leaves out. Hold the query budget fixed, run the source-text map-reduce with the same number of tokens the community summaries use, and see whether the graph still wins. If it does, the structure is contributing. If it does not, the win came from summarisation and the graph is an expensive way to decide what to summarise.
We would also like a corpus with human-written global questions and human judgments, even a small one. A hundred questions judged by three people would settle more than 250 judged by GPT-4-turbo, because the paper's whole claim is about a kind of question that models were not good at answering, and we are not sure they are good at judging it either.
Sources
From the foundation