GraphRAG: Unlocking Enterprise Knowledge with Knowledge Graphs
By Alexandre Achten and Thibaud Vanmechelen
1. Introduction & Context
Bringing Large Language Models (LLMs) into production often hits a snag when dealing with company-specific knowledge. LLMs are great at general text tasks, but without direct access to internal context, they can hallucinate or miss critical operational details.
Standard Retrieval-Augmented Generation (RAG) was the first attempt to fix this. It chops documents into text chunks, turns them into vector embeddings, and stores them in a vector database. When a user asks a question, the system finds similar text chunks and sends them to the LLM. While this works well for straightforward lookups, standard RAG relies purely on text similarity in a flat vector space. It struggles when answering complex questions that require connecting dots across multiple files, teams, or processes.
Consider a real-world enterprise supply chain scenario where an operational manager asks: "If our main parts supplier delays their next shipment, which of our top-tier customer orders will be impacted first?"
Answering this requires connecting a chain of facts: finding the supplier, reading contract terms, mapping delayed component numbers to inventory levels across warehouses, checking assembly schedules, and filtering affected purchase orders by customer tier. A standard vector RAG system might grab chunks mentioning "supplier delay" or "customer order," but because vector search lacks structural links between data points, it can't trace multi-step relationships reliably. The result is often an incomplete or incorrect answer.
GraphRAG fixes this gap by replacing isolated text chunks with a Knowledge Graph, where entities become nodes and relationships become edges. By structuring corporate knowledge as a relational network, GraphRAG helps LLMs follow connections directly, reason across multiple documents, and deliver accurate, fully grounded answers.
2. Knowledge Graph Construction
Turning unstructured files like PDFs, technical specs, and legal contracts—into a graph database requires an automated ingestion pipeline.
Our document pipeline starts by ingesting raw PDFs and parsing them into Markdown with Docling. Next, LangChain splitters break the text into semantic chunks, and an LLM extracts entities and relationships. We then post-process the data to merge duplicates and resolve references before saving everything into a graph database like Neo4j.
When building a Knowledge Graph pipeline, you generally choose between two main design approaches:
- Schema-Constrained (Ontology-Driven): Defines a strict schema ahead of time. A fixed domain ontology specifies allowed entity types and relationship labels during extraction, preventing the LLM from creating redundant or messy labels.
- Schema-Free (Ontology-Free): Extracts entities and relationships dynamically without upfront constraints. While flexible, this can generate thousands of slightly different labels, so the pipeline uses caching, iterative extraction, and clustering algorithms to clean up the graph afterwards.
From a cost standpoint, graph construction volume scales directly with LLM token usage. Defining a clear domain ontology up front keeps prompts concise, which substantially cuts down processing costs:
| Construction Methodology | Token Pricing | Est. Per-Page Cost (~1k tokens) | Operational & Financial Impact |
|---|---|---|---|
| Schema-Free (Ontology-Free) | $16.00 – $17.00 / 1M tokens | ~$0.016 per page | Higher cost due to open-ended extraction and extensive post-processing clustering. |
| Schema-Constrained (Ontology-Driven) | $8.00 / 1M tokens | ~$0.008 per page | 50% Lower Processing Cost thanks to structured prompt constraints and lower token overhead. |
Using a schema-constrained ontology allows teams to cut ingestion costs in half while keeping the graph clean and consistent.
3. Retrieval Strategies
With the graph in place, the core focus shifts to retrieval: pulling the exact relevant subgraph for a user's question without cluttering the prompt context with noise.
We evaluated four different retrieval architectures across relational graph topologies:
Random Baseline: Selects a random sample of nodes and edges. This acts as a control to measure chance performance and highlight gains from smarter methods.
Graph Traversal (Text2Cypher): Uses an LLM (like GPT-4o) to turn user queries into schema-aware Cypher code for Neo4j. It includes a self-correction loop where database syntax errors are sent back to the LLM to refine the query.
Semantic Agent: Combines vector search with graph traversal. It first finds seed nodes via semantic search, then explores neighbouring connections to collect structural context.
Chain-of-Thought (CoT) Agent: Builds on the semantic agent by giving the LLM specialised graph utilities. The agent can count neighbours, find shared connections, and perform targeted lookups, step by step.
4. Benchmarking & Quality Evaluation of Retrieval
A major drawback of standard GraphRAG evaluations (like RAGAS, ARES, or subjective LLM judges) is that they only score the final generated text. This leads to diagnostic conflation: it’s impossible to tell whether an error came from bad context retrieval or a generation mistake by the LLM. Standard text metrics are also topologically blind, meaning they can't check if the retriever followed the correct database paths.
To fix this, we created an automated evaluation framework that measures retrieval quality directly on the graph structure itself.
The framework generates test datasets by populating Cypher query templates with real database entities, creating verifiable ground-truth targets. An LLM then translates these structured queries into natural language test questions across different phrasing styles to evaluate robustness.
To guarantee thorough performance tracking, the benchmark categorises queries across nine operational query types:
Direct Node & Edge Retrieval: Tests basic entity and relationship lookups.
Aggregation Node & Aggregation Edge: Evaluate quantitative counting and filtering operations.
Negative Node & Edge: Tests whether the retriever incorrectly pulls data when no valid link exists.
Out-of-Scope Queries: Test system stability when confronted with questions entirely unrelated to the corporate database.
Multi-Hop Path & Intersection Queries: Tests multi-step reasoning across connected paths and converging nodes.
We tested this on a biochemical Knowledge Graph from Hetionet containing 38,584 nodes and 1,488,879 edges.
| Architecture | Precision | Recall | F1 Score | Path Hit Rate | ROUGE-L | Avg. Time | Avg. Calls |
|---|---|---|---|---|---|---|---|
| Random Baseline | 0.10 | 0.56 | 0.10 | 0.00 | 0.00 | 2.68 s | 0.00 |
| Graph Traversal (Text2Cypher) | 0.57 | 0.73 | 0.56 | 0.53 | 0.46 | 4.26 s | 1.08 |
| Semantic Agent | 0.52 | 0.81 | 0.55 | 0.81 | 0.62 | 6.89 s | 2.10 |
| Chain-of-Thought (CoT) Agent | 0.60 | 0.89 | 0.61 | 0.95 | 0.72 | 7.56 s | 4.02 |
Results show that the Chain-of-Thought (CoT) Agent performed best on graph retrieval, reaching a 0.89 Recall and 0.95 Path Hit Rate. This demonstrates that giving agents specialized graph tools significantly improves their navigation accuracy.
The benchmarks also provided an interesting insight into precision. Agentic retrievers often show lower strict precision because they fetch nearby context entities alongside direct edges. However, this extra context actually helps the downstream LLM write better answers. As a result, Recall and Path Hit Rate are better indicators of real-world retrieval quality. Furthermore, query phrasing had minimal impact on performance, showing that structural complexity is what primarily drives GraphRAG challenge.
5. Conclusion
Building production-grade GraphRAG systems relies on combining efficient, ontology-driven graph construction with agentic retrievers designed for multi-hop navigation. By evaluating retrieval deterministically, teams can identify bottlenecks, reduce hallucinations, and ensure reliable answers.
Key takeaways and adoption guidelines for engineering teams include:
De-Risking Ingestion with Schemas: Schema-constrained architectures that reduce token usage by 50% while generating clean graph databases.
Optimized Retrievers: Tool-equipped agentic retrievers built to handle multi-hop relational queries across connected enterprise data.
Deterministic Benchmarking: Dual-layer evaluation frameworks that isolate retrieval performance from text generation, giving clear, measurable insights into system accuracy.
Should you need assistance with implementing GraphRAG systems in your organization, feel free to reach out to our team.