PERSONAL RESEARCH / RAG & retrieval
Choosing retrieval.
For the question.
RAG, agentic search and GraphRAG solve different retrieval problems. I explore how to choose between them using the questions, sources and failure modes that matter.
- Article date
- Base notes dated
- Sources checked
01RESEARCH NOTE
Start with the question
Retrieval-augmented generation brings external evidence into a model’s working context. Vector search is one implementation; an agent can also search keywords, inspect document structure and read relevant passages through tools. These approaches can be combined. [01] [02]
For my curated Markdown notes, I treat agentic file search as a useful baseline. It lets an agent refine a query and inspect the original text without maintaining an embedding index. Subramanian and colleagues found that tool-based keyword search approached conventional RAG performance on the evaluated tasks. That supports testing a simpler approach, with results checked on the actual collection. [01]
I would compare approaches against representative questions: exact lookups, terminology mismatches, questions spanning several sources and questions that require a current structured record.
02RESEARCH NOTE
Improve the evidence before adding complexity
A relevant passage can exist in the source collection and still fail to reach the model. This is a retrieval failure. It differs from an answer misusing evidence that was successfully retrieved.
My first checks would be the query, document boundaries, missing surrounding context and ranking. Hybrid retrieval combines lexical and semantic matching; reranking can help select relevant candidates. Anthropic’s Contextual Retrieval work also tests adding document-specific context to chunks before indexing. Its reported improvements belong to the evaluated setup. [02]
I want to measure both evidence retrieval and the resulting answer. More passages help when they improve correctness, completeness or traceability. Latency and maintenance matter too, especially for a voice interface.
03RESEARCH NOTE
Use graphs when the relationships matter
GraphRAG becomes interesting when questions depend on relationships or broader themes across documents. Microsoft’s original approach builds an entity graph and community summaries to support questions about the collection as a whole. [03]
GraphRAG-Bench offers a useful qualification: graph methods showed advantages on some complex reasoning and summarization tasks, while basic RAG remained competitive on simple fact retrieval. [04]
My interpretation is to introduce a graph when representative questions demonstrate that relationships are the missing capability. The graph also needs to stay aligned with its sources. For frequently edited notes, I would first test clearer structure, better retrieval and an iterative search loop.
This is a working architectural position. I would revisit it when the questions, collection or measured failure patterns change.
REFERENCES / 04
Sources.
This article builds on dated personal research notes. The source check covers the cited claims and links, rather than an exhaustive literature review. Document dates and any access limits appear below.
- Subramanian et al. — Keyword search is all you need (opens in a new tab)
AAAI 2026; agentic keyword search evaluated against RAG baselines.
- Anthropic — Introducing Contextual Retrieval (opens in a new tab)
Contextual embeddings, contextual BM25 and reranking; vendor experiment.
- Edge et al. — From Local to Global: A Graph RAG Approach to Query-Focused Summarization (opens in a new tab)
First submission; revision 19 February 2025. Entity graphs and community summaries.
- Xiang et al. — When to use Graphs in RAG (opens in a new tab)
GraphRAG-Bench; version 3, 22 February 2026. Results vary by task category.