In our previous blogs, we learned about:
- Chunking
- Embeddings
- Vector Databases
- Semantic Search
- Similarity Search
We learned how AI finds information that is relevant to a user's question.
But one important question remains:
After finding the relevant information, how does the system actually collect it and use it to generate an answer?
This process is called Retrieval.
Let's understand Retrieval in RAG with a simple real-world example.
What Is Retrieval?
Retrieval is the process of finding and collecting relevant information from stored data based on a user's query.
In a RAG system, retrieval usually means selecting the most useful document chunks from a vector database and sending them as context to the LLM.
Simple definition:
Retrieval is the process of bringing the most relevant information from a knowledge source for answering a question.
Real-World Example: College Student Handbook
Imagine a college has a large student handbook.
It contains information about:
- Attendance
- Exams
- Hostel
- Library
- Fees
- Leave Rules
The handbook is divided into smaller chunks and stored in a vector database.
Now a student asks:
"How much attendance do I need to write my final exam?"
The system needs to find the correct information from the handbook.
How Retrieval Works Step by Step
Step 1: The User Asks a Question
The student asks:
"How much attendance do I need to write my final exam?"
This is called the user query.
Step 2: The Query Is Converted Into an Embedding
The question is converted into a numerical representation called an embedding.
User Question
↓
Query Embedding
This helps the system compare the question with stored document embeddings.
Step 3: Similarity Search Finds Relevant Chunks
The system compares the query embedding with the stored embeddings in the vector database.
For example:
Attendance Rule → Highly Relevant
Exam Rule → Relevant
Hostel Rule → Not Relevant
Library Rule → Not Relevant
Fee Rule → Not Relevant
Similarity Search helps identify which chunks are closest to the user's question.
Step 4: Relevant Chunks Are Retrieved
Now the system selects the useful information.
For example, it retrieves:
"Students must maintain a minimum of 75% attendance to appear for the final examination."
This act of selecting and collecting the relevant chunk is called Retrieval.
Similarity Search vs Retrieval
These two concepts are closely connected, but they have different roles.
Similarity Search
Similarity Search helps answer:
"Which stored information is most similar to the user's question?"
Retrieval
Retrieval helps answer:
"Which relevant information should be collected and passed to the next stage?"
In many RAG systems, similarity search is one of the techniques used during the retrieval process.
What Is Retrieved Information?
The information retrieved from the knowledge source is usually called a:
- Relevant chunk
- Retrieved document
- Retrieved context
- Search result
For example:
User Question:
"How much attendance is required?"
Retrieved Chunk:
"Students must maintain 75% attendance..."
This retrieved chunk contains the information needed to answer the question.
What Is Top-K Retrieval?
Sometimes the system does not retrieve only one result.
It retrieves the top few most relevant results.
This is called Top-K Retrieval.
Here, K means the number of results to retrieve.
For example:
Top-1 → Retrieve 1 result
Top-3 → Retrieve 3 results
Top-5 → Retrieve 5 results
For the college question, Top-3 Retrieval might return:
- Attendance Rule
- Exam Rule
- Examination Eligibility Rule
The system can then use these results to prepare a better answer.
Why Not Retrieve Every Document?
Imagine the college handbook contains 500 pages.
If the system sends all 500 pages to the LLM, it would be:
- Slow
- Expensive
- Unnecessary
- Difficult for the LLM to process
- More likely to include unrelated information
Instead, Retrieval selects only the relevant sections.
500 Pages
↓
Retrieve Relevant Chunks
↓
Only 2–5 Useful Chunks
↓
LLM
This makes the system more efficient.
Retrieval in the RAG Pipeline
Let's connect Retrieval with the complete RAG process.
Retrieval is the bridge between the stored knowledge and the LLM.
What Happens After Retrieval?
After the relevant chunks are retrieved, they are added to the prompt as context.
For example:
User Question:
"How much attendance do I need?"
Retrieved Context:
"Students must maintain a minimum of 75% attendance..."
LLM:
Uses the question + retrieved context
↓
Generates the answer
The LLM can now answer using the actual information from the college handbook.
Another Real-World Example: Company HR Policy
Imagine an employee asks:
"How many days of maternity leave are available?"
A company RAG system may search its HR policy documents.
It retrieves the relevant leave policy section instead of sending every company document to the LLM.
Employee Question
↓
Search HR Documents
↓
Retrieve Leave Policy
↓
Send Relevant Context to LLM
↓
Generate Answer
This is Retrieval in a practical business application.
Why Is Retrieval Important in RAG?
Retrieval is important because an LLM may not know the latest or private information stored in a company's documents.
For example:
- College rules
- Company HR policies
- Product manuals
- Internal documents
- Customer support information
- Legal or financial documents
Retrieval brings the required information from these sources before the LLM generates the answer.
Retrieval Does Not Generate the Final Answer
This is an important point.
Retrieval only finds and collects the relevant information.
It does not explain the answer in natural language.
Retrieval
↓
Finds relevant information
LLM
↓
Understands the context
and generates the answer
So, Retrieval and Generation are two different stages.
Retrieval vs Generation
Retrieval
Finds information from stored knowledge.
Generation
Creates a natural-language answer using the retrieved information.
Retrieval → Finds information
Generation → Creates the answer
That is why RAG stands for:
Retrieval-Augmented Generation
- Retrieval → Find relevant information
- Augmented → Add that information to the prompt
- Generation → LLM generates the answer
Complete Example
Let's see the complete process using the college handbook.
User Question
"Can I write my final exam with 70% attendance?"
System Process
1. User asks the question
↓
2. Question becomes an embedding
↓
3. Similarity Search checks stored chunks
↓
4. Attendance rule is identified
↓
5. Relevant chunk is retrieved
↓
6. Retrieved chunk is added as context
↓
7. LLM reads the context
↓
8. LLM generates the answer
Final Answer
"According to the college policy, students need a minimum of 75% attendance to appear for the final examination."
The answer is based on the retrieved college information.
Simple Retrieval Flow
User Question
↓
Query Embedding
↓
Similarity Search
↓
Relevant Chunks Selected
↓
Retrieved Context
↓
LLM
↓
Answer
Key Points to Remember
- Retrieval means finding and collecting relevant information.
- Similarity Search helps identify the closest matching chunks.
- Retrieval is an important stage in RAG.
- Only useful information is passed to the LLM.
- Retrieval does not generate the final answer.
- The LLM uses the retrieved context to create the response.
- Top-K Retrieval means selecting the top few relevant results.
Conclusion
Retrieval is one of the most important steps in a RAG system.
It helps the AI system find the right information from a large collection of documents and provide it to the LLM.
Instead of asking the LLM to answer only from its memory, RAG first retrieves relevant information and then generates the answer using that context.
The simple idea to remember is:
Retrieval = Find and collect the right information before generating the answer.
The complete flow is:
Question
↓
Search
↓
Retrieve Relevant Information
↓
Add Context
↓
LLM
↓
Answer
In our next blog, we can explore:






No comments:
Post a Comment