In our previous blogs, we learned how RAG works step by step.
We learned about:
- Chunking
- Embeddings
- Vector Database
- Semantic Search
- Similarity Search
- Retrieval
We now know that Retrieval helps find the relevant information from a large collection of documents.
But what happens after that?
The relevant information is given to the LLM so that it can use that information to generate an answer.
This relevant information is called Context.
Let's understand what context means in RAG with a simple real-world example.
What Is Context?
Context is the relevant information provided to an AI model to help it understand and answer a user's question.
In RAG, context usually comes from information retrieved from external documents or a knowledge base.
Simple definition:
Context is the useful information given to the LLM along with the user's question so it can generate a better answer.
Real-World Example: College Student Handbook
Imagine a college has a large student handbook.
It contains information about:
- Attendance
- Exams
- Hostel
- Library
- Fees
- Leave Rules
Now a student asks:
"Can I write my final exam if my attendance is 70%?"
The AI system needs to find the correct information.
Step 1: User Asks a Question
The student asks:
"Can I write my final exam if my attendance is 70%?"
This is the user query.
User Question
↓
"Can I write my final exam
if my attendance is 70%?"
Step 2: The System Searches for Relevant Information
The question is converted into an embedding.
The system then uses similarity-based search to find relevant information from the vector database.
It may find:
"Students must maintain a minimum of 75% attendance to appear for the final examination."
This information is relevant to the student's question.
Step 3: Retrieved Information Becomes Context
The relevant information retrieved from the college handbook is provided to the LLM.
This information acts as context.
User Question
↓
Retrieved Information
↓
Context
↓
LLM
So the LLM receives something like:
Question:
Can I write my final exam if my attendance is 70%?
Context:
Students must maintain a minimum of 75% attendance to appear for the final examination.
Now the LLM has the information it needs to formulate an answer.
Why Does an LLM Need Context?
An LLM has learned from large amounts of data during training.
But it may not have access to your specific or current private information.
For example, an LLM may not automatically know:
- Your college's latest attendance policy
- Your company's internal HR rules
- A product's latest internal manual
- A private organization's documents
RAG helps by retrieving the relevant information and providing it as context.
RAG Without Context
Imagine asking:
"What is my college's attendance requirement?"
Without access to the college handbook, the AI may not know the specific rule.
It could potentially give a generic answer that does not match your college's actual policy.
RAG With Context
Now imagine the system retrieves this information:
"Students must maintain a minimum of 75% attendance to appear for the final examination."
This information is provided to the LLM as context.
The LLM can then use that information to formulate the answer.
Question
+
Relevant Context
↓
LLM
↓
Answer
This is the basic idea behind Retrieval-Augmented Generation.
What Is the Difference Between Context and Query?
These two terms can sometimes be confusing.
Query
The query is what the user asks.
Example:
"Can I write my exam with 70% attendance?"
Context
The context is the relevant information provided to help answer the query.
Example:
"Students need a minimum of 75% attendance to appear for the final examination."
So:
Query → What the user wants to know
Context → Information that helps answer it
How Is Context Created in RAG?
Context in RAG generally comes from the retrieved information.
The basic process is:
Documents
↓
Chunking
↓
Embeddings
↓
Vector Database
↓
User Query
↓
Similarity Search
↓
Retrieval
↓
Relevant Chunks
↓
Context
↓
LLM
↓
Answer
The retrieved chunks become the information the LLM can use as context.
Can There Be More Than One Context?
Yes.
Sometimes one chunk may not contain everything needed to answer a question.
The system may retrieve multiple relevant chunks.
For example, a student asks:
"What attendance do I need, and what happens if my attendance is below the requirement?"
The system might retrieve:
Chunk 1:
Students need a minimum of 75% attendance.
Chunk 2:
Students below the required attendance may need to follow the college's eligibility or condonation policy.
Both pieces of information can be provided to the LLM as context.
Relevant Chunk 1
+
Relevant Chunk 2
↓
Context
↓
LLM
Too Little Context
Context needs to be relevant and sufficient.
Suppose the user asks:
"What happens if my attendance is below 75%?"
But the system retrieves only:
"Students must maintain 75% attendance."
This may not provide enough information to answer the complete question.
The system may need another relevant chunk containing the consequences or applicable policy.
Too Much Unnecessary Context
On the other hand, imagine giving the LLM hundreds of unrelated pages.
For example:
Attendance Rule
Hostel Rule
Library Rule
Fee Rule
Bus Rule
Canteen Rule
Sports Rule
...
500 pages
Most of this information is unrelated to the question.
A good RAG system tries to retrieve relevant context instead of simply giving everything to the LLM.
Large Document Collection
↓
Relevant Retrieval
↓
Useful Context
↓
LLM
Context in a Real-World Company
Let's take another example.
Imagine an employee asks:
"How many days of leave can I take?"
The company has hundreds of documents.
The RAG system searches the company's HR knowledge base and retrieves the relevant leave policy.
For example:
"Employees are eligible for 18 days of annual leave per year."
That information becomes context for the LLM.
Employee Question
↓
Retrieve HR Policy
↓
Relevant Context
↓
LLM
↓
Answer
The same idea can be used for:
- Customer support
- Product manuals
- Company policies
- College documents
- Research documents
- Internal knowledge bases
Context Does Not Mean the LLM Is Retrained
This is an important point.
When we provide retrieved information as context, we are not training the LLM again.
The information is simply provided to the model while it is answering the current question.
Think of it like this:
Training
The model learns patterns from training data.
RAG Context
The model receives relevant information for a particular question.
Training
↓
Model learns
RAG Context
↓
Model gets useful information for the current question
So RAG can provide external information without retraining the entire model for every new document.
Context and LLM: Simple Example
Let's imagine the LLM is like a student taking an exam.
The student knows many things already.
But the teacher gives the student a relevant page from a textbook.
That page acts like context.
The student can use that information to answer the question.
Similarly:
User Question
+
Relevant Context
↓
LLM
↓
Generated Answer
Where Does Context Fit in the RAG Pipeline?
Let's connect everything we have learned so far.
Documents
↓
Chunking
↓
Embeddings
↓
Vector Database
↓
User Question
↓
Query Embedding
↓
Similarity Search
↓
Retrieval
↓
Relevant Context
↓
LLM
↓
Generated Answer
Now the complete flow becomes much easier to understand.
Why Is Good Context Important?
Good context should be:
Relevant
It should relate to the user's question.
Useful
It should contain information that helps answer the question.
Focused
It should avoid unnecessary information whenever possible.
For example, if the question is about attendance, the system should prioritize attendance-related information instead of unrelated hostel or library rules.
Simple Everyday Example
Imagine you ask your friend:
"When is my doctor's appointment?"
Your friend checks your calendar and tells you:
"Your appointment is tomorrow at 10 AM."
The calendar information is the context your friend used to answer your question.
Similarly, in RAG:
Question
↓
Find Relevant Information
↓
Give Information to LLM
↓
Generate Answer
The relevant information helps the AI answer the question.
Context vs Knowledge
Context and knowledge are not exactly the same thing.
Knowledge
Information the model has learned during training.
Context
Information provided to the model while answering a particular question.
For example:
Model's Learned Knowledge
+
Retrieved Context
↓
LLM
↓
Answer
RAG mainly helps bring external or specific information into the current context.
One Complete RAG Example
Let's put everything together.
User Question
"Can I write my final exam if my attendance is 70%?"
Search
The system searches the college knowledge base.
Retrieval
It retrieves:
"Students must maintain a minimum of 75% attendance to appear for the final examination."
Context
The retrieved information becomes context.
Generation
The LLM uses the question and context to generate an answer.
User Question
↓
Query Embedding
↓
Similarity Search
↓
Retrieval
↓
Relevant Information
↓
Context
↓
LLM
↓
Answer
Simple Recap
Let's quickly remember what we learned.
What is Context?
Relevant information provided to the LLM to help answer a question.
Where does RAG context come from?
Usually from information retrieved from external documents or a knowledge base.
Is the user question the context?
No.
The user question is the query.
The relevant supporting information is the context.
Does providing context retrain the LLM?
No.
The information is provided to the model while generating the current response.
Why is good context important?
Relevant and focused context helps the LLM use the right information when generating an answer.
Conclusion
Context is a very important part of RAG.
We can think of it as the useful information that is brought to the LLM at the right time.
The overall idea is simple:
User asks → Relevant information is retrieved → Retrieved information becomes context → LLM uses the context → Answer is generated.
So the RAG process we have learned so far looks like this:
Documents
↓
Chunking
↓
Embeddings
↓
Vector Database
↓
User Question
↓
Similarity Search
↓
Retrieval
↓
Context
↓
LLM
↓
Answer
Now we have understood how RAG finds information and gives that information to the LLM.
In the next blog, we can explore an important question:



No comments:
Post a Comment