Tuesday, 29 September 2026

What Is Reranking in RAG? A Simple Guide for Beginners

We have already learned how Retrieval, Semantic Search, Similarity Search, and Hybrid Search work in RAG.

But there is one important question:

What if the search system finds many relevant results?

Which result should come first?

Which information is actually the most relevant to the user's question?

This is where Reranking can help.


What Is Reranking?

Reranking is the process of reordering the retrieved search results so that the most relevant results appear at the top.

In simple words:

Search finds possible answers. Reranking helps put the most relevant answers first.

Think of it like this:

Search → Find relevant results

Reranking → Reorder those results by relevance


A Simple Real-World Example

Let's use our familiar college handbook example.

Imagine a college has a large handbook containing information about:

  • Attendance
  • Exams
  • Fees
  • Hostel
  • Library
  • Leave
  • Scholarships

Now a student asks:

“Can I write my final exam if my attendance is 70%?”

The search system looks through the college documents.

It may find several potentially related sections:

Result 1

Attendance Policy

Students must maintain a minimum of 75% attendance to appear for the final examination.

Result 2

Leave Policy

Students can apply for leave under certain conditions.

Result 3

Examination Rules

Students must complete the required examination registration before the final examination.

Result 4

Hostel Rules

Students must follow the college hostel timings.

All these documents may be related to college rules.

But which one is most relevant to the student's question?

Clearly, the Attendance Policy contains the most directly relevant information.

This is where reranking can help.


Before Reranking

The initial search might return results in an order like:

1. Examination Rules

2. Leave Policy

3. Attendance Policy

4. Hostel Rules

The relevant attendance rule is there, but it is not at the top.


After Reranking

A reranking step can reconsider the retrieved results in relation to the user's question.

The order might become:

1. Attendance Policy ⭐

2. Examination Rules

3. Leave Policy

4. Hostel Rules

Now the most relevant information is at the top.


Why Do We Need Reranking?

Search systems are designed to find potentially relevant information.

But finding something that is somewhat related is not always the same as finding the most useful result.

For example, a query about:

“Attendance required for final exam”

could retrieve documents about:

  • Attendance
  • Exams
  • Leave
  • Academic rules
  • Student registration

Several may be related.

But the system needs to identify which result answers the question most directly.

Reranking helps improve this ordering.


How Does Reranking Work?

Let's keep it simple.

A typical process can look like this:

Step 1: User asks a question

“Can I write my final exam with 70% attendance?”

↓

Step 2: Search finds possible results

The system retrieves several potentially relevant chunks.

↓

Step 3: Reranker looks at the query and retrieved results

It examines how relevant each result is to the question.

↓

Step 4: Results are reordered

The most relevant results move to the top.

↓

Step 5: Top results are used as context

The selected information can then be provided to the LLM.

↓

Step 6: LLM generates the answer


Simple Reranking Flow

User Query

↓

Initial Search

↓

Retrieved Results

↓

Reranking

↓

Most Relevant Results

↓

Context

↓

LLM

↓

Answer


Where Does Reranking Come in RAG?

Let's connect it with what we already learned.

A simplified RAG pipeline can look like:

Documents

↓

Chunking

↓

Embeddings

↓

Vector Database

↓

User Query

↓

Search

↓

Initial Results

↓

Reranking

↓

Relevant Results

↓

Context

↓

LLM

↓

Answer

Reranking happens after the initial retrieval/search and before the final context is given to the LLM.

The exact architecture can vary depending on the RAG system.


What Is Initial Retrieval?

Before reranking can happen, the system first needs to find some possible results.

This is called initial retrieval.

For example, imagine there are 10,000 chunks in a knowledge base.

It would not be practical to deeply evaluate every chunk for every question.

So the search system can first retrieve a smaller group of candidates.

For example:

10,000 chunks

↓

Initial Search

↓

Top 20 candidates

↓

Reranking

↓

Top 5 most relevant results

↓

LLM

This is a simplified example, but it explains the basic idea.


What Is Top-K Retrieval?

We previously learned about Top-K Retrieval.

Here, K means the number of results we want to retrieve.

For example:

Top-3 → Retrieve 3 results

Top-5 → Retrieve 5 results

Top-10 → Retrieve 10 results

Reranking can then reorder those retrieved results.

Example

Initial search:

Top 5 results

↓

Reranking

↓

Best 3 results

↓

Context for LLM

So retrieval and reranking can work together.


Retrieval vs Reranking

These two concepts are closely related, but they are not the same.

Retrieval asks:

“Which information should we bring back?”

Reranking asks:

“Among the retrieved information, which results are more relevant and should come first?”

A simple way to remember:

Retrieval = Find candidates

Reranking = Reorder candidates


Reranking vs Similarity Search

We have also learned about Similarity Search.

Similarity search compares representations, often embeddings, to find information that is similar to the query.

For example:

“Can I attend the exam with low attendance?”

may be similar in meaning to:

“Students need 75% attendance to appear for the final examination.”

Reranking is a further relevance step that can reconsider the retrieved candidates and put the strongest matches higher in the result list.

So:

Similarity Search → Find similar candidates

Reranking → Reorder candidates based on relevance


Reranking vs Hybrid Search

We recently learned about Hybrid Search.

Hybrid Search combines approaches such as:

Keyword Search + Semantic Search

This can produce an initial set of results.

Then reranking can be applied to those results.

Example

User Query

↓

Keyword Search + Semantic Search

↓

Initial Results

↓

Reranking

↓

Best Results

This means Hybrid Search and Reranking can work together.


Another Real-World Example: Online Shopping

Imagine you search for:

“comfortable running shoes for daily use”

An online shopping system might find many products related to:

  • Running shoes
  • Sports shoes
  • Walking shoes
  • Training shoes
  • Casual shoes

Several products may be relevant.

But some may match the user's query more closely than others.

A ranking or reranking process can help place the more relevant products higher.

The same basic idea is useful in information retrieval systems.


Another Example: Customer Support

Imagine a customer asks:

“My payment was successful but my order is still showing as pending. What should I do?”

The company's knowledge base might contain articles about:

  • Payment failed
  • Payment pending
  • Refunds
  • Order cancellation
  • Order tracking

Several articles may contain related words.

But the article about payment pending is likely to be the most directly relevant to this question.

A reranking stage can help place the most relevant result higher among the retrieved candidates.


Why Is Reranking Useful in RAG?

RAG systems need to provide useful information to the LLM.

If irrelevant or less relevant information is included, it can make it harder for the model to focus on the information needed for the question.

Reranking can help by:

  • Improving the order of retrieved results
  • Bringing highly relevant information to the top
  • Reducing the number of less relevant results passed forward
  • Helping select better context for the LLM

However, reranking is not a guarantee of a correct answer. The quality of the documents, retrieval system, reranker, and overall RAG design still matters.


Does Reranking Replace Retrieval?

No.

Reranking usually works after an initial retrieval step.

Think of it like shopping for a product.

Retrieval

You search for:

“Laptop for programming”

The search system finds several possible laptops.

Reranking

The system then reorders those results based on how well they match the query.

So:

Retrieval finds the candidates.

Reranking improves their order.


An Easy Everyday Analogy

Imagine you ask a friend:

“Which restaurant nearby is good for a family dinner?”

Your friend gives you 10 options.

That's like initial retrieval.

Then your friend thinks:

  • Which restaurants are family-friendly?
  • Which are actually nearby?
  • Which have suitable food?
  • Which best match your requirements?

Your friend rearranges the list and gives you the most suitable options first.

That's similar to the idea of reranking.


Reranking Does Not Generate the Final Answer

This is an important point.

Reranking is mainly about organizing and selecting relevant search results.

It does not normally generate the final natural-language answer to the user.

The simplified flow is:

Search

↓

Reranking

↓

Relevant Context

↓

LLM

↓

Final Answer

The LLM is the component that generates the final response.


Reranking in a RAG Example

Let's put everything together.

User Question

“Can I write my final exam with 70% attendance?”

Initial Search Finds:

  1. Examination rules
  2. Hostel rules
  3. Attendance policy
  4. Leave policy
  5. Fee policy

Reranking

The system evaluates the relevance of these results to the question.

New Order:

  1. Attendance policy
  2. Examination rules
  3. Leave policy
  4. Fee policy
  5. Hostel rules

Context

The most relevant information is selected.

LLM

The LLM uses the retrieved context to generate the response.


Complete RAG Flow With Reranking

Here is the complete simplified flow:

Documents

↓

Chunking

↓

Embeddings

↓

Vector Database

↓

User Query

↓

Keyword / Semantic / Hybrid Search

↓

Initial Retrieval

↓

Reranking

↓

Top Relevant Results

↓

Context

↓

LLM

↓

Answer

Now you can see where Reranking fits into the bigger RAG picture.


One-Line Definition

Reranking is the process of reordering retrieved search results so that the most relevant information appears first.


Key Takeaway

Remember these three concepts:

Retrieval

Find potentially relevant information.

Reranking

Reorder the retrieved information based on relevance.

Generation

Use the selected context to generate the final answer.

In short:

Retrieve → Rerank → Provide Context → Generate


Conclusion

Reranking is an important concept in many RAG systems.

A search system may retrieve several potentially useful results, but not all of them are equally relevant.

Reranking helps put the most relevant results first.

It can work together with:

  • Keyword Search
  • Semantic Search
  • Similarity Search
  • Hybrid Search
  • Top-K Retrieval

The overall goal is simple:

Find the right information and make sure the most useful information gets priority.


What Will We Learn Next?

Now we know how RAG can:

Search → Retrieve → Rerank → Provide Context → Generate

But there is another important question:

What happens when an AI gives an answer that is not supported by the available information?

This leads us to our next topic:

How Does RAG Reduce Hallucination?

We will learn what AI hallucination means, why it happens, and how retrieval and context can help reduce unsupported answers.

Sunday, 27 September 2026

What Is Hybrid Search in RAG? A Simple Guide for Beginners

When we search for information, sometimes we know the exact words we are looking for.

Sometimes, we may ask the same thing using different words, but the meaning is still the same.

For example:

“75% attendance”

and

“How much attendance is required to write the exam?”

These two searches are different in wording, but they can refer to the same information.

This is where Hybrid Search becomes useful.


What Is Hybrid Search?

Hybrid Search is a search approach that combines keyword search and semantic search to find relevant information.

In simple words:

Hybrid Search = Keyword Search + Semantic Search

It tries to use the strengths of both types of search.


First, What Is Keyword Search?

Keyword search looks for specific words or terms in the available information.

For example, imagine a college handbook contains:

“Students must maintain a minimum of 75% attendance to appear for the final examination.”

If a student searches:

“75% attendance”

the system can find the document because the exact terms “75%” and “attendance” appear in the document.

This is the basic idea behind keyword search.

Simple Flow

User Query

↓

Find matching words

↓

Relevant Documents


What Is Semantic Search?

Semantic search focuses more on the meaning of the query rather than only matching exact words.

For example, the student may ask:

“Can I write my final exam if my attendance is low?”

The handbook may not contain exactly those words.

But the meaning is related to:

“Students must maintain a minimum of 75% attendance to appear for the final examination.”

Semantic search can use embeddings to understand that these two pieces of text are related.

Simple Flow

User Query

↓

Understand Meaning

↓

Find Similar Information

↓

Relevant Documents


So Why Do We Need Hybrid Search?

Keyword search and semantic search each have different strengths.

Keyword Search is useful when:

  • Exact words matter
  • Product names matter
  • Names or IDs are important
  • Technical terms need to match
  • Specific numbers are important

Semantic Search is useful when:

  • The user uses different words
  • The meaning is more important than exact wording
  • The query is written naturally
  • Similar concepts need to be found

Instead of relying on only one approach, we can combine both.

That is Hybrid Search.


Real-World Example: College Handbook

Let's continue with our college handbook example.

Imagine the handbook contains this rule:

“Students must maintain a minimum of 75% attendance to appear for the final examination.”

Now imagine three different student questions.

Question 1

“75% attendance”

This contains the exact terms from the document.

Keyword Search can be very useful here.


Question 2

“Can I attend my final exam with low attendance?”

The exact words may not appear in the handbook.

But the meaning is related to the attendance requirement.

Semantic Search can help here.


Question 3

“Is 75% attendance compulsory for final exam?”

This contains important exact terms like:

75% + attendance + final exam

and it also has a clear meaning related to the rule.

Using both keyword and semantic signals can help the system identify the relevant information.


How Hybrid Search Works

At a high level, Hybrid Search can work like this:

The exact implementation can vary between systems.


Hybrid Search in RAG

Now let's connect Hybrid Search to everything we have already learned about RAG.

Our RAG pipeline was:

With Hybrid Search, the search stage can use both keyword and semantic search.

So it becomes:

Documents

↓

Chunking

↓

Embeddings + Searchable Text

↓

Search System

↓

Keyword Search + Semantic Search

↓

Combine Results

↓

Retrieval

↓

Context

↓

LLM

↓

Answer


Why Is Keyword Search Still Important?

You may wonder:

“If semantic search understands meaning, why do we still need keyword search?”

Because exact words can sometimes be very important.

Imagine a customer asks:

“What is the status of order AB12345?”

The order number AB12345 is very specific.

A keyword-based search can directly look for that exact identifier.

Semantic similarity alone may not be the best way to handle such exact identifiers.

The same idea applies to:

  • Product IDs
  • Order numbers
  • Employee IDs
  • Model numbers
  • Error codes
  • Names
  • Technical terms

So keyword search still has an important role.


Another Real-World Example: Online Shopping

Imagine you are searching for a product.

You type:

“Nike running shoes size 9”

There are different types of information in this query.

“Nike” → Brand

“running shoes” → Product type

“size 9” → Exact attribute

A search system can use keyword matching for specific terms and semantic understanding for the overall meaning.

This combination can help retrieve relevant products.


Another Example: Customer Support

Imagine a company's support documents contain:

“Customers can request a refund within 30 days of purchase.”

A customer asks:

“I bought this product 20 days ago. Can I get my money back?”

The exact phrase “request a refund” may not appear in the customer's question.

Semantic search can recognize the relationship.

But if the customer asks:

“What is the 30-day refund policy?”

keyword matching can also be useful because “30-day” and “refund” are important terms.

Hybrid Search can use both types of signals.

Does Hybrid Search Always Give Better Results?

Not automatically.

The effectiveness depends on:

  • The quality of the documents
  • How the queries are written
  • Search configuration
  • How keyword and semantic results are combined
  • Ranking methods
  • The quality of the retrieval system

So Hybrid Search is not simply:

“Use two searches and everything will be perfect.”

The goal is to combine useful signals to improve information retrieval for a particular application.


What Happens After Hybrid Search?

Hybrid Search helps find relevant information.

But finding information is not the final step in a RAG system.

The retrieved information can become context.

Then the LLM can use that context to generate an answer.

Example

Student Question:

“Can I write my final exam with 70% attendance?”

↓

Hybrid Search

↓

Find relevant attendance rule

↓

Retrieved Context:

“Students must maintain a minimum of 75% attendance…”

↓

LLM

↓

Answer:

The student does not meet the stated 75% attendance requirement.

The exact answer should, of course, follow the actual college policy contained in the source documents.


Simple Everyday Analogy

Imagine you are searching for a particular book in a large library.

You remember the exact title.

You can search using the title.

That's similar to keyword search.

But what if you don't remember the title?

You remember:

“It was a book about learning AI for beginners.”

Now you are searching based on the meaning or topic.

That's similar to semantic search.

If the library uses both:

Exact title + Topic/Meaning

that is similar to the basic idea of Hybrid Search.


Where Can Hybrid Search Be Used?

Hybrid Search can be useful in many applications, such as:

🛒 E-commerce

Finding products using product names, attributes, and meaning.

🏢 Company Knowledge Bases

Finding relevant HR, finance, or company policy documents.

🎓 Education

Finding information from textbooks, student handbooks, and course materials.

💬 Customer Support

Finding the right help article for a customer's question.

📄 Document Search

Searching large collections of PDFs, reports, manuals, and other documents.

🤖 RAG Applications

Retrieving relevant information before giving it to an LLM.


Hybrid Search and Embeddings

We have already learned about embeddings in our previous posts.

Embeddings convert text into numerical representations that capture aspects of its meaning.

Semantic search can use these embeddings to find semantically related information.

Hybrid Search can combine this semantic search with traditional keyword-based search.

So our concepts are connected:

Text

↓

Chunking

↓

Embeddings

↓

Semantic Search

Keyword Search

↓

Hybrid Search

↓

Relevant Results


Hybrid Search in One Simple Example

Let's take everything in one example.

Document

“Students must maintain a minimum of 75% attendance to appear for the final examination.”

Query A

“75% attendance”

Keyword Search → Strong exact match

Query B

“Can I write my exam if I have low attendance?”

Semantic Search → Strong meaning-based match

Query C

“Is 75% attendance required for final exam?”

Keyword + Semantic Search → Both signals can be useful

This is the basic idea behind Hybrid Search.


The Complete RAG Journey So Far

We have now learned many important RAG concepts:

1. Chunking
Break large documents into smaller meaningful pieces.

↓

2. Embeddings
Convert text into numerical representations.

↓

3. Vector Database
Store and search those representations.

↓

4. Semantic Search
Find information based on meaning.

↓

5. Similarity Search
Find representations that are similar.

↓

6. Retrieval
Collect relevant information.

↓

7. Context
Give relevant information to the LLM.

↓

8. RAG vs Fine-Tuning
Understand two different ways of adapting AI systems.

↓

9. Hybrid Search
Combine keyword and semantic search approaches.


One-Line Definition

Hybrid Search is a search approach that combines keyword search and semantic search to find relevant information using both exact terms and meaning.


Conclusion

Searching for information is an important part of a RAG system.

Keyword Search is useful when exact words, numbers, names, or identifiers matter.

Semantic Search is useful when the meaning of the query matters more than exact wording.

Hybrid Search combines both approaches.

That can be especially useful when users may search using a mixture of:

Exact terms + Natural language + Meaning

And this is one reason Hybrid Search is an important concept when building practical RAG applications.


What Will We Learn Next?

We have learned how a RAG system can find relevant information.

But what if the search returns many results?

How does the system decide:

“Which result should come first?”

That's where our next concept comes in:

What Is Reranking in RAG?

We will learn how retrieved results can be reordered to identify the most relevant information before it is given to the LLM.


Tuesday, 22 September 2026

RAG vs Fine-Tuning: What Is the Difference? A Simple Guide for Beginners Artificial Intelligence models like Large

Artificial Intelligence models like Large Language Models (LLMs) can answer questions, write content, summarize information, and perform many other tasks.

But sometimes we want an AI model to work with our own information or behave in a specific way.

Two common approaches are:

  • RAG (Retrieval-Augmented Generation)
  • Fine-Tuning

Both can customize how we use AI, but they work in very different ways.

Let’s understand the difference with simple real-world examples.


What Is RAG?

RAG stands for Retrieval-Augmented Generation.

In RAG, we give the AI access to relevant external information when the user asks a question.

The AI does not need to learn that information permanently.

Instead, it:

  1. Receives the user's question
  2. Searches the available documents
  3. Retrieves relevant information
  4. Provides that information as context to the LLM
  5. Generates an answer using that context

Simple RAG Flow

Documents → Chunking → Embeddings → Vector Database → User Question → Search → Retrieval → Context → LLM → Answer


Real-World Example: College Handbook

Imagine a college has a student handbook containing information about:

  • Attendance rules
  • Exam rules
  • Leave policies
  • Hostel rules
  • Fee details

A student asks:

“How much attendance do I need to write my final exam?”

Instead of expecting the AI model to already know the college's rules, a RAG system searches the college handbook.

It may retrieve information such as:

“Students must maintain a minimum of 75% attendance to appear for the final examination.”

This retrieved information becomes context.

The LLM then uses that context to generate the answer.

So:

RAG = Find the relevant information → Give it to the AI → Generate an answer


What Is Fine-Tuning?

Fine-tuning means further training an existing AI model using a specific set of examples.

Instead of simply giving the model information during a question, we train the model to improve its behavior for a particular task or style.

For example, imagine a company wants an AI assistant to consistently respond to customer questions in a particular format.

The company can provide many examples of:

Customer Question → Desired Answer

The model learns patterns from these examples during fine-tuning.


Simple Fine-Tuning Example

Imagine a company wants its AI assistant to respond like this:

Customer:
“My product has not arrived yet.”

Desired style:
“Sorry for the delay. Please share your order number so we can check the delivery status.”

If the company provides many examples like this during fine-tuning, the model can learn the desired response patterns and style.

So:

Fine-Tuning = Train the model with examples → Adapt its behavior for a specific task


RAG vs Fine-Tuning: The Main Difference

The easiest way to understand the difference is:

RAG

Changes the information available to the AI at answer time.

Fine-Tuning

Changes how the model behaves by training it with examples.

Think of it like this:

RAG = Give the AI a book to refer to.

Fine-Tuning = Teach the AI how you want it to perform a task.

Simple Example

Let's take our college example again.

Suppose a college changes its attendance rule.

Previously:

Minimum attendance = 75%

Later, the college changes it to:

Minimum attendance = 80%

With RAG, we can update the college handbook or knowledge base.

The next time a student asks about attendance, the system can retrieve the updated information.

We don't necessarily need to retrain the entire model just because the document changed.


What About Fine-Tuning?

Fine-tuning is useful when we want the model to learn a particular behavior, style, format, or task pattern from examples.

For example, a company may want its AI assistant to:

  • Follow a specific response format
  • Classify customer messages
  • Follow a particular writing style
  • Perform a specialized task consistently

Fine-tuning can help the model learn those patterns from training examples.


Another Real-World Example: Company HR Assistant

Imagine a company has an internal HR assistant.

Employees ask questions such as:

“How many days of leave can I take?”

“What is the work-from-home policy?”

“What is the maternity leave policy?”

These rules may change over time.

A RAG system can connect the AI to the company's latest HR documents.

Flow:

HR Documents → Chunking → Embeddings → Vector Database → Employee Question → Retrieval → Context → LLM → Answer

The AI retrieves the relevant policy before answering.


When Can RAG Be Useful?

RAG can be useful when the AI needs access to external or changing information.

For example:

  • Company documents
  • College handbooks
  • Product manuals
  • Customer-support knowledge bases
  • Internal policies
  • Research documents
  • Frequently updated information

The important idea is:

The information can be stored outside the model and retrieved when needed.


When Can Fine-Tuning Be Useful?

Fine-tuning can be useful when we want an AI model to become more consistent at a particular task or response pattern.

For example:

  • Specific classification tasks
  • Consistent output formats
  • Specialized writing styles
  • Task-specific behavior
  • Repeated domain-specific patterns

The important idea is:

The model learns from examples during additional training.


Can We Use RAG and Fine-Tuning Together?

Yes.

They are not necessarily competing technologies.

A system can use both.

For example, imagine a customer-support AI.

Fine-Tuning can help with:

How should the AI respond?

RAG can help with:

What current information should the AI use?

So the system could work like:

Fine-Tuned Model + RAG Knowledge Base → Context → Answer

This can be useful when an application needs both specialized behavior and access to external information.


An Easy Everyday Analogy

Imagine you are preparing for an exam.

You already know how to answer questions because you have practiced many examples.

That is somewhat like Fine-Tuning.

Now imagine you are allowed to take your textbook into the exam and look up a specific fact when needed.

That is somewhat like RAG.

So:

Fine-Tuning → Learn a particular way of doing something

RAG → Look up relevant information when needed

This is only an analogy, but it makes the basic difference easier to remember.


One Important Point

RAG does not mean that the AI permanently learns the documents.

Suppose we connect an AI system to a company's employee handbook.

The handbook can be retrieved as context when an employee asks a question.

The model is not automatically retrained every time the handbook changes.

This is one reason RAG can be useful for information that changes frequently.


Does Fine-Tuning Store Your Whole Knowledge Base?

Not in the same way as a database.

Fine-tuning trains a model using examples so that its parameters are adjusted to learn patterns.

It is therefore important not to think of fine-tuning as simply:

“Uploading a document into the AI's memory.”

For large collections of changing documents, a retrieval-based approach such as RAG may be more suitable depending on the application.


RAG and Fine-Tuning in One Simple Diagram

RAG

Your Documents
↓
Retrieve Relevant Information
↓
Give Information as Context
↓
LLM
↓
Answer


Fine-Tuning

Training Examples
↓
Additional Training
↓
Adapted Model
↓
User Question
↓
Answer


The Simplest Way to Remember

If your main question is:

“How can I give the AI the latest information?”

Think about RAG.

If your main question is:

“How can I make the model behave differently for a particular task?”

Think about Fine-Tuning.

The exact choice depends on the application, the data, how frequently the information changes, and the behavior you want from the model.


RAG + Fine-Tuning + LLM

Now our RAG learning journey looks like this:

Documents
↓
Chunking
↓
Embeddings
↓
Vector Database
↓
Semantic Search
↓
Similarity Search
↓
Retrieval
↓
Context
↓
LLM
↓
Answer

And now we also understand that:

RAG helps an LLM use relevant external information.

Fine-Tuning helps adapt a model using additional training examples.


Conclusion

RAG and Fine-Tuning are two different ways of adapting AI systems for specific needs.

RAG focuses on bringing relevant information to the model when it is needed.

Fine-Tuning focuses on adapting the model's behavior through additional training.

A simple way to remember:

RAG = Give the AI the right information.
Fine-Tuning = Teach the AI a particular way to perform a task.

And in real AI applications, both approaches can sometimes be used together.


What Will We Learn Next?

Now that we understand RAG vs Fine-Tuning, the next interesting concept is:

What Is Hybrid Search in RAG?

We will learn how AI can combine keyword search + semantic search to find more relevant information.



Friday, 18 September 2026

What Is Context in RAG? A Simple Guide for Beginners

In our previous blogs, we learned how RAG works step by step.

We learned about:

  • Chunking
  • Embeddings
  • Vector Database
  • Semantic Search
  • Similarity Search
  • Retrieval

We now know that Retrieval helps find the relevant information from a large collection of documents.

But what happens after that?

The relevant information is given to the LLM so that it can use that information to generate an answer.

This relevant information is called Context.

Let's understand what context means in RAG with a simple real-world example.


What Is Context?

Context is the relevant information provided to an AI model to help it understand and answer a user's question.

In RAG, context usually comes from information retrieved from external documents or a knowledge base.

Simple definition:

Context is the useful information given to the LLM along with the user's question so it can generate a better answer.


Real-World Example: College Student Handbook

Imagine a college has a large student handbook.

It contains information about:

  • Attendance
  • Exams
  • Hostel
  • Library
  • Fees
  • Leave Rules

Now a student asks:

"Can I write my final exam if my attendance is 70%?"

The AI system needs to find the correct information.


Step 1: User Asks a Question

The student asks:

"Can I write my final exam if my attendance is 70%?"

This is the user query.

User Question
      ↓
"Can I write my final exam
if my attendance is 70%?"

Step 2: The System Searches for Relevant Information

The question is converted into an embedding.

The system then uses similarity-based search to find relevant information from the vector database.

It may find:

"Students must maintain a minimum of 75% attendance to appear for the final examination."

This information is relevant to the student's question.


Step 3: Retrieved Information Becomes Context

The relevant information retrieved from the college handbook is provided to the LLM.

This information acts as context.

User Question
      ↓
Retrieved Information
      ↓
Context
      ↓
LLM

So the LLM receives something like:

Question:

Can I write my final exam if my attendance is 70%?

Context:

Students must maintain a minimum of 75% attendance to appear for the final examination.

Now the LLM has the information it needs to formulate an answer.


Why Does an LLM Need Context?

An LLM has learned from large amounts of data during training.

But it may not have access to your specific or current private information.

For example, an LLM may not automatically know:

  • Your college's latest attendance policy
  • Your company's internal HR rules
  • A product's latest internal manual
  • A private organization's documents

RAG helps by retrieving the relevant information and providing it as context.


RAG Without Context

Imagine asking:

"What is my college's attendance requirement?"

Without access to the college handbook, the AI may not know the specific rule.

It could potentially give a generic answer that does not match your college's actual policy.


RAG With Context

Now imagine the system retrieves this information:

"Students must maintain a minimum of 75% attendance to appear for the final examination."

This information is provided to the LLM as context.

The LLM can then use that information to formulate the answer.

Question
   +
Relevant Context
   ↓
LLM
   ↓
Answer

This is the basic idea behind Retrieval-Augmented Generation.


What Is the Difference Between Context and Query?

These two terms can sometimes be confusing.

Query

The query is what the user asks.

Example:

"Can I write my exam with 70% attendance?"

Context

The context is the relevant information provided to help answer the query.

Example:

"Students need a minimum of 75% attendance to appear for the final examination."

So:

Query   → What the user wants to know

Context → Information that helps answer it

How Is Context Created in RAG?

Context in RAG generally comes from the retrieved information.

The basic process is:

Documents
    ↓
Chunking
    ↓
Embeddings
    ↓
Vector Database
    ↓
User Query
    ↓
Similarity Search
    ↓
Retrieval
    ↓
Relevant Chunks
    ↓
Context
    ↓
LLM
    ↓
Answer

The retrieved chunks become the information the LLM can use as context.


Can There Be More Than One Context?

Yes.

Sometimes one chunk may not contain everything needed to answer a question.

The system may retrieve multiple relevant chunks.

For example, a student asks:

"What attendance do I need, and what happens if my attendance is below the requirement?"

The system might retrieve:

Chunk 1:

Students need a minimum of 75% attendance.

Chunk 2:

Students below the required attendance may need to follow the college's eligibility or condonation policy.

Both pieces of information can be provided to the LLM as context.

Relevant Chunk 1
       +
Relevant Chunk 2
       ↓
    Context
       ↓
      LLM

Too Little Context

Context needs to be relevant and sufficient.

Suppose the user asks:

"What happens if my attendance is below 75%?"

But the system retrieves only:

"Students must maintain 75% attendance."

This may not provide enough information to answer the complete question.

The system may need another relevant chunk containing the consequences or applicable policy.


Too Much Unnecessary Context

On the other hand, imagine giving the LLM hundreds of unrelated pages.

For example:

Attendance Rule
Hostel Rule
Library Rule
Fee Rule
Bus Rule
Canteen Rule
Sports Rule
...
500 pages

Most of this information is unrelated to the question.

A good RAG system tries to retrieve relevant context instead of simply giving everything to the LLM.

Large Document Collection
          ↓
Relevant Retrieval
          ↓
Useful Context
          ↓
LLM

Context in a Real-World Company

Let's take another example.

Imagine an employee asks:

"How many days of leave can I take?"

The company has hundreds of documents.

The RAG system searches the company's HR knowledge base and retrieves the relevant leave policy.

For example:

"Employees are eligible for 18 days of annual leave per year."

That information becomes context for the LLM.

Employee Question
       ↓
Retrieve HR Policy
       ↓
Relevant Context
       ↓
LLM
       ↓
Answer

The same idea can be used for:

  • Customer support
  • Product manuals
  • Company policies
  • College documents
  • Research documents
  • Internal knowledge bases

Context Does Not Mean the LLM Is Retrained

This is an important point.

When we provide retrieved information as context, we are not training the LLM again.

The information is simply provided to the model while it is answering the current question.

Think of it like this:

Training

The model learns patterns from training data.

RAG Context

The model receives relevant information for a particular question.

Training
   ↓
Model learns

RAG Context
   ↓
Model gets useful information for the current question

So RAG can provide external information without retraining the entire model for every new document.


Context and LLM: Simple Example

Let's imagine the LLM is like a student taking an exam.

The student knows many things already.

But the teacher gives the student a relevant page from a textbook.

That page acts like context.

The student can use that information to answer the question.

Similarly:

User Question
      +
Relevant Context
      ↓
     LLM
      ↓
Generated Answer

Where Does Context Fit in the RAG Pipeline?

Let's connect everything we have learned so far.

Documents
    ↓
Chunking
    ↓
Embeddings
    ↓
Vector Database
    ↓
User Question
    ↓
Query Embedding
    ↓
Similarity Search
    ↓
Retrieval
    ↓
Relevant Context
    ↓
LLM
    ↓
Generated Answer

Now the complete flow becomes much easier to understand.


Why Is Good Context Important?

Good context should be:

Relevant

It should relate to the user's question.

Useful

It should contain information that helps answer the question.

Focused

It should avoid unnecessary information whenever possible.

For example, if the question is about attendance, the system should prioritize attendance-related information instead of unrelated hostel or library rules.


Simple Everyday Example

Imagine you ask your friend:

"When is my doctor's appointment?"

Your friend checks your calendar and tells you:

"Your appointment is tomorrow at 10 AM."

The calendar information is the context your friend used to answer your question.

Similarly, in RAG:

Question
   ↓
Find Relevant Information
   ↓
Give Information to LLM
   ↓
Generate Answer

The relevant information helps the AI answer the question.


Context vs Knowledge

Context and knowledge are not exactly the same thing.

Knowledge

Information the model has learned during training.

Context

Information provided to the model while answering a particular question.

For example:

Model's Learned Knowledge
          +
Retrieved Context
          ↓
        LLM
          ↓
       Answer

RAG mainly helps bring external or specific information into the current context.


One Complete RAG Example

Let's put everything together.

User Question

"Can I write my final exam if my attendance is 70%?"

Search

The system searches the college knowledge base.

Retrieval

It retrieves:

"Students must maintain a minimum of 75% attendance to appear for the final examination."

Context

The retrieved information becomes context.

Generation

The LLM uses the question and context to generate an answer.

User Question
      ↓
Query Embedding
      ↓
Similarity Search
      ↓
Retrieval
      ↓
Relevant Information
      ↓
Context
      ↓
LLM
      ↓
Answer

Simple Recap

Let's quickly remember what we learned.

What is Context?

Relevant information provided to the LLM to help answer a question.

Where does RAG context come from?

Usually from information retrieved from external documents or a knowledge base.

Is the user question the context?

No.

The user question is the query.

The relevant supporting information is the context.

Does providing context retrain the LLM?

No.

The information is provided to the model while generating the current response.

Why is good context important?

Relevant and focused context helps the LLM use the right information when generating an answer.


Conclusion

Context is a very important part of RAG.

We can think of it as the useful information that is brought to the LLM at the right time.

The overall idea is simple:

User asks → Relevant information is retrieved → Retrieved information becomes context → LLM uses the context → Answer is generated.

So the RAG process we have learned so far looks like this:

Documents
   ↓
Chunking
   ↓
Embeddings
   ↓
Vector Database
   ↓
User Question
   ↓
Similarity Search
   ↓
Retrieval
   ↓
Context
   ↓
LLM
   ↓
Answer

Now we have understood how RAG finds information and gives that information to the LLM.

In the next blog, we can explore an important question:

RAG vs Fine-Tuning: What Is the Difference?

Monday, 14 September 2026

What Is Retrieval in RAG? A Simple Guide for Beginners

In our previous blogs, we learned about:

  • Chunking
  • Embeddings
  • Vector Databases
  • Semantic Search
  • Similarity Search

We learned how AI finds information that is relevant to a user's question.

But one important question remains:

After finding the relevant information, how does the system actually collect it and use it to generate an answer?

This process is called Retrieval.

Let's understand Retrieval in RAG with a simple real-world example.


What Is Retrieval?

Retrieval is the process of finding and collecting relevant information from stored data based on a user's query.

In a RAG system, retrieval usually means selecting the most useful document chunks from a vector database and sending them as context to the LLM.

Simple definition:

Retrieval is the process of bringing the most relevant information from a knowledge source for answering a question.


Real-World Example: College Student Handbook

Imagine a college has a large student handbook.

It contains information about:

  • Attendance
  • Exams
  • Hostel
  • Library
  • Fees
  • Leave Rules

The handbook is divided into smaller chunks and stored in a vector database.

Now a student asks:

"How much attendance do I need to write my final exam?"

The system needs to find the correct information from the handbook.


How Retrieval Works Step by Step

Step 1: The User Asks a Question

The student asks:

"How much attendance do I need to write my final exam?"

This is called the user query.


Step 2: The Query Is Converted Into an Embedding

The question is converted into a numerical representation called an embedding.

User Question
      ↓
Query Embedding

This helps the system compare the question with stored document embeddings.


Step 3: Similarity Search Finds Relevant Chunks

The system compares the query embedding with the stored embeddings in the vector database.

For example:

Attendance Rule → Highly Relevant
Exam Rule       → Relevant
Hostel Rule     → Not Relevant
Library Rule    → Not Relevant
Fee Rule        → Not Relevant

Similarity Search helps identify which chunks are closest to the user's question.


Step 4: Relevant Chunks Are Retrieved

Now the system selects the useful information.

For example, it retrieves:

"Students must maintain a minimum of 75% attendance to appear for the final examination."

This act of selecting and collecting the relevant chunk is called Retrieval.


Similarity Search vs Retrieval

These two concepts are closely connected, but they have different roles.

Similarity Search

Similarity Search helps answer:

"Which stored information is most similar to the user's question?"

Retrieval

Retrieval helps answer:

"Which relevant information should be collected and passed to the next stage?"

In many RAG systems, similarity search is one of the techniques used during the retrieval process.


What Is Retrieved Information?

The information retrieved from the knowledge source is usually called a:

  • Relevant chunk
  • Retrieved document
  • Retrieved context
  • Search result

For example:

User Question:
"How much attendance is required?"

Retrieved Chunk:
"Students must maintain 75% attendance..."

This retrieved chunk contains the information needed to answer the question.


What Is Top-K Retrieval?

Sometimes the system does not retrieve only one result.

It retrieves the top few most relevant results.

This is called Top-K Retrieval.

Here, K means the number of results to retrieve.

For example:

Top-1 → Retrieve 1 result
Top-3 → Retrieve 3 results
Top-5 → Retrieve 5 results

For the college question, Top-3 Retrieval might return:

  1. Attendance Rule
  2. Exam Rule
  3. Examination Eligibility Rule

The system can then use these results to prepare a better answer.


Why Not Retrieve Every Document?

Imagine the college handbook contains 500 pages.

If the system sends all 500 pages to the LLM, it would be:

  • Slow
  • Expensive
  • Unnecessary
  • Difficult for the LLM to process
  • More likely to include unrelated information

Instead, Retrieval selects only the relevant sections.

500 Pages
    ↓
Retrieve Relevant Chunks
    ↓
Only 2–5 Useful Chunks
    ↓
LLM

This makes the system more efficient.


Retrieval in the RAG Pipeline

Let's connect Retrieval with the complete RAG process.

Retrieval is the bridge between the stored knowledge and the LLM.


What Happens After Retrieval?

After the relevant chunks are retrieved, they are added to the prompt as context.

For example:

User Question:
"How much attendance do I need?"

Retrieved Context:
"Students must maintain a minimum of 75% attendance..."

LLM:
Uses the question + retrieved context
      ↓
Generates the answer

The LLM can now answer using the actual information from the college handbook.


Another Real-World Example: Company HR Policy

Imagine an employee asks:

"How many days of maternity leave are available?"

A company RAG system may search its HR policy documents.

It retrieves the relevant leave policy section instead of sending every company document to the LLM.

Employee Question
      ↓
Search HR Documents
      ↓
Retrieve Leave Policy
      ↓
Send Relevant Context to LLM
      ↓
Generate Answer

This is Retrieval in a practical business application.


Why Is Retrieval Important in RAG?

Retrieval is important because an LLM may not know the latest or private information stored in a company's documents.

For example:

  • College rules
  • Company HR policies
  • Product manuals
  • Internal documents
  • Customer support information
  • Legal or financial documents

Retrieval brings the required information from these sources before the LLM generates the answer.


Retrieval Does Not Generate the Final Answer

This is an important point.

Retrieval only finds and collects the relevant information.

It does not explain the answer in natural language.

Retrieval
    ↓
Finds relevant information

LLM
    ↓
Understands the context
and generates the answer

So, Retrieval and Generation are two different stages.


Retrieval vs Generation

Retrieval

Finds information from stored knowledge.

Generation

Creates a natural-language answer using the retrieved information.

Retrieval → Finds information

Generation → Creates the answer

That is why RAG stands for:

Retrieval-Augmented Generation

  • Retrieval → Find relevant information
  • Augmented → Add that information to the prompt
  • Generation → LLM generates the answer

Complete Example

Let's see the complete process using the college handbook.

User Question

"Can I write my final exam with 70% attendance?"

System Process

1. User asks the question
          ↓
2. Question becomes an embedding
          ↓
3. Similarity Search checks stored chunks
          ↓
4. Attendance rule is identified
          ↓
5. Relevant chunk is retrieved
          ↓
6. Retrieved chunk is added as context
          ↓
7. LLM reads the context
          ↓
8. LLM generates the answer

Final Answer

"According to the college policy, students need a minimum of 75% attendance to appear for the final examination."

The answer is based on the retrieved college information.


Simple Retrieval Flow

User Question
      ↓
Query Embedding
      ↓
Similarity Search
      ↓
Relevant Chunks Selected
      ↓
Retrieved Context
      ↓
LLM
      ↓
Answer

Key Points to Remember

  • Retrieval means finding and collecting relevant information.
  • Similarity Search helps identify the closest matching chunks.
  • Retrieval is an important stage in RAG.
  • Only useful information is passed to the LLM.
  • Retrieval does not generate the final answer.
  • The LLM uses the retrieved context to create the response.
  • Top-K Retrieval means selecting the top few relevant results.

Conclusion

Retrieval is one of the most important steps in a RAG system.

It helps the AI system find the right information from a large collection of documents and provide it to the LLM.

Instead of asking the LLM to answer only from its memory, RAG first retrieves relevant information and then generates the answer using that context.

The simple idea to remember is:

Retrieval = Find and collect the right information before generating the answer.

The complete flow is:

Question
   ↓
Search
   ↓
Retrieve Relevant Information
   ↓
Add Context
   ↓
LLM
   ↓
Answer

In our next blog, we can explore:

What Is Context in RAG? How Retrieved Information Helps the LLM

Wednesday, 9 September 2026

What Is Similarity Search? How AI Finds the Closest Information

In our previous blog, we learned about Semantic Search.

Semantic Search helps AI find information based on meaning, instead of looking only for exact keywords.

But now an important question comes:

How does AI decide which information is most similar to our question?

The answer is Similarity Search.

Similarity Search is one of the important techniques used behind semantic search and many RAG systems.

In this blog, let's understand it with a simple real-world example.


What Is Similarity Search?

Similarity Search is the process of finding the most similar information from a collection of stored information.

In AI applications, similarity is often calculated by comparing embeddings.

Simple definition:

Similarity Search finds the information that is closest to a given query based on its representation.


Let's Take a Real-World Example

Imagine a college has a student handbook.

The handbook contains different information:

  • Attendance Rules
  • Exam Rules
  • Hostel Rules
  • Library Rules
  • Fee Rules

These sections are converted into smaller chunks and then into embeddings.

Now a student asks:

"How much attendance do I need for my final exam?"

The system needs to find which information is closest to this question.

It may compare the question with different stored chunks:

User Question
      ↓
"How much attendance do I need
for my final exam?"
      ↓
Compare with stored information
      ↓
Attendance Rule → Very Relevant
Exam Rule       → Relevant
Hostel Rule     → Not Relevant
Library Rule    → Not Relevant
Fee Rule        → Not Relevant

The attendance rule is the best match.

This process of finding the clo


sest matches is called Similarity Search.


But How Does AI Know What Is Similar?

This is where Embeddings come in.

We already learned that embeddings convert text into numerical representations.

For example:

"How much attendance do I need?"
              ↓
        Embedding
              ↓
     [0.21, 0.67, 0.45, ...]

The attendance rule also has an embedding:

"Students must maintain 75% attendance..."
              ↓
        Embedding
              ↓
     [0.22, 0.65, 0.48, ...]

The system compares these representations.

If they are close according to the chosen similarity method, the information is considered more similar.


What Is a Similarity Score?

When two embeddings are compared, the system can produce a similarity score.

The score helps the system understand how closely two items match according to the selected similarity method.

For example:

Question ↔ Attendance Rule
        ↓
     High Similarity

Question ↔ Hostel Rule
        ↓
     Low Similarity

So the system can select the most relevant information.

Remember:

Higher similarity generally means the information is more closely related.

The exact score and its range depend on the similarity method being used.


Similarity Does Not Mean Exact Words

This is very important.

Consider these two sentences:

Sentence 1:

"How much attendance is required for the final exam?"

Sentence 2:

"Students need a minimum of 75% attendance to appear for the examination."

The words are different.

But the meaning is closely related.

Embeddings help represent that meaning, and similarity search can help identify that relationship.

That's why AI systems can find relevant information even when the user's words don't exactly match the document.


One Common Method: Cosine Similarity

One commonly used method for comparing embeddings is Cosine Similarity.

You don't need to learn the mathematics right now.

The basic idea is:

It compares the direction of two vectors.

Think about two arrows.

If they point in a similar direction, they are more similar.

Vector A  ↗
          /

         ↗ Vector B

If they point in very different directions, they are less similar.

Vector A  ↗

          ↘ Vector B

So cosine similarity is one way an AI system can measure how similar two embeddings are.


Where Does Similarity Search Fit in RAG?

Now let's connect this with the RAG concepts we have already learned.

A simplified RAG flow looks like this:

Similarity Search helps answer this question:

"Which stored chunks are closest to the user's question?"


What Does the Vector Database Do?

A Vector Database stores embeddings and allows the system to search through them efficiently.

For example:

Vector Database

Attendance Chunk → Embedding
Exam Chunk       → Embedding
Hostel Chunk     → Embedding
Library Chunk    → Embedding
Fee Chunk        → Embedding

When the student asks a question:

User Question
      ↓
Query Embedding
      ↓
Similarity Search
      ↓
Vector Database
      ↓
Most Relevant Chunks

The relevant chunks can then be passed to the LLM.


Semantic Search vs Similarity Search

This is where many beginners get confused.

They are closely connected, but we can explain them differently.

Semantic Search

Semantic Search is the search approach.

Its goal is:

Find information based on meaning.

Similarity Search

Similarity Search is the process of comparing representations to find the closest matches.

Its goal is:

Find which stored items are most similar to the query.

So, in many AI systems:

Semantic Search
      ↓
Search based on meaning
      ↓
Embeddings
      ↓
Similarity Search
      ↓
Relevant Information

Important: These terms can overlap in real-world AI discussions. We are separating them here only to make the concepts easier for beginners to understand.


A Simple Everyday Example

Imagine you are searching for a product online.

You type:

"Comfortable shoes for daily walking"

The search system may find products described as:

"Lightweight sneakers for everyday walking."

The words are not exactly the same.

But the meaning is closely related.

The system can use embeddings and similarity-based techniques to find relevant results.

This is the basic idea behind many modern AI-powered search systems.


Similarity Search vs Keyword Search

Let's compare them.

Keyword Search

User:

"attendance"

The system mainly looks for matching words such as:

attendance

Similarity-Based Search

User:

"How much attendance do I need for my exam?"

The system can find information such as:

"Students must maintain a minimum of 75% attendance to appear for the final examination."

Even though the wording is different, the meaning is closely related.


Why Is Similarity Search Important?

Similarity Search is useful because people don't always ask questions using the exact words present in a document.

For example:

Document:

"Minimum attendance requirement is 75%."

User:

"Can I write my final exam if my attendance is 70%?"

The wording is different.

But the question is related to the attendance requirement.

Similarity-based retrieval can help find the relevant attendance information.


Similarity Search in One Simple Flow

Let's remember the entire idea with our college example.

Student asks a question
          ↓
Question is converted into embedding
          ↓
Compared with stored embeddings
          ↓
Similarity is calculated
          ↓
Most relevant information is selected
          ↓
Relevant chunk is retrieved
          ↓
LLM uses the information
          ↓
Final answer

One Important Point

Similarity Search does not generate the final answer.

Its main job is to find relevant information.

The LLM then uses that information to generate a human-readable answer.

So remember:

Similarity Search
       ↓
Finds relevant information

LLM
       ↓
Generates the answer

How Similarity Search Connects With What We Learned

So far, our concepts are building one after another:

Embeddings
    ↓
Vector Database
    ↓
Semantic Search
    ↓
Similarity Search
    ↓
Relevant Information
    ↓
Retrieval
    ↓
Context
    ↓
LLM
    ↓
Answer

Each concept has a different role.

That's why understanding similarity search helps us understand how RAG retrieves the right information.


Simple Recap

What is Similarity Search?

It finds the most similar or relevant information from stored data.

What does it compare?

Usually, embeddings represented as vectors.

What is a similarity score?

A value used to indicate how closely two representations match according to a chosen similarity method.

What is Cosine Similarity?

One common method for comparing the similarity of vectors.

Does Similarity Search generate an answer?

No.

It finds relevant information. The LLM generates the final answer.


Conclusion

Similarity Search is an important part of many modern AI systems.

It helps the system find information that is closely related to the user's question.

In RAG, similarity search helps identify relevant chunks from a vector database. Those chunks can then be provided to the LLM as context.

The simple idea to remember is:

Similarity Search = Find the closest relevant information.

And the overall process is:

Question
   ↓
Embedding
   ↓
Similarity Search
   ↓
Relevant Information
   ↓
LLM
   ↓
Answer

Now we understand how the system finds relevant information.

But another important question comes next:

After finding the relevant information, how does RAG actually retrieve it and give it to the LLM?

That takes us to our next topic:

What Is Retrieval in RAG?

Sunday, 6 September 2026

What Is Semantic Search? A Simple Guide for Beginners

Have you ever searched for something online using different words, but still got the result you were looking for?



For example, you search:

"How much attendance do I need for exams?"

But the document says:

"Students must maintain a minimum of 75% attendance to be eligible for the examination."

The words are different, but the meaning is similar.

How can an AI system understand this?

The answer is Semantic Search.


🔍 What Is Semantic Search?

Semantic Search is a search method that looks at the meaning of a query instead of matching only the exact words.

In simple words:

Keyword Search looks for matching words.

Semantic Search looks for matching meaning.

This makes semantic search especially useful for AI applications such as RAG systems, chatbots, recommendation systems, and question-answering systems.


📝 Keyword Search vs Semantic Search

Let's understand with a simple example.

Suppose a college document contains:

"Students must maintain a minimum of 75% attendance to appear for the final examination."

Now a student searches:

"What attendance is required for the exam?"

The words are not exactly the same.

The document says:

  • maintain
  • minimum
  • attendance
  • appear
  • examination

The student asks:

  • attendance
  • required
  • exam

A traditional keyword search may focus on exact word matches.

Semantic search tries to understand that both are talking about the same idea.


🎓 Real-World Example: College Handbook

Imagine your college has a large student handbook.

It contains information about:

  • Attendance
  • Exams
  • Fees
  • Hostel
  • Leave
  • Scholarships
  • Library
  • Discipline

A student asks:

"Can I write the exam if my attendance is only 70%?"

The handbook may contain:

"Students must maintain at least 75% attendance to be eligible to appear for the examination."

The question and the document use different wording.

But they have the same meaning.

Semantic search can help identify this relevant information.

The process can look like this:

Student Question
      ↓
"What attendance is required for the exam?"
      ↓
Understand the meaning
      ↓
Search relevant information
      ↓
Attendance Rule
      ↓
Relevant Result

🧠 How Does Semantic Search Understand Meaning?

This is where our previous topic, Embeddings, becomes important.

We already learned that embeddings convert text into numerical representations that capture semantic meaning.

For example:

"What attendance is required?"
            ↓
        Embedding
            ↓
      [0.21, 0.74, ...]

And:

"Minimum attendance needed for examination"
            ↓
        Embedding
            ↓
      [0.23, 0.71, ...]

Because the two sentences have similar meanings, their vector representations can also be relatively close in vector space.

Semantic search uses this idea to find information that is semantically relevant.


🔗 Semantic Search and Embeddings

Let's connect this with what we already learned.

Step 1: Document

We have a college handbook.

Step 2: Chunking

The handbook is divided into smaller meaningful chunks.

Handbook
   ↓
Attendance Chunk
Exam Chunk
Hostel Chunk
Leave Chunk

Step 3: Embeddings

Each chunk is converted into an embedding.

Attendance Chunk
       ↓
   Embedding

Step 4: Store

The embeddings can be stored in a vector database.

Step 5: User Question

The student asks:

"What attendance do I need for the exam?"

The question is also converted into an embedding.

Step 6: Semantic Search

The system compares the question's meaning with the stored information.

Step 7: Relevant Chunk

The attendance-related chunk is retrieved.

So the overall flow becomes:

Document
   ↓
Chunking
   ↓
Embeddings
   ↓
Vector Database
   ↓
User Question
   ↓
Question Embedding
   ↓
Semantic Search
   ↓
Relevant Chunk
   ↓
LLM
   ↓
Final Answer

🛒 Another Real-World Example: Online Shopping

Semantic search is not limited to RAG.

Imagine you are shopping online.

You search:

"comfortable shoes for walking"

But a product description says:

"Lightweight sneakers designed for everyday walking and long-distance comfort."

The exact words may not be identical.

But the meaning is closely related.

Semantic search can understand the relationship between:

"comfortable shoes for walking"

and

"lightweight sneakers for everyday walking."

So it can help return relevant products.


🔎 Keyword Search Example

Suppose you search:

"car repair"

A keyword search mainly looks for pages containing words like:

car + repair

If a useful page says:

"vehicle maintenance and servicing"

it may not match as strongly because the exact words are different.


🧠 Semantic Search Example

Semantic search can recognize that:

car repair

and

vehicle maintenance

can be related concepts.

So even when the exact words are different, the system can find information with a similar meaning.


⚖️ Keyword Search vs Semantic Search

Let's make it simple.


🤖 Why Is Semantic Search Important in RAG?

RAG systems need to find the right information before asking the LLM to generate an answer.

Imagine a company has thousands of documents.

An employee asks:

"How many days can I take off for parental leave?"

The answer might be inside an HR policy document.

The employee may use completely different words from the document.

Semantic search helps connect the question with the relevant information based on meaning.

Employee Question
       ↓
"What leave can I take after having a baby?"
       ↓
Semantic Search
       ↓
HR Parental Leave Policy
       ↓
Relevant Context
       ↓
LLM
       ↓
Answer

This is one reason semantic search is an important part of many RAG systems.


🧩 Semantic Search in Simple Words

Think about talking to a friend.

You say:

"I'm feeling very tired today."

Your friend understands that you need rest.

You don't have to say:

"My energy level is currently low."

The exact words are different, but the meaning is clear.

Semantic search tries to achieve something similar when searching information.

It focuses on:

"What does this query mean?"

rather than only:

"Which exact words appear in the document?"


🔗 How Semantic Search Connects to Our Previous Topics

Our learning journey is now becoming connected.

We learned:

Embeddings

Embeddings represent the meaning of text as vectors.

Vector Database

A vector database stores these vector representations and supports similarity-based retrieval.

RAG

RAG retrieves relevant information and provides it to an LLM.

RAG Pipeline

The pipeline connects document processing, retrieval, and generation.

Chunking

Chunking divides large documents into smaller meaningful pieces.

Semantic Search

Semantic search helps find information based on meaning.

So now our flow looks like:


⚠️ Is Semantic Search Always Perfect?

No.

Semantic search can sometimes retrieve information that is related but not exactly what the user needs.

For example, a company may have separate policies for:

  • Maternity leave
  • Paternity leave
  • Parental leave

A question may be semantically related to all three.

The system still needs to retrieve the most relevant information.

This is why modern RAG systems can use additional techniques such as:

  • Similarity search
  • Metadata filtering
  • Hybrid search
  • Reranking

We will learn these concepts step by step.


📌 Simple Example to Remember

Imagine you search for:

"How do I fix my laptop?"

A document contains:

"Troubleshooting common notebook computer problems."

The words are different.

But the meaning is related.

Keyword Search:

"laptop" → looks for the word "laptop"

Semantic Search:

"laptop" → understands that "notebook computer" can refer to a similar concept

That is the basic idea behind semantic search.


🚀 Final Takeaway

Semantic Search is a search technique that focuses on the meaning of a query rather than only matching exact keywords.

In simple words:

Keyword Search → Find matching words

Semantic Search → Find matching meaning

In RAG, semantic search works together with:

Chunking → Embeddings → Vector Database → Retrieval → LLM

This helps an AI system find relevant information even when the user's question and the stored information use different words.


📌 One-Line Definition

Semantic Search = Finding relevant information based on meaning rather than exact keyword matching.


🎯 What Will We Learn Next?

Now we know that semantic search uses meaning to find relevant information.

But another important question comes up:

How does the system decide which result is more similar to the user's question?

That leads us to our next topic:

What Is Similarity Search? A Simple Guide for Beginners

What Is Reranking in RAG? A Simple Guide for Beginners

We have already learned how Retrieval , Semantic Search , Similarity Search , and Hybrid Search work in RAG. But there is one important qu...