In our previous blog post, we saw how AI can use external information with the help of Embeddings, Vector Databases, Retrieval, and RAG.
Now let's take one step deeper.
What exactly is an embedding?
You may have seen something like this:
«“Students must maintain 75% attendance.”»
How can a computer understand the meaning of this sentence and compare it with another sentence?
For example:
«“What percentage of attendance is required?”»
The words are different, but both sentences are talking about the same idea.
This is where embeddings become usefull
What Is an Embedding?
An embedding is a numerical representation of information that helps an AI system understand and compare its meaning.
Text can be converted into a list of numbers called a vector.
For example:
«“Students must maintain 75% attendance.”»
can be converted into something conceptually like:
«"[0.12, -0.45, 0.78, 0.31, ...]"»
These numbers are not meant for humans to read.
They are used by AI systems to represent information in a mathematical form.
So the basic idea is:
Text
⬇️
Embedding Model
⬇️
Numbers / Vector
The actual vectors can contain many more dimensions than the simple example above.
Why Does AI Need Embeddings?
Computers work very well with numbers.
But humans communicate using words.
For example:
«“I love dogs.”»
and:
«“Dogs are my favorite animals.”»
These sentences use different words, but their meanings are related.
A good embedding system can represent these sentences in a way that allows their semantic similarity to be measured.
That's the important idea.
«Embeddings help AI work with the meaning of information using numbers.»
A Real-Time Example: College Handbook 🎓
Let's continue with the same example from our previous post.
Imagine a college has a large student handbook.
The handbook contains:
- Attendance rules
- Exam rules
- Fee information
- Hostel rules
- Library rules
- Leave policies
One section says:
«“Students must maintain at least 75% attendance to be eligible for the final examination.”»
Now a student asks:
«“How much attendance do I need to write my final exam?”»
Notice something interesting.
The student's question doesn't use exactly the same words as the handbook.
The handbook says:
«“75% attendance”»
The student asks:
«“How much attendance do I need?”»
The wording is different.
But the meaning is related.
Embeddings help systems work with this kind of semantic relationship.
How Does This Work?
Let's simplify the process.
Step 1: Take the Text
The system receives:
«“Students must maintain at least 75% attendance.”»
Step 2: Create an Embedding
An embedding model converts the text into a numerical vector.
Conceptually:"[0.12, -0.45, 0.78, ...]
Step 3: Store the Vector
The embedding can be stored in a system such as a vector database.
Step 4: User Asks a Question
The student asks:
«“How much attendance do I need?”»
This question can also be converted into an embedding.
Step 5: Compare Meaning
The system searches for information whose embedding is relevant or similar to the question.
It can find:
«“Students must maintain at least 75% attendance.”»
Step 6: Use the Information
The retrieved information can then be passed to an LLM.
The LLM can generate:
«“You need at least 75% attendance to be eligible for the final examination.”»
Embeddings Are Not Just for Text
Embeddings are not limited to sentences.
AI systems can create embeddings for different types of information, depending on the model and application.
For example:
📝 Text - Articles, documents, questions, emails, etc.
🖼️ Images- Images can also be represented in vector form.
🎵 Audio- Audio information can also be represented as embeddings.
This allows AI systems to compare and search different types of information.
Another Real-Time Example: Online Shopping 🛒
Imagine you are using an online shopping application.
You search:
«“Comfortable shoes for running.”»
The system doesn't necessarily need to find only products containing the exact words “comfortable” and “running.”
It can use semantic representations to find products related to your search.
For example, it might find:
«“Lightweight running shoes with cushioned soles.”»
The wording is different.
But the meaning is related to what you searched for.
Embeddings can help make this kind of semantic search possible.
Embeddings vs Keywords
This is an important difference.
Suppose you search:
«“How can I learn programming?”»
A simple keyword-based system may focus heavily on words such as:
- learn
- programming
But an embedding-based system can help identify content that is semantically related, such as:
«“A beginner's guide to coding.”»
The exact words are different, but the meaning is similar.
That's one reason embeddings are useful for modern AI search systems.
For example:
«Dog 🐶»
might be closer to:
«Puppy»
than to:
«Car»
This is a simplified analogy, but it gives us an intuitive idea of how embeddings can represent relationships between concepts.
What Is a Vector?
You will often hear the word vector when learning about embeddings.
A vector is simply a collection of numbers.
For example:
«"[0.2, 0.8, -0.4, 0.6]"»
In AI, an embedding is represented as a vector.
Real AI systems can use vectors with many dimensions.
You don't need to understand the mathematics immediately.
For now, remember:
«Embedding = numerical representation»«Vector = the collection of numbers used to represent it»
How Do We Know Two Embeddings Are Similar?
Suppose we have two sentences:
Sentence A:
«“I want to learn programming.”»
Sentence B:
«“I want to learn coding.”»
Their words aren't exactly identical.
But their meanings are very similar.
The vectors created from these sentences may therefore be positioned relatively close together in the embedding space.
AI systems can use mathematical methods such as cosine similarity to compare vectors.
You don't need to understand the formula yet.
The simple idea is:
«Closer / more similar vectors → potentially more similar meaning»
This is one of the ideas behind semantic search.
Now let's connect this to what we learned in the previous post.
Remember our flow?
«Documents → Chunks → Embeddings → Vector Database → Retrieval → LLM → Answer»
Embeddings are an important part of this process.
This is one of the ways embeddings can support a RAG system.
Why Are Embeddings Important?
Embeddings are useful because they allow AI applications to work with information based on meaning, rather than relying only on exact words.
They can help with:
- Semantic search
- Document search
- Recommendation systems
- Question answering
- RAG applications
- Similarity search
- Image search
- Content matching
This makes embeddings an important building block in many modern AI applications.
One More Simple Example
Imagine you have 10,000 documents.
You ask:
«“What is the company's work-from-home policy?”»
You don't want the AI to read all 10,000 documents every time.
Instead, the system can:
1. Convert documents into embeddings.
2. Store them.
3. Convert your question into an embedding.
4. Search for relevant vectors.
5. Retrieve the most relevant information.
6. Give that information to the LLM.
The LLM can then generate a useful answer.
This is one of the ways embeddings help make AI-powered knowledge systems practical.
What Embeddings Do NOT Mean
It's important not to misunderstand embeddings.
An embedding isn't simply:
«“The meaning of a sentence stored perfectly inside a list of numbers.”»
It's better to think of it as a learned numerical representation that captures useful patterns and relationships for a particular model and task.
Different embedding models can represent information differently.
So embeddings are not a magical universal dictionary of meaning.
They are a tool that helps AI systems perform tasks such as similarity search and retrieval.
The Big Picture
Let's summarize the idea.
This allows AI applications to work with information in a more meaning-oriented way.
Conclusion
So, what are embeddings?
«Embeddings are numerical representations of information that help AI systems compare and work with meaning and relationships.»
For example:
«“How much attendance do I need?”»
and:
«“Students must maintain at least 75% attendance.”»
use different words, but they are related in meaning.
Embeddings help AI systems represent these pieces of information in a form that can be compared mathematically.
And when embeddings are combined with vector databases, retrieval, and LLMs, they become an important part of systems such as RAG.
The next question naturally becomes:
«“If embeddings are vectors, where do we store all these vectors, and how do we search them efficiently?”»
That's where our next topic comes in:
What Is a Vector Database? 🗄️
In the next post, we'll understand vector databases with the same college handbook example, so the whole RAG architecture becomes easier to visualize.



No comments:
Post a Comment