You will be able to explain what an embedding is and how similarity search uses it.
Daniel's first attempt at search was the one built into his file storage. A staff member asked "Can I get reimbursed for a Grab ride home after overtime?" and the search came back empty. The policy existed. It just said "transport claims for employees working beyond 10pm" and never used the words reimbursed, Grab, ride or overtime. Keyword search matches words. The question and the answer meant the same thing in different words, so it found nothing.
Retrieval needs a search that matches meaning, and embeddings are how you get one.
An embedding is a list of numbers that stands for the meaning of a piece of text. You get one by sending the text to an embedding model, a different kind of model from the one that writes replies. Most major providers offer one through their API, and open-source embedding models exist that you can run yourself.
You send a sentence, a paragraph or a short passage, and you get back a long list of numbers, often hundreds or thousands of them. That list is called a vector. On its own, no single number means anything you could name. What matters is how one vector compares with another.
The embedding model was trained on enormous amounts of text so that texts with similar meaning produce similar vectors. "Transport claims for employees working beyond 10pm" and "Can I get reimbursed for a Grab ride home after overtime" share almost no words, but they are about the same thing, so their vectors come out close together. "The office pantry is restocked on Mondays" comes out far from both.
It helps to picture a simpler version. Imagine every piece of text placed as a dot on a map, where dots that mean similar things sit near each other. Leave policies cluster in one area, claims in another, IT rules in a third. A question about a taxi home lands in the claims area, near the transport policy, even if it uses none of the same words.
Real embeddings are like that map, but with hundreds or thousands of directions instead of two. You cannot picture it, but the maths of distance and direction works the same way.
This is also why embeddings handle different phrasings well. "Childcare leave", "time off to look after my kid" and "days off for parents" land near one another. Many embedding models also place text in different languages near its translation, so a question in Chinese can find a passage in English. Check your model's documentation before relying on that.
Searching with embeddings has two parts.
Ahead of time, you embed every passage in your document set and store each vector alongside the passage text. Daniel did this once for the whole handbook, and he repeats it for any document that changes.
When a question arrives, you embed the question with the same model. Then you compare the question's vector with every stored vector and return the passages whose vectors are closest.
The most common way to measure closeness is cosine similarity, which looks at the direction each vector points in and ignores how long it is. Vectors pointing exactly the same way score 1, and vectors at right angles to each other score 0, which means they have nothing in common. In practice, related passages score noticeably higher than unrelated ones, and you take the top few. The exact scores depend on the model, so judge them relative to each other rather than against a fixed cut-off you read somewhere.
You do not write this comparison yourself. Vector databases and search libraries do it for you, quickly, even across millions of passages. Lesson 3.3, Split, store and search your documents, covers where the vectors live.
There is one rule you cannot break. The question and the documents must be embedded with the same embedding model. Different models produce different maps. A vector from one model compared with a vector from another is like comparing a postcode with a phone number: both are numbers, but the comparison means nothing. The vectors may not even be the same length.
So if you move to a newer embedding model later, every stored passage has to be embedded again, including the ones that have not changed. Record which model you used next to your stored vectors, so that you, or the next person, can tell.
Embeddings are also not magic. They are weak with exact strings that carry little meaning on their own: product codes, invoice numbers, people's names, specific figures. Lesson 3.3 shows how to cover that gap with plain keyword search alongside.
The best way to build intuition is to see the scores for yourself. In the activity below you embed a few sentences and questions and look at which ones come out closest.
Embed ten short sentences and three questions with an embedding API, and check which sentences come back closest to each question.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).