Split, store and search your documents

You will be able to chunk documents, store their embeddings with metadata, and retrieve the best matches.

Once Daniel understood embeddings, he tried embedding each whole policy document as one vector. The search got worse, not better. A question about childcare leave matched the 40-page leave policy, which was right, but then he had to send all 40 pages to the model, which was the problem he started with. One vector for a whole document blurs every topic in it into one average meaning.

So before you embed anything, you cut documents into smaller pieces. Then you store each piece carefully enough that you can find it again and say where it came from.

Split documents into chunks

A chunk is a piece of a document small enough to be about one thing, usually a few paragraphs. Each chunk gets its own embedding, and search returns chunks, not documents.

Size is a trade-off. Chunks that are too small, like single sentences, lose context: "This applies only to confirmed staff" means nothing without the paragraph before it. Chunks that are too large go back to blurring several topics into one vector. A few paragraphs is a sensible starting point, and you adjust after testing.

Split along the document's own structure where you can. Headings, sections and numbered clauses are natural boundaries, because the author already grouped related text under them. Daniel's handbook has clear section headings, so each section became one chunk, and very long sections were split into a few paragraphs at a time.

Add overlap between neighbouring chunks, repeating the last part of one chunk at the start of the next. Without overlap, an answer that spans a boundary gets cut in half, with the rule in one chunk and its exception in another, and search may only return one of them. A sentence or two of overlap is usually enough.

Most document libraries can split by headings or by length with overlap. Ask your coding assistant to use one rather than writing the splitting yourself.

Store every chunk with its source

Each stored chunk needs more than text and a vector. It needs metadata, labels describing where it came from: the file name, the page number or section heading, the date the document was issued, and anything else you might want to filter by later, such as department.

Metadata does two jobs. It lets answers cite their sources, which lesson 3.4 relies on, because "see Claims Policy, section 4.2, page 7" is something a staff member can check. And it lets you filter before searching. When Daniel adds the 2025 claims policy, he can mark the 2024 version as superseded and exclude it, so the tool never answers from an old rule.

Where the vectors live

You need somewhere to store vectors and search them by similarity. There are two common choices.

A vector database is built for this job. Several exist, some hosted for you and some you run yourself, and most have free tiers that are enough for a practice project.

A database extension adds vector search to a database you may already use. The best-known is pgvector, which adds a vector column type and similarity search to PostgreSQL. If your app already stores data in PostgreSQL, keeping the vectors in the same database means one less service to run.

For a few hundred chunks, either choice is fine, and even a simple file loaded into memory works for practice. Whatever you choose, the search call looks the same: give it the question's vector, ask for the top few matches, and get back the chunk text with its metadata. Daniel fetches the top five. Fewer risks missing the answer. Many more fills the prompt with noise and cost.

Add keyword search for exact terms

Embeddings struggle with text whose meaning is in the exact characters: a form number like HR-07, an invoice reference, a person's name, a specific dollar figure. To an embedding model, HR-07 and HR-08 look almost identical, but to a staff member they are different forms.

Plain keyword search is good at exactly that. So many retrieval systems run both searches and combine the results, an approach called hybrid search. A question like "Where do I find form HR-07?" gets the exact match from keyword search, while "What can I claim after working late?" gets the meaning match from vector search. Many vector databases and PostgreSQL itself offer keyword search, so this rarely needs another tool.

Daniel tested both on his own questions before deciding. Vector search alone missed every question that named a form number. With hybrid search, those came back first.

Now try it on your own documents. In the activity below you chunk one document, store it with metadata, and look at what comes back for each of your five questions from lesson 3.1.

Chunk one of your documents, store the chunks with metadata, and print the top three matches for each of your five questions.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).