Python Embeddings and Semantic Search: Build It from Scratch

Turn text into vectors, store them in a NumPy matrix, and rank by cosine similarity. A complete semantic search you can run locally, with tests.

Executive Summary: Semantic search in Python needs an embedding model, a NumPy matrix of vectors, and a dot product to rank results, about sixty lines that scale to roughly a hundred thousand documents. This post builds it from scratch, explains where vectors miss exact strings like error codes, and shows when a vector database is actually needed.

A help center has an article titled “Refunds are issued to the original payment method”. A customer types “how do I get my money back” into the search box and gets zero results. Every word in the query is ordinary, and none of them appears in the article. Keyword search compares letters. The customer was comparing meaning.

Semantic search in Python closes that gap with embeddings. An embedding is a list of numbers, a vector, produced by a model so that texts with similar meanings get vectors that point in similar directions. If the idea is new, read about what embeddings are first. Search then becomes geometry: embed the query, and find the stored vectors nearest to it.

This guide uses Python 3.12 and NumPy. The main example runs a small open source model on your own machine, so it needs no API key and costs nothing. A second option calls a hosted embedding API, which does cost money. The tests use a fake embedder and need neither.

My position: you do not need a vector database to start. You need to understand the sixty lines that a vector database replaces, because that understanding is what lets you debug bad results later.

How semantic search works

INDEXING (once, and again when documents change)

  documents  --->  embedding model  --->  matrix of vectors
  "Refunds are issued..."                 [ 0.03, -0.11, ... ]   row 0
  "To reset your password..."             [ 0.08,  0.02, ... ]   row 1

SEARCHING (every query)

  "how do I get my money back"  --->  same embedding model  --->  query vector
                                                                      |
                    similarity of the query to every row  <----------+
                                     |
                                     v
                    sort by score, return the top k documents

Two rules follow from the picture. The same model must embed both documents and queries, because vectors from different models live in unrelated spaces. And the quality of the search can never exceed the quality of the embedding model.

Cosine similarity is a dot product

The standard measure of closeness is cosine similarity: the cosine of the angle between two vectors. It ranges from 1 (same direction) through 0 (unrelated) to -1 (opposite). The formula divides the dot product by the two vector lengths.

If you scale every vector to length 1 in advance, a step called normalization, the division disappears. Cosine similarity becomes a plain dot product, and comparing one query against a whole matrix is a single NumPy operation:

scores = matrix @ query_vector      # one score per document

That line is the entire search engine. Everything else is bookkeeping.

Set up

python --version
python -m venv .venv
source .venv/bin/activate          # macOS and Linux
.venv\Scripts\Activate.ps1         # Windows PowerShell
python -m pip install numpy sentence-transformers pytest

sentence-transformers pulls in PyTorch, so the install is large, and the first run downloads a model of roughly 90 MB. After that, everything works offline. If you only want to run the tests, numpy and pytest are enough.

Build semantic search in Python

Save this complete file as search.py. It defines three interchangeable embedders and one index class.

import hashlib
import re
import sys
from typing import Protocol

import numpy as np

LOCAL_MODEL = 'all-MiniLM-L6-v2'
OPENAI_MODEL = 'text-embedding-3-small'

DOCS = [
    'To reset your password, open Settings and choose Security.',
    'Refunds are issued to the original payment method within five days.',
    'You can export all tasks as a CSV file from the Reports page.',
    'Invite teammates by entering their email address on the Members page.',
    'Two-factor authentication adds a code from your phone at sign-in.',
]


class Embedder(Protocol):
    def embed(self, texts: list[str]) -> np.ndarray: ...


class LocalEmbedder:
    """Runs an open source model on this machine. No key, no cost."""

    def __init__(self, model_name: str = LOCAL_MODEL) -> None:
        from sentence_transformers import SentenceTransformer

        self._model = SentenceTransformer(model_name)

    def embed(self, texts: list[str]) -> np.ndarray:
        return np.asarray(self._model.encode(texts), dtype=np.float32)


class OpenAIEmbedder:
    """Calls a hosted API. Needs OPENAI_API_KEY, and each call is billed."""

    def __init__(self, model: str = OPENAI_MODEL) -> None:
        from openai import OpenAI

        self._client = OpenAI(timeout=30.0, max_retries=2)
        self._model = model

    def embed(self, texts: list[str]) -> np.ndarray:
        response = self._client.embeddings.create(model=self._model, input=texts)
        return np.asarray([item.embedding for item in response.data], dtype=np.float32)


class FakeEmbedder:
    """Deterministic word-count vectors for tests. No model, no network."""

    def __init__(self, dimensions: int = 256) -> None:
        self.dimensions = dimensions

    def embed(self, texts: list[str]) -> np.ndarray:
        vectors = np.zeros((len(texts), self.dimensions), dtype=np.float32)
        for row, text in enumerate(texts):
            for word in re.findall(r'[a-z0-9]+', text.lower()):
                digest = hashlib.sha256(word.encode('utf-8')).digest()
                vectors[row, int.from_bytes(digest[:4], 'big') % self.dimensions] += 1.0
        return vectors


def normalize(vectors: np.ndarray) -> np.ndarray:
    norms = np.linalg.norm(vectors, axis=1, keepdims=True)
    norms[norms == 0] = 1.0
    return vectors / norms


class SemanticIndex:
    def __init__(self, embedder: Embedder) -> None:
        self._embedder = embedder
        self._texts: list[str] = []
        self._vectors: np.ndarray | None = None

    def add(self, texts: list[str]) -> None:
        new_vectors = normalize(self._embedder.embed(texts))
        if self._vectors is None:
            self._vectors = new_vectors
        else:
            self._vectors = np.vstack([self._vectors, new_vectors])
        self._texts.extend(texts)

    def search(self, query: str, k: int = 3) -> list[tuple[float, str]]:
        if self._vectors is None:
            return []
        query_vector = normalize(self._embedder.embed([query]))[0]
        scores = self._vectors @ query_vector
        top = np.argsort(scores)[::-1][:k]
        return [(float(scores[i]), self._texts[i]) for i in top]


def main() -> None:
    query = ' '.join(sys.argv[1:]) or 'How do I get my money back?'
    index = SemanticIndex(LocalEmbedder())
    index.add(DOCS)
    for score, text in index.search(query):
        print(f'{score:.2f}  {text}')


if __name__ == '__main__':
    main()
python search.py

Example output from my machine. Your scores may differ slightly:

0.48  Refunds are issued to the original payment method within five days.
0.13  To reset your password, open Settings and choose Security.
0.09  You can export all tasks as a CSV file from the Reports page.

The refund article wins by a wide margin, although the query shares no meaningful word with it. Try your own queries by passing them as arguments, such as python search.py "add a colleague to my workspace".

What each part does

  • The Embedder protocol says that anything with an embed method fits. The index does not know or care which model produced the vectors, which is what makes the fake embedder possible.
  • normalize scales each row to length 1, so the later dot product equals cosine similarity. The guard against zero norms avoids a division by zero for empty text.
  • add embeds texts in one batch. Batching matters: one call for 500 texts is far faster than 500 calls.
  • search embeds the query, multiplies, and sorts. np.argsort(scores)[::-1][:k] gives the positions of the k highest scores.

The imports for the two real embedders sit inside their constructors. As a result, the tests can import this file on a machine that has neither package installed.

Choosing an embedding model

Model Runs Dimensions Trade-off
all-MiniLM-L6-v2 Locally, through sentence-transformers 384 Free, private, and fast. Smaller and mostly English-focused
text-embedding-3-small OpenAI API 1,536, and can be shortened Paid per token. No model to host
text-embedding-3-large OpenAI API 3,072, and can be shortened Higher quality at higher cost and storage
text-embedding-ada-002 OpenAI API 1,536 The previous generation. Prefer the newer models for new projects

The text-embedding-3 models were announced last month. They accept a dimensions parameter that returns shorter vectors, which trades a little accuracy for less storage. Model names change quickly, so keep the name in one constant, as the script does.

To use the hosted model, install openai, set OPENAI_API_KEY in your environment, and replace LocalEmbedder() with OpenAIEmbedder(). Remember that this sends your documents to a third party and bills you for every token. Configure the client carefully, as described in this guide to an OpenAI client with timeouts and retries.

Switching models is not free. Vectors from two models are not comparable, so a change of model means embedding every document again. Store the model name next to the vectors, and refuse to search an index built with a different one.

The myth: semantic search needs a vector database

Most tutorials start by installing a vector database. For a large share of real projects, that is unnecessary, and you can prove it with multiplication. This is the calculation I do before any infrastructure decision:

memory = number of vectors x dimensions x 4 bytes (float32)
Vectors 384 dimensions 1,536 dimensions
10,000 15 MB 61 MB
100,000 154 MB 614 MB
1,000,000 1.5 GB 6.1 GB

A hundred thousand vectors fit in the memory of a small server. Brute-force search over them is one matrix-vector product. Measure it yourself with this illustrative script:

import time

import numpy as np


def main() -> None:
    rng = np.random.default_rng(0)
    matrix = rng.standard_normal((100_000, 384), dtype=np.float32)
    matrix /= np.linalg.norm(matrix, axis=1, keepdims=True)
    query = matrix[0]

    start = time.perf_counter()
    for _ in range(20):
        scores = matrix @ query
        top = np.argpartition(scores, -5)[-5:]
    elapsed = (time.perf_counter() - start) / 20
    print(f'{elapsed * 1000:.1f} ms per search over {len(matrix):,} vectors')


if __name__ == '__main__':
    main()

On my laptop, each search takes a few milliseconds. In practice, embedding the query takes longer than searching the matrix. np.argpartition finds the top five without sorting everything, which helps at this size.

So when do you need more? My heuristic: stay with NumPy below roughly 100,000 vectors, and move when one of these becomes true.

  • The matrix no longer fits in memory.
  • Documents change constantly, and rebuilding a file is too slow.
  • You must filter by metadata, such as tenant or date, together with similarity.
  • Several processes need to read and write the index at once.
Option What it is Choose it when
NumPy matrix Exact brute-force search in memory Up to about 100,000 vectors and simple needs
FAISS A library of fast exact and approximate indexes Millions of vectors in one process
Chroma An embedded vector store with metadata Prototypes and small applications that want persistence
pgvector A PostgreSQL extension for vector columns Your data already lives in PostgreSQL and you need filters and transactions

Approximate indexes trade a small loss of accuracy for a large gain in speed. Below a million vectors, exact search is often fast enough, and it has no tuning parameters to get wrong.

Save and load the index

Embedding is the expensive step, so do it once and keep the result. NumPy’s own format is enough.

import json
from pathlib import Path

import numpy as np


def save_index(folder: Path, model: str, texts: list[str], vectors: np.ndarray) -> None:
    folder.mkdir(parents=True, exist_ok=True)
    np.save(folder / 'vectors.npy', vectors)
    payload = {'model': model, 'texts': texts}
    (folder / 'meta.json').write_text(json.dumps(payload), encoding='utf-8')


def load_index(folder: Path, expected_model: str) -> tuple[list[str], np.ndarray]:
    payload = json.loads((folder / 'meta.json').read_text(encoding='utf-8'))
    if payload['model'] != expected_model:
        raise ValueError(f"index was built with {payload['model']}, not {expected_model}")
    return payload['texts'], np.load(folder / 'vectors.npy')

The model check is the important part. Without it, a model upgrade silently produces nonsense rankings, because the new query vectors are compared with old document vectors.

Real documents usually arrive as a table. Loading a CSV into a pandas DataFrame and passing df['body'].tolist() to index.add() works well. Keep the row IDs alongside the texts, so that a search result leads back to a record.

Test without a model or a network

You cannot assert exact scores from a real embedding model, since a model update changes them. Test your own logic instead: ranking, limits, and edge cases. The FakeEmbedder turns words into counts in fixed positions, so texts that share words score higher. Save this as test_search.py.

from search import DOCS, FakeEmbedder, SemanticIndex


def make_index() -> SemanticIndex:
    index = SemanticIndex(FakeEmbedder())
    index.add(DOCS)
    return index


def test_best_match_shares_the_query_words():
    results = make_index().search('reset password')
    assert results[0][1].startswith('To reset your password')


def test_results_are_sorted_and_limited():
    results = make_index().search('export tasks to a file', k=2)
    scores = [score for score, _ in results]
    assert len(results) == 2
    assert scores == sorted(scores, reverse=True)


def test_empty_index_returns_nothing():
    assert SemanticIndex(FakeEmbedder()).search('anything') == []


def test_identical_text_scores_one():
    results = make_index().search(DOCS[1], k=1)
    assert abs(results[0][0] - 1.0) < 1e-5
python -m pytest -q test_search.py
....                                                                     [100%]
4 passed in 0.15s

The fake embedder understands no meaning, and that is fine. It is deterministic, instant, and free, which is what a unit test needs. Checking the quality of a real model is a separate job, done with a small set of real queries and expected documents.

Where semantic search goes wrong

  • Exact strings. Embeddings blur identifiers. A search for error code E4012 or a product SKU may rank a similar-looking code above the right one. Keyword search handles these far better.
  • Long documents. One vector for a 30-page document averages many topics into mush. Split long texts into passages of a few hundred words, and embed each passage.
  • No natural cutoff. The search always returns k results, even when nothing is relevant. A score that means “good match” for one model means nothing for another, so tune any threshold on your own data.
  • Negation and numbers. “Supports refunds” and “does not support refunds” land close together. Do not rely on embeddings for logic.

Because of the first point, production search is usually hybrid. It runs a keyword search and a vector search, and merges the two result lists.

How real systems use embeddings

  • Offline indexing, online querying. A batch job embeds new and changed documents. The request path embeds only the query.
  • Content hashing. Systems store a hash of each text and skip texts that have not changed, which avoids paying to embed the same content twice.
  • Passage-level vectors with metadata. Each vector carries a document ID, a position, and fields such as language or tenant, so results can be filtered and traced.
  • Hybrid ranking. Keyword scores and vector scores are combined, which protects exact-match queries.
  • A golden query set. Teams keep a list of real queries with the documents that should appear, and they rerun it whenever the model or the splitting rules change.

A mistake I have seen in production is an index that was built with one embedding model and queried with another after a routine configuration change. Nothing crashed. Search quality simply collapsed, and results looked random. It took two days to connect the complaints to the change, because both models returned vectors of the same length. Since then, I store the model name with every index and fail loudly on a mismatch.

Choosing your approach: a decision framework

  1. Do users search for exact names, codes, or IDs? Start with keyword search. Add vectors later as a second signal.
  2. Do users describe what they want in their own words? Use embeddings.
  3. Can your documents leave your infrastructure? If not, use a local model.
  4. How many vectors will you store? Multiply by dimensions and by 4 bytes. If the result fits in memory, start with NumPy.
  5. Do you need filters, concurrent writes, or millions of vectors? Move to pgvector, FAISS, or a vector store.
  • Structured lookups. Finding orders by customer and date is a database query. Embeddings add cost and remove precision.
  • Small, well-labeled collections. Twenty help articles with good titles are served well by a list and a keyword filter.
  • Queries that need exact matching. Legal citations, part numbers, and log message IDs must match character for character.

Common mistakes

  • Mixing embedding models. Documents embedded with one model and queries with another produce meaningless scores, with no error.
  • Forgetting to normalize. A raw dot product favors long vectors, so rankings reflect vector length instead of meaning.
  • Embedding one text per call. Thousands of small calls are slow and hit rate limits. Send texts in batches.
  • Embedding whole documents. A single vector for a long text matches everything weakly. Split into passages.
  • Treating the score as a probability. A score of 0.8 is not “80 percent relevant”. Scores are comparable only within one model and one collection.
  • Re-embedding unchanged content. Without content hashes, every index run pays for the full collection again.

Key takeaways

  • Semantic search is: embed the documents, embed the query, and rank by similarity.
  • Normalize vectors once, and cosine similarity becomes a dot product.
  • A NumPy matrix handles about 100,000 vectors with millisecond searches.
  • Use the same embedding model for documents and queries, and store its name with the index.
  • Local models are free and private. Hosted models cost per token and need no hosting.
  • Vectors miss exact strings, so combine them with keyword search where IDs matter.
  • Test your logic with a fake embedder, and evaluate quality with real queries.

FAQ

What is semantic search?

Semantic search finds documents by meaning instead of matching words. It converts text into vectors with an embedding model and returns the documents whose vectors are closest to the query’s vector.

How do I calculate cosine similarity in Python?

With NumPy, divide the dot product of two vectors by the product of their norms. If you normalize the vectors to length 1 first, the cosine similarity is just a @ b.

Do I need a vector database for semantic search?

Not at small scale. A NumPy matrix searches about 100,000 vectors in milliseconds. A vector database becomes useful for millions of vectors, metadata filters, or frequent concurrent updates.

Which embedding model should I use?

Use a local model such as all-MiniLM-L6-v2 when you want zero cost and privacy. Use a hosted model such as text-embedding-3-small when you prefer not to run a model yourself.

What is the difference between semantic search and keyword search?

Keyword search matches the words in the query against the words in documents. Semantic search matches meaning, so it finds related text with different wording, but it is weaker at exact strings such as codes and IDs.

Start with a matrix and a dot product

Semantic search looks like specialized infrastructure and is mostly linear algebra you can read in one sitting. Build the small version, check its results against real queries, and let measured limits tell you when to add more. The sixty lines stay useful either way, because every vector store does the same thing at a larger scale.

Rule of thumb: if your vectors fit in memory, your search engine is one line of NumPy, and the hard part is choosing what to embed.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *