Contents
Every semantic search tutorial starts with the same broken keyword-search example: searching "doctor visit" and getting zero results for a document that says "physician appointment." That failure is exactly the itch Sentence Transformers exist to scratch, and it's genuinely one of the most satisfying "aha" moments in NLP once you see it work firsthand.
I've used this library across several projects specifically because it's free, runs entirely locally, and doesn't require an API key or a monthly bill. That combination makes it my default recommendation whenever someone asks how to get started with semantic search without committing to a paid service first.
By the end of this tutorial, you'll have a working semantic search system running on your own machine, and you'll actually understand why it finds "physician" when you searched "doctor." IMO, that understanding matters more than the code itself :)
Figure 1: Sentence Transformers enable semantic search by understanding text meaning
What Sentence Transformers Actually Are
Sentence Transformers, often abbreviated SBERT, is a Python library for computing dense vector embeddings from text — sentences, paragraphs, even images and audio depending on the model. Unlike raw BERT, which produces token-level embeddings that don't play nicely with sentence-level comparison, SBERT is specifically fine-tuned to produce embeddings where cosine similarity directly correlates with semantic similarity.
That distinction matters more than it sounds. Averaging raw BERT token outputs gives you poor semantic representations — similar sentences can end up with surprisingly dissimilar vectors. SBERT fixes this by training on natural language inference pairs using a siamese network architecture, producing embeddings that actually behave the way you'd want for search and comparison tasks.
- Over 25,000 pre-trained models available immediately through Hugging Face.
- Works entirely locally, no API calls or per-token costs required.
- Supports semantic search, clustering, paraphrase mining, and reranking, all through one library.
Ever wondered why "physician" and "doctor" land close together in vector space despite sharing zero letters? This is exactly the mechanism responsible.
Installing the Library
Getting set up takes about thirty seconds, assuming your Python environment is already in order.
python -m venv env
source env/bin/activate # Windows: env\Scripts\activate
pip install sentence-transformers
A virtual environment isn't optional here in my opinion. This library pulls in PyTorch and a handful of other dependencies, and isolating your project avoids the classic "why did installing one package break three other things" headache.
Loading Your First Model
Let's load a lightweight, genuinely fast model and confirm everything's working.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
print("Model loaded successfully")
all-MiniLM-L6-v2 is a great starting point — small, fast, and surprisingly capable for its size. It's the model I reach for first in almost every prototype, since it strikes a genuinely good balance between speed and quality.
Encoding Some Sentences
Now let's actually convert text into vectors and see how similarity plays out.
from sentence_transformers import util
sentences = [
"The quick brown fox jumps over the lazy dog",
"A fast auburn fox leaps above a sleepy canine",
"The stock market closed higher today",
]
embeddings = model.encode(sentences, convert_to_tensor=True)
cos_sim = util.cos_sim(embeddings, embeddings)
print(f"Sentences 0 and 1 similarity: {cos_sim[0][1]:.4f}")
print(f"Sentences 0 and 2 similarity: {cos_sim[0][2]:.4f}")
Run this, and you'll see the first two sentences score high similarity despite sharing almost no exact words — they mean the same thing, just phrased differently. The stock market sentence, unsurprisingly, scores far lower against both. That gap between the numbers is the entire value proposition of semantic search in one glance.
Building an Actual Semantic Search Function
Encoding pairs is neat, but let's build something genuinely useful: searching a small corpus for the most relevant match to a query.
corpus = [
"Python is a high-level programming language",
"Machine learning requires large datasets",
"Neural networks are inspired by the human brain",
"Flask is a lightweight web framework",
]
corpus_embeddings = model.encode(corpus, convert_to_tensor=True)
def semantic_search(query, top_k=2):
query_embedding = model.encode(query, convert_to_tensor=True)
hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=top_k)
for hit in hits[0]:
print(f"{corpus[hit['corpus_id']]} (score: {hit['score']:.4f})")
semantic_search("web development frameworks")
Notice that util.semantic_search helper — it handles the similarity computation and ranking for you, sparing you from writing that logic by hand every time. For small corpora, this is genuinely all you need. No FAISS, no vector database, just Python and PyTorch tensors doing the work.
Why This Beats Keyword Search
If your corpus contained "Flask is a lightweight web framework" and someone searched "web development frameworks," a basic keyword matcher might miss it entirely — zero exact word overlap beyond "web." Semantic search finds it anyway, because the model understands these phrases occupy similar conceptual territory, not just similar spelling.
Symmetric vs. Asymmetric Search
This is the part beginners consistently skip, and it genuinely changes which model you should pick. There are two distinct flavors of semantic search, and treating them the same hurts your results.
- Symmetric search: your query and corpus entries are roughly the same length and content type — think "find similar questions." Searching "How to learn Python online?" to find "How to learn Python on the web?" is symmetric.
- Asymmetric search: a short query paired with longer content — think "What is Python?" matching against a full paragraph explaining Python. This is the far more common case in real RAG and FAQ systems.
Using a model tuned for symmetric search on an asymmetric task (or vice versa) quietly hurts retrieval quality, even though nothing throws an error. I made this mistake early on and spent longer than I'd like to admit wondering why results felt "off" before realizing the model-task mismatch was the actual culprit.
Choosing the Right Model for Your Task
With over 25,000 pre-trained models available, picking one can feel paralyzing. Here's a practical shortlist to actually start from.
- all-MiniLM-L6-v2 — fast, lightweight, great for prototyping and lower-resource environments.
- all-mpnet-base-v2 — noticeably higher quality, at the cost of somewhat slower inference; a strong general-purpose default.
- Domain-specific and multilingual variants exist too — check the Hugging Face model hub if your content isn't standard English text.
My honest take? Start with MiniLM to get things working, then upgrade to mpnet if retrieval quality genuinely becomes the bottleneck. Don't reach for the heaviest model on day one just because it scores higher on a benchmark you haven't validated against your actual data.
Building a Real FAQ Search Example
Let's put this together into something closer to a real use case: a healthcare FAQ matcher, where users phrase the same question a dozen different ways.
faqs = [
"How do I schedule a doctor appointment?",
"What insurance plans do you accept?",
"Where can I pick up my lab results?",
"How do I reset my patient portal password?",
]
faq_embeddings = model.encode(faqs, convert_to_tensor=True)
def find_best_faq(user_question, threshold=0.5):
query_embedding = model.encode(user_question, convert_to_tensor=True)
hits = util.semantic_search(query_embedding, faq_embeddings, top_k=1)[0]
if hits[0]["score"] < threshold:
return "No confident match found — consider routing to a human."
return faqs[hits[0]["corpus_id"]]
print(find_best_faq("Can I see a physician this week?"))
Notice how "physician" correctly matches "doctor appointment" even though those exact words never overlap. That's semantic search doing exactly what it's supposed to do — understanding intent rather than pattern-matching text.
That Threshold Check Matters
I added that threshold check deliberately. Without it, your search will confidently return the "best" match even when nothing in your corpus is actually relevant — that's a recipe for confusing, unhelpful results. Better to admit uncertainty and route to a human than to fake confidence with a genuinely poor match.
Scaling Beyond a Tiny Corpus
The examples above work great for a handful of sentences, but computing similarity against thousands of documents one-by-one gets slow fast. At that point, pairing Sentence Transformers with a proper vector index — FAISS for a lightweight local option, or Qdrant/pgvector for something with persistence — becomes the natural next step.
The good news: your embedding generation code barely changes. Sentence Transformers handles the "convert text to meaning" part regardless of what you use to store and search those vectors afterward. Swapping the storage layer later doesn't mean rewriting your embedding logic.
Common Mistakes Beginners Make
I've made a couple of these myself, so treat this as a shortcut past the confusion.
- Mismatching symmetric vs. asymmetric search patterns. Using a model tuned for one on the other silently degrades results without any obvious error.
- Skipping a relevance threshold. Returning the "best" match even when nothing is genuinely relevant produces confidently wrong results.
- Assuming a bigger model is always better. MiniLM is often plenty for prototyping — jumping straight to a heavier model wastes compute for no measurable benefit at small scale.
- Recomputing corpus embeddings every run. Cache them to disk once your corpus is stable — recalculating unchanged embeddings wastes time for zero benefit.
Frequently Asked Questions
What are Sentence Transformers?
Sentence Transformers (SBERT) is a Python library for computing dense vector embeddings from text. It produces embeddings where cosine similarity directly correlates with semantic meaning, making it ideal for search and comparison.
Is Sentence Transformers free to use?
Yes, Sentence Transformers is completely free and open-source. It runs entirely locally with no API calls, no per-token costs, and no account required. Over 25,000 pre-trained models are available on Hugging Face.
What is the difference between symmetric and asymmetric search?
Symmetric search pairs similar-length queries with corpus entries (finding similar questions). Asymmetric search pairs short queries with longer content (FAQ matching). Using the wrong model type silently degrades results.
Which Sentence Transformer model should I use?
Start with all-MiniLM-L6-v2 for fast prototyping. Use all-mpnet-base-v2 for higher quality general-purpose search. Check Hugging Face for domain-specific or multilingual models.
How does semantic search differ from keyword search?
Keyword search matches exact words. Semantic search understands meaning — "physician" matches "doctor" even with zero shared letters. Sentence Transformers capture conceptual similarity, not just text matching.
How do I scale semantic search beyond a small corpus?
Pair Sentence Transformers with a vector index like FAISS for local use, or Qdrant/pgvector for persistence. Generate embeddings once, store them, and use the vector index for fast similarity search.
Wrapping This Up
Sentence Transformers gives you a genuinely fast, free, local path to semantic search: load a model, encode your text, and compare using cosine similarity. No API keys, no per-token costs, and no cloud dependency required to get something working today.
Remember to match your model choice to whether your task is symmetric or asymmetric, set a relevance threshold so your system knows when to admit uncertainty, and cache your corpus embeddings once they're stable. FYI, this exact library quietly powers a huge chunk of the "semantic search" features you've probably already used without realizing it :)
Now go swap in your own FAQ list or document corpus instead of my four sample questions about doctor visits and passwords. That's genuinely where this stuff starts feeling useful instead of theoretical.