Sam Austin AI

Chroma DB Tutorial: Local Vector Store for LLM Apps

September 1, 2026 10 min read Updated September 2, 2026 Sam Austin
Contents

Not every project needs a fully managed cloud vector database with a monthly bill attached. Sometimes you just want to store some embeddings locally, query them, and move on with your life. That's exactly the gap ChromaDB fills, and it does it without demanding a credit card or a weekend of infrastructure setup.

I reach for Chroma constantly when prototyping — it's become my default "let me just test this idea real quick" tool before deciding whether a project deserves something heavier. That's not a knock against it; it's genuinely the point of the thing.

By the end of this tutorial, you'll have a working local vector store, some documents loaded in, and a semantic search query returning real results. IMO, this is the fastest path from "I have an idea" to "I have working code" in the entire vector database space :)

What ChromaDB Actually Is

ChromaDB is an open-source vector database built specifically for AI applications. It stores embeddings locally, runs without a separate server process, and drops into any Python project in minutes.

The defining feature here isn't raw performance — it's friction removal. No Docker required for basic use, no account signup, no API key. You install a package and you're already querying vectors.

  • Runs fully offline for local development.
  • Uses SQLite for local persistent storage, so your data survives between runs.
  • Supports a client-server mode for when you outgrow purely local use.
  • Ships SDKs for Python, JavaScript, and more, with Python being the most commonly used.

Ever wondered why so many tutorials default to Chroma for teaching RAG concepts? This is exactly why — nobody wants to explain cloud account setup before they've even shown you what an embedding does.

ChromaDB Local Vector Database Tutorial
ChromaDB Local Vector Database Tutorial

Figure 1: ChromaDB provides a lightweight local vector store for AI applications

Installing ChromaDB

Let's get this running. I'd strongly recommend a virtual environment here, since Chroma pulls in a fair number of dependencies underneath the hood.

python -m venv venv
source venv/bin/activate  # on Windows: venv\Scripts\activate
pip install chromadb

Skipping the virtual environment isn't a fatal mistake, but installing directly into your global Python environment risks dependency conflicts with other projects down the line. I've had this bite me once, and cleaning up a broken global environment is far more annoying than typing three extra setup commands.

Your First Collection

In Chroma, a collection is the core organizing unit — think of it like a table in a relational database, except it holds vectors, documents, and metadata together.

import chromadb

client = chromadb.Client()

collection = client.create_collection(name="my_first_collection")

That chromadb.Client() call creates an in-memory client — perfect for quick experiments, but everything disappears the moment your script ends. We'll fix that persistence problem in a minute.

Adding Documents

Now let's actually put something in the collection. Chroma handles the embedding generation automatically using a default model, so you don't need to call a separate embedding API yourself.

collection.add(
    documents=[
        "ChromaDB is an open-source vector database for AI applications.",
        "Python is a high-level programming language created by Guido van Rossum.",
        "The Eiffel Tower was completed in 1889 in Paris, France.",
    ],
    ids=["doc1", "doc2", "doc3"]
)

Notice you're just passing plain text strings — no manual embedding step required. Chroma converts these to vectors behind the scenes using its default embedding function. That's the whole point: less plumbing, faster iteration.

Querying Your Collection

Here's where the "vector database" part actually earns its keep. Let's search using a phrase that doesn't share exact words with our stored documents.

results = collection.query(
    query_texts=["What programming language did Guido create?"],
    n_results=2
)

print(results["documents"])

Even though the query doesn't contain "Python" verbatim, the semantic match still surfaces that document near the top. That's the difference between keyword search and vector search in one tiny example — meaning matters more than exact phrasing.

Understanding the Results

The query response includes a few useful pieces beyond just the matched text:

  • documents — the actual text content that matched.
  • distances — how close each match is conceptually; lower usually means more relevant, depending on the distance metric.
  • metadatas — any extra metadata you attached when adding documents.
  • ids — the identifiers you assigned when inserting the data.

I'll be honest, the first time I saw distance scores I assumed higher meant "more relevant" — that assumption cost me a confused afternoon. Double-check your distance metric's direction before trusting your intuition on this.

Making Data Actually Persist

An in-memory client is fine for a five-minute experiment, but real projects need data that survives a restart. This is where PersistentClient comes in.

import chromadb

client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_or_create_collection(name="my_docs")

collection.add(
    documents=["This data will survive a restart."],
    ids=["persistent_doc1"],
)

print(f"Collection has {collection.count()} documents")

That path argument tells Chroma where to write its SQLite-backed storage on disk. Restart your script, run the same code again, and your data is still there — no re-embedding, no reloading from scratch.

Adding Metadata for Filtering

Raw semantic search is powerful, but real applications almost always need to narrow results by category, date, or source. Metadata makes that possible.

collection.add(
    documents=[
        "Refund requests must be submitted within 30 days.",
        "New employees receive 15 days PTO in year one.",
    ],
    metadatas=[
        {"category": "policy"},
        {"category": "hr"},
    ],
    ids=["policy1", "hr1"]
)

results = collection.query(
    query_texts=["how many vacation days do I get"],
    n_results=2,
    where={"category": "hr"}
)

That where clause scopes your search to just the HR category before it even considers similarity. This combination of semantic search plus hard filtering is exactly the pattern production RAG systems rely on — you're just seeing it at a much smaller, friendlier scale here.

Running Chroma in Client-Server Mode

Local, in-process usage is great for prototyping, but sometimes you need multiple processes or applications hitting the same Chroma instance. That's what server mode solves.

chroma run --path ./chroma_db

Then connect from your Python code using an HTTP client instead of a local one:

import chromadb

client = chromadb.HttpClient(host="localhost", port=8000)
collection = client.get_or_create_collection("my_docs")

This is basically the same API surface, just talking over HTTP instead of running in-process. Switching between local and server mode barely changes your application code — which honestly deserves more credit than it usually gets.

Common Mistakes Beginners Make

I've made a few of these myself while getting comfortable with Chroma, so treat this as a shortcut past my own trial and error.

  • Using the in-memory client and expecting data to persist. It won't — switch to PersistentClient the moment you need data to survive between runs.
  • Installing directly into a global Python environment. Chroma pulls in a fair number of dependencies; isolate your projects with a virtual environment to avoid conflicts.
  • Forgetting unique IDs when adding documents. Duplicate IDs will overwrite existing entries silently, which is a confusing bug to track down later.
  • Assuming Chroma scales to production workloads at massive size. It's genuinely built for development speed, not operating at hundreds of millions of vectors — that's a different tool's job.

When Should You Move Past Chroma?

Chroma is fantastic for prototyping, personal projects, and small-to-medium production workloads. But it's honestly not designed for production at 50 million or 100 million vectors — that's not a criticism, just an accurate description of what it optimizes for.

Once you outgrow it, teams typically migrate to something like Qdrant, Pinecone, or Milvus for the scale and operational tooling those systems provide. My honest opinion? That migration path is a feature, not a flaw — Chroma lets you validate an idea fast without committing to heavier infrastructure before you know if the project even works.

Frequently Asked Questions

What is ChromaDB used for?

ChromaDB is an open-source vector database used for storing and querying embeddings in AI applications. It's ideal for prototyping RAG pipelines, semantic search, and LLM-powered apps without needing cloud infrastructure.

Is ChromaDB free to use?

Yes, ChromaDB is completely open-source and free. You can run it locally without any API keys, accounts, or subscription fees.

How is ChromaDB different from Pinecone?

ChromaDB runs locally and is free, making it ideal for prototyping. Pinecone is a managed cloud service better suited for production at scale. Chroma is simpler but Pinecone handles infrastructure for you.

Does ChromaDB persist data between restarts?

ChromaDB has two modes: in-memory (data lost on restart) and PersistentClient (data saved to disk). Use PersistentClient for data that needs to survive restarts.

Can ChromaDB be used in production?

ChromaDB works for small to medium production workloads. For massive scale (50M+ vectors), consider Qdrant, Pinecone, or Milvus instead.

What embedding model does ChromaDB use?

ChromaDB uses a default embedding model (all-MiniLM-L6-v2) automatically. You can also configure custom embedding functions for different models.

Wrapping This Up

ChromaDB gives you a genuinely fast path from idea to working vector search: install the package, create a collection, add documents, and query by meaning instead of exact keywords. No servers, no API keys, no cloud account required to get started.

Remember to switch to PersistentClient the moment you need data to survive a restart, and don't be afraid to outgrow Chroma once your project actually demands production-scale infrastructure. FYI, plenty of successful RAG prototypes start exactly here before graduating to something heavier — there's no shame in starting small :)

Now go load in your own documents instead of my three sentences about towers and Python's creator. That's genuinely where this stuff starts feeling useful.

Share this article X Facebook LinkedIn Reddit WhatsApp