Sam Austin AI

Supabase Vector Tutorial: pgvector Made Easy

September 23, 2026 12 min read Sam Austin
Contents

So you want to build a chatbot that actually understands your data, huh? Maybe a semantic search engine that doesn't return garbage results when someone types cheap laptop instead of affordable notebook computer? Good news, friend: Supabase and pgvector solve exactly this problem, and today I'll walk you through it like we're grabbing coffee together.

I've spent a silly amount of hours wrestling with vector databases over the past couple of years. I tried Pinecone, dabbled with Weaviate, and once spent an entire weekend debugging a Chroma setup that refused to persist anything. Then I tried pgvector on Supabase, and honestly? I felt like I'd been overcomplicating my life. :/ Let me save you that weekend.

What Even Is a Vector Database?

Picture this: you take a sentence, run it through an embedding model, and get back a list of numbers. Something like [0.12, -0.45, 0.88, ...]. That list of numbers captures the meaning of your sentence.

Embeddings let machines compare concepts instead of just matching keywords. The sentences I adore pizza and I love pizza produce vectors that sit right next to each other in mathematical space, even though they share only one word.

Regular keyword search works fine for exact matches. But have you ever searched for something and gotten results that technically contain your words but completely miss your intent? Yeah, me too. That's the classic "I searched for apple recipes and got iPhone cases" problem.

A vector database stores those number lists and finds the nearest neighbors to your query. It answers what's semantically close? instead of what contains this exact string?

Why Supabase + pgvector? (Spoiler: It's Just Postgres)

Here's the part where I get a little evangelical. pgvector is a Postgres extension, and Supabase supports it out of the box. That means:

  • One database for everything — your app data AND your vectors live together
  • No extra service to pay for — you skip yet another SaaS subscription with a usage dashboard designed to induce anxiety
  • SQL you already know — no proprietary query language to learn
  • Supabase's generous free tier — perfect for prototypes and hobby projects

IMO, the "just use Postgres" approach wins for most small-to-medium projects. Why juggle two databases when one does the job?

My Honest Comparison

Pinecone is great, don't get me wrong. It's fast and it scales beautifully. But for a solo dev or a small team, adding a managed vector DB feels like renting a forklift to move a couch. With pgvector, you already have the couch dolly — it's called SQL. And compared to the best vector databases compared in this series' earlier roundup, pgvector's real edge is zero extra infrastructure, not a feature checklist win.

Supabase pgvector Vector Database Semantic Search Tutorial

Figure 1: Supabase pgvector — semantic search without leaving Postgres

Image Alt Text: "Supabase vector tutorial with pgvector HNSW index and SQL semantic search"

Setting Up Your Supabase Project

Time to build. Head over to supabase.com, create a free project, and give it a name. Once your project spins up, open the SQL Editor in the dashboard.

Step 1: Enable the Extension

CREATE EXTENSION IF NOT EXISTS vector;

That's it. One line. Compare that to spinning up an entire vector database cluster, and you'll see why I'm so fond of this approach. The vector extension comes pre-installed on Supabase, so this command just flips the switch.

Step 2: Create a Table

Let's build something practical: a table for documents we want to search.

CREATE TABLE documents (
  id BIGSERIAL PRIMARY KEY,
  content TEXT,
  embedding VECTOR(1536)
);

Important: that 1536 isn't random. It matches the output dimension of OpenAI's text-embedding-3-small model. If you use a different embedding model, change this number to match. Mismatched dimensions cause errors that will make you question your career choices.

Step 3: Add a Vector Index

Without an index, Postgres scans every row on every query. That works fine for testing with ten rows, but it becomes a horror show at scale. Add an HNSW index:

CREATE INDEX ON documents
USING hnsw (embedding vector_cosine_ops);

HNSW gives you fast approximate nearest-neighbor search. It trades a tiny bit of accuracy for a massive speed boost — a trade I'd take every single time, same HNSW-versus-exact-search tension covered in the Elasticsearch vector search tutorial.

Generating Embeddings

Vectors don't appear by magic — you create them with an embedding model. Here's a Node.js example using OpenAI's API:

import OpenAI from "openai";

const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

async function getEmbedding(text) {
  const response = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: text,
  });
  return response.data[0].embedding;
}

A Quick Tip From Hard-Earned Mistakes

Always use the same embedding model for storing and querying. I once mixed two models in the same table, and my search results looked like they came from a random number generator. Consistency is everything here — same principle the embedding models compared article made about picking one model and sticking with it. Also, FYI, embedding models cost money per token — so chunk long documents instead of embedding entire novels in one go.

Inserting Data into Supabase

Now let's put everything together using the Supabase client:

import { createClient } from "@supabase/supabase-js";

const supabase = createClient(
  process.env.SUPABASE_URL,
  process.env.SUPABASE_ANON_KEY
);

async function addDocument(text) {
  const embedding = await getEmbedding(text);
  const { error } = await supabase.from("documents").insert({
    content: text,
    embedding: embedding,
  });
  if (error) console.error("Insert failed:", error);
}

Not glamorous, but it works. And honestly, boring and reliable is exactly what you want in a data pipeline.

Here's where the magic happens. Supabase exposes a match_documents pattern through Postgres functions. Create this function in the SQL Editor:

CREATE FUNCTION match_documents(query_embedding VECTOR(1536))
RETURNS TABLE (id BIGINT, content TEXT, similarity FLOAT)
LANGUAGE sql AS $$
  SELECT
    documents.id,
    documents.content,
    1 - (documents.embedding <=> query_embedding) AS similarity
  FROM documents
  ORDER BY documents.embedding <=> query_embedding
  LIMIT 5;
$$;

That <=> operator computes cosine distance between vectors. Subtract it from 1, and you get a similarity score where 1.0 means basically the same thing.

Then call it from JavaScript:

async function search(query) {
  const embedding = await getEmbedding(query);
  const { data, error } = await supabase.rpc("match_documents", {
    query_embedding: embedding,
  });
  return data;
}

Run search("how do I reset my password") against a pile of FAQ entries, and watch it pull up your password-reset doc — even if that doc never contains the word "reset." Satisfying, right? :)

Distance Functions: Pick the Right One

pgvector offers three distance operators, and choosing correctly matters:

Operator Distance When to use
<-> L2 Vector magnitude carries meaning (some image embeddings)
<=> Cosine Default for text — compares direction, not length
<#> Inner product Certain normalized or specialized models

For text embeddings, cosine distance wins nine times out of ten. I default to it unless something breaks, and I recommend you do the same.

Scaling Tips (Because You'll Need Them)

Once your table grows past a few thousand rows, keep these tips handy:

  • Chunk your documents into 200–500 token pieces for better retrieval quality — same sizing guidance as the building a RAG pipeline tutorial
  • Store metadata alongside vectors — source URLs, titles, timestamps — so you can filter results with plain SQL WHERE clauses
  • Use the match_threshold pattern to filter out low-similarity junk results
  • Monitor index performance with EXPLAIN ANALYZE before assuming you need to pay for more compute

That last one deserves emphasis: pgvector gives you all of Postgres alongside your vectors. You can join embeddings against your users table, filter by tenant, and enforce row-level security — all in one place. Try doing that with a standalone vector DB without writing glue code until sunrise.

Where This Really Shines

Where does this combo actually deliver? A few favorites:

  • RAG chatbots that answer questions from your documentation — the classic use case, pair it with the RAG beginner's guide from earlier in this series
  • Semantic product search for e-commerce stores
  • Recommendation engines based on content similarity
  • Deduplication — finding near-identical text across a database

If RAG is your goal, pair this setup with an LLM call: retrieve the top matches, stuff them into the prompt, and let the model write the answer. The whole pipeline fits in well under a hundred lines of code.

  • Designing Data-Intensive Applications by Martin Kleppmann — the deeper Postgres and storage-engine context that makes pgvector's "just another index" story click.
  • AI-Powered Search by Trey Grainger et al. — semantic retrieval patterns that map directly onto match_documents-style SQL functions.
  • Retrieval-Augmented Generation titles on RAG architecture — pairing pgvector retrieval with LLM prompting is the exact pipeline this tutorial sets you up to build.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is pgvector in Supabase?

pgvector is a PostgreSQL extension that adds a VECTOR column type plus distance operators and indexing (HNSW, IVFFlat) for similarity search. Supabase supports it out of the box, so your app data and embeddings live in one database you already query with SQL.

How do I enable pgvector on Supabase?

Run CREATE EXTENSION IF NOT EXISTS vector; in the Supabase SQL Editor. The extension is pre-installed on Supabase, so the command just activates it — no cluster to spin up or separate service to configure.

What dimension should my VECTOR column be?

Match the output dimension of your embedding model. VECTOR(1536) matches OpenAI text-embedding-3-small; all-MiniLM-L6-v2 produces 384 dimensions. Mismatched dimensions cause insert and query errors, so keep storage and query models consistent.

When should I use HNSW vs no index with pgvector?

Without an index, Postgres scans every row — fine for testing with a handful of rows, slow at scale. HNSW gives fast approximate nearest-neighbor search, trading a small amount of recall for a large speed win. Monitor with EXPLAIN ANALYZE before paying for more compute.

Which distance operator should I use for text embeddings?

Cosine distance (<=>) is the default choice for text — it compares direction rather than length. L2 (<->) fits image embeddings where magnitude carries meaning, and inner product (<#>) suits certain normalized or specialized models.

Can pgvector do hybrid search or filtering?

You can filter with plain SQL WHERE clauses on metadata columns alongside vector distance ordering, and join embeddings against other tables with row-level security. Dedicated vector databases often offer native keyword+vector fusion; with pgvector you compose hybrid behavior in SQL or application code.

Wrapping Up

Let's recap the journey: you enabled pgvector with one line of SQL, created a table with a VECTOR column, generated embeddings, added an HNSW index, and built semantic search with a single Postgres function. No new services, no new languages, no weekend-long debugging sessions. Just Postgres doing what Postgres does — with a superpower bolted on.

My honest take? If you're building anything with AI features and you're not already committed to a dedicated vector platform, start with Supabase and pgvector. You can always migrate later if you outgrow it, and migrating one table beats untangling a distributed architecture.

So go spin up that free Supabase project, paste in the SQL from this tutorial, and embed something. Your future self — the one with a working semantic search feature — will thank you. And hey, if you blow up the vector dimensions like I did, no judgment. We've all been there. ;)

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles