Contents
Figure 1: Running large language models locally on your own hardware — zero cloud dependency, complete privacy
Here's the entire pitch for Ollama in one sentence: one command to install, one command to chat, and your conversation never leaves your machine. No API keys, no per-token billing anxiety, no "the model changed overnight" surprises. As of late August 2026, the current stable release sits at v0.33.1, and the install process genuinely hasn't gotten any more complicated than that one-liner promises.
I'm writing this as the practical follow-up to the local LLM tools comparison from earlier — that article told you why Ollama is the default starting point for developers. This one actually gets you from zero to a working local model, plus the configuration details that separate "it technically works" from "it's actually fast on your specific hardware."
By the end of this guide, you'll have Ollama installed, a model running, and it wired into either your terminal, a Python script, or your IDE. IMO, watching your first local model respond with zero network activity happening is a genuinely satisfying moment the first time :)
What Ollama Actually Does
Ollama wraps llama.cpp inference behind a clean command-line interface and an OpenAI-compatible REST API, abstracting away model quantization, GPU memory allocation, and file management that you'd otherwise handle manually.
You install it once, pull a model, and immediately get a chat interface in your terminal or an API endpoint any library can call.
The API runs locally at http://localhost:11434 — any tool expecting an OpenAI-style endpoint can usually point at this instead with minimal changes.
The tool itself is completely free. An optional Pro plan exists specifically for routing requests to datacenter hardware when a model exceeds your local RAM or VRAM — but nothing about basic local use requires paying anything.
Installing Ollama
macOS and Linux
curl -fsSL https://ollama.com/install.sh | sh
This single line installs Ollama and, on distributions using systemd, sets up a background service automatically — meaning Ollama starts on boot and stays running without you needing to launch it manually each session.
Windows
irm https://ollama.com/install.ps1 | iex
The Windows build installs as a background service with a system tray icon. If you have an Nvidia GPU with a current driver, Ollama detects and uses it automatically at the next model run — recent installer updates specifically improved GPU auto-detection tuned for RTX 40-series cards.
If you're planning to run larger models, a mid-range GPU like the RTX 4070 with 12GB VRAM handles 7B-13B models comfortably. For headroom with bigger models, the RTX 4070 Ti Super with 16GB VRAM is a solid investment.
Docker (For Containerized Setups)
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
Drop the --gpus=all flag if you don't have Nvidia's container toolkit installed. Ollama falls back to CPU inference without complaint — just noticeably slower.
Verifying Your Install
ollama --version
If ollama isn't recognized right after installation, open a new terminal session — PATH changes sometimes need a fresh shell to take effect. Ollama ships new point releases every few days, so don't worry if the version number has already moved past whatever's current when you read this.
Checking Your GPU Situation
Before pulling any models, it's worth confirming what hardware Ollama will actually use.
nvidia-smi
This should display your driver version — 525+ is required, 550+ recommended — and list your available GPUs. If the command isn't found or throws errors, update your driver before expecting GPU-accelerated inference to work. If no supported GPU is found, or a model doesn't fit in available VRAM, Ollama silently falls back to CPU. Output is still correct either way — just significantly slower.
Pulling and Running Your First Model
ollama run llama3.2
This pulls the default 3B-parameter tag (roughly 2.0GB) if it's not already local, then drops you into an interactive chat. Type /bye to exit whenever you're done.
Choosing a Model for Your Hardware
Not every model fits every machine, and picking the right size upfront saves you a frustrating first impression.
ollama run llama3.2:1b # ~1.3GB, fast on limited hardware
ollama run gemma3:4b # ~3.3GB, multimodal
ollama run deepseek-r1:7b # ~4.7GB, reasoning-focused model
On a machine with 8GB of RAM, start with the compact Llama 3.2 tag rather than anything larger. Start small and only move up to a bigger model once you've confirmed you actually have spare RAM or VRAM to spend — a model that barely fits will run, but painfully slowly as your system fights for memory.
One-Shot Prompts Without the Chat Session
If you just need a quick answer rather than an ongoing conversation:
ollama run llama3.2 "Explain the CAP theorem in two sentences"
This is genuinely useful for scripting — piping a single prompt through Ollama from a shell script or cron job without needing to manage an interactive session at all.
Managing Models: The Package-Manager Mental Model
Ollama's model commands behave a lot like a package manager, and that mental model genuinely helps once you're juggling several models.
ollama list # see what's downloaded locally
ollama ps # see what's currently loaded in memory
ollama pull qwen3:8b # download without running immediately
ollama rm qwen3:8b # delete a model you no longer need
ollama ps is worth checking regularly if you're running multiple models across different projects — it tells you exactly what's currently occupying memory, which matters once you start layering local LLM work with anything else memory-intensive on the same machine.
The REST API: Wiring Ollama Into Your Own Code
Once Ollama is running, it exposes an OpenAI-compatible API at localhost:11434 that any HTTP client or compatible library can hit directly.
import requests
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": "llama3.2",
"prompt": "Summarize the plot of a heist movie in one paragraph.",
"stream": False
}
)
print(response.json()["response"])
Because the API is OpenAI-compatible, plenty of existing tooling built for OpenAI's API works against Ollama with minimal changes — often just swapping the base URL and dropping the API key requirement entirely.
IDE Integration: Local Code Completion
If you're trying to replace a cloud coding assistant with something fully private, Ollama paired with an IDE extension is genuinely the most common path people take.
Pull a code-focused model: ollama pull codellama:code (or a comparable current coding model from Ollama's library).
Extensions like Continue.dev support VS Code and JetBrains, using your local Ollama server for code completions, chat, and inline edits.
Recent Ollama releases added one-command IDE integration — launching supported apps directly against a local model without manually configuring endpoints yourself.
Everything here runs offline once the model's downloaded — your code genuinely never leaves your machine, which is the entire point for anyone choosing this path over a cloud coding assistant.
Configuration: Getting the Most Out of Your Hardware
A few environment variables and settings genuinely matter once you're past the "does it work at all" stage.
OLLAMA_LOAD_TIMEOUT — worth adjusting if you're running larger models or image generation models that take longer than the default timeout to load into memory.
GPU selection — if you have multiple GPUs, Ollama's documentation covers environment variables for controlling which GPU(s) handle inference, rather than defaulting to automatic selection.
WSL2 users on AMD hardware: ROCm support under WSL2 is experimental, hardware-dependent, and not officially supported — check AMD's WSL2 ROCm support matrix for your specific GPU before assuming this path will work smoothly.
What's Genuinely New in 2026 Worth Knowing About
A few developments since Ollama's earlier days are worth flagging specifically, since older tutorials won't mention them.
Apple Silicon inference now runs on MLX rather than llama.cpp for Ollama specifically — a meaningful speed improvement since MLX exploits Apple's unified memory architecture more directly.
Speculative decoding (MTP) ships for supported models, delivering roughly 2x speed gains — worth checking whether your chosen model supports this before assuming you're at peak achievable speed.
Cloud models exist within Ollama too — if privacy is your actual reason for going local, double-check you're selecting a genuinely local model tag rather than one of Ollama's cloud-routed options, since the two coexist in the same interface.
Common Mistakes People Make
Pulling a model too large for available RAM/VRAM on a first attempt. Start with a compact tag like llama3.2:1b, confirm things work, and scale up only once you know you have headroom.
Not checking nvidia-smi before assuming GPU acceleration is active. Ollama fails silently into CPU mode rather than erroring — if responses feel unexpectedly slow, this is the first thing to check.
Forgetting a new terminal session is needed after install for PATH changes to register — a genuinely common "command not found" moment right after installation.
Assuming every model tag is local-only. With cloud models now coexisting inside Ollama's interface, verify you're actually running a local tag if privacy or offline operation is the reason you're here in the first place.
Skipping ollama ps when juggling multiple models — it's the fastest way to see what's actually loaded and consuming memory right now, versus what's just downloaded and sitting idle on disk.
Next Steps: Deepen Your AI Knowledge
Once you've got Ollama running locally, you might want to deepen your understanding of LLMs, fine-tuning, and building production AI applications. Educative offers hands-on, interactive courses on LLMs, RAG pipelines, and AI engineering — learn by doing rather than just watching videos, with sandboxed environments that let you practice without setting up complex local infrastructure.
Wrapping This Up
Getting Ollama running is genuinely as simple as the marketing promises: one install command, one ollama run command, and you've got a private, offline-capable LLM responding on your own hardware. The REST API and IDE integrations are what turn that from a fun terminal toy into something you'd actually build workflows around.
Remember to match model size to your actual available RAM or VRAM before assuming something's broken, and confirm your GPU is actually being detected if performance feels off. FYI, Ollama ships updates every few days, so it's worth running ollama --version occasionally and skimming release notes — the tool's genuinely still evolving quickly even though the core workflow has stayed remarkably stable :)
Now go pull a small model, disconnect your wifi, and confirm it still answers you. That's genuinely the moment local LLM tooling stops being an abstract idea and starts feeling like something real.