Contents
Here's the pitch in one sentence: download the weights once, and generate as many images as you want, forever, with nobody watching, nobody rate-limiting you, and nobody's content filter deciding what you're allowed to create. No subscription, no per-image credits, no queue. That's genuinely the entire appeal of running Stable Diffusion locally instead of through a cloud service.
This is the visual-generation counterpart to the local LLM tools comparison earlier in this series — same underlying philosophy (own the weights, own the compute), applied to image generation instead of text. The landscape has shifted meaningfully in 2026 though, and picking the interface everyone recommended two years ago will actively hold you back from running the models setting the current quality bar.
By the end of this guide, you'll understand which interface actually fits your needs in 2026, have it installed, and know exactly which model to download first. IMO, watching your first locally-generated image finish rendering with zero network activity is a genuinely similar thrill to the local LLM moment from earlier in this series :)
The Interface Decision: ComfyUI vs. Automatic1111 vs. Forge
This is genuinely the first and most consequential choice, and the honest answer has shifted since most existing tutorials were written.
- Automatic1111 (A1111) was the original, most-extended Stable Diffusion web UI — the largest extension ecosystem, the most familiar tabbed interface, and the deepest back-catalog of community tutorials. Development has genuinely stalled, though, and it cannot run Flux, the model family currently setting the quality bar for local generation.
- ComfyUI is now the maintained default. It picks up new model architectures first, runs faster, and uses less VRAM on the same card compared to A1111 — genuinely the more future-proof choice if you're starting fresh.
- Forge exists as a middle ground for people who want A1111's familiar interface with better performance and more current architecture support underneath.
The practical guidance for 2026: if you're specifically working with SDXL, SD 1.5, Pony, or Illustrious-family models, A1111 remains genuinely excellent and its extension ecosystem is unmatched. For Flux, SD 3.5, or video generation, ComfyUI or Forge are the correct choice — A1111 simply can't run these newer architectures.
Your models, LoRAs, and embeddings carry over to any of these interfaces — the .safetensors files are interchangeable, and only the frontend changes. This means switching later isn't the disruptive migration it sounds like.
Python Version Matters More Than You'd Expect
Before installing anything, get this exactly right — it's genuinely the single most common cause of installation failures.
# Ubuntu/Debian
sudo add-apt-repository ppa:deadsnakes/ppa
sudo apt update
sudo apt install python3.10 python3.10-venv python3.10-dev
python3.10 -m venv sd-env
source sd-env/bin/activate
Python 3.11 and newer will cause errors — several dependencies, notably xformers, aren't compatible with anything past 3.10.6 specifically. Use pyenv or conda to pin the exact version if your system's default Python is newer, rather than fighting dependency resolution errors that trace back to this mismatch.
Installing ComfyUI (The Recommended Default)
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
python3.10 -m venv venv
source venv/bin/activate
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
ComfyUI ships no models of its own. You download checkpoints separately and place them into models/checkpoints — the interface is genuinely just the engine and node-based workflow editor, not a bundled model package.
Downloading Your First Models
cd ~/ComfyUI/models/checkpoints
wget https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0/resolve/main/sd_xl_base_1.0.safetensors
SDXL base is roughly 6.9GB and remains a genuinely solid baseline that still holds up well. For sharper, more current results, the all-in-one fp8 build of Flux.1 schnell is worth grabbing too — it packs the text encoders and VAE into a single file, meaning there are no additional component downloads to chase down separately.
Installing Automatic1111 (If SDXL/SD 1.5 Is Your Focus)
git clone https://github.com/AUTOMATIC1111/stable-diffusion-webui.git
cd stable-diffusion-webui
python launch.py
The launch script handles most of the remaining dependency setup automatically on first run. This remains the right choice specifically if your work centers on SDXL, SD 1.5, or similar established architectures and you want access to A1111's unmatched extension catalog — ControlNet, LoRA management, and inpainting tooling are all genuinely more mature here than in newer alternatives.
Managing VRAM: The Flags That Actually Matter
Not everyone has a flagship GPU, and both interfaces support meaningful VRAM-constrained operation.
--lowvram— the most aggressive memory-saving flag, genuinely usable on GPUs with as little as 4GB of VRAM, paired specifically with SD 1.5 models rather than SDXL.--medvram— a middle-ground flag for GPUs with moderate VRAM, trading some speed for meaningfully reduced memory pressure.- SDXL genuinely needs more VRAM than SD 1.5 — if you're on a 4GB card, stick with SD 1.5-family models rather than fighting SDXL into a memory budget it wasn't designed for.
python launch.py --medvram --xformers
That --xformers flag is worth adding regardless of your VRAM tier — it's a memory-efficient attention implementation that speeds up generation and reduces memory usage simultaneously, genuinely close to a free performance win once installed correctly.
AMD GPU Support: ROCm
If you're on AMD rather than NVIDIA hardware, the setup path diverges meaningfully — this mirrors the exact same ROCm caveat from the Ollama tutorial earlier in this series.
- ROCm support exists but is genuinely less mature than the NVIDIA/CUDA path, with more hardware-specific troubleshooting typically required.
- Confirm your specific AMD GPU model is actually on ROCm's supported hardware list before investing significant setup time — not every AMD card is equally well supported.
Your First Generation
Once either interface is running and a model is downloaded, generation follows the same basic pattern regardless of which UI you chose.
- txt2img: the standard starting point — type a text prompt, generate an image from nothing but that description.
- img2img: start from an existing image and transform it according to a prompt, useful for style transfer or iterative refinement of a composition you already like.
- Inpainting/outpainting: selectively regenerate part of an image (inpainting) or extend it beyond its original borders (outpainting) — genuinely powerful editing capability that has no clean cloud-service equivalent with the same level of control.
Expected Performance: Setting Realistic Numbers
Worth knowing concretely what to expect rather than guessing based on marketing claims. On an RTX 4090, generation times and VRAM usage are well-documented enough to set genuine expectations — SDXL generations typically complete in a handful of seconds per image, with Flux.1 schnell (built specifically for fast inference despite its higher quality ceiling) landing in a similar practical range.
If you're evaluating hardware for local image generation, the RTX 5070 and RTX 5080 both handle SDXL and Flux comfortably — the VRAM headroom matters more than raw TFLOPS for generation speed at typical resolutions.
Your actual numbers will vary meaningfully based on GPU tier, resolution, and step count — don't assume flagship-GPU benchmarks translate directly to a mid-range card without adjustment.
Extending Your Setup: ControlNet and LoRAs
Once basic generation works, two extensions genuinely transform what's possible.
- ControlNet lets you condition generation on structural input — a pose skeleton, an edge map, a depth map — giving you dramatically more compositional control than prompt text alone can provide.
- LoRAs (Low-Rank Adaptation models) are small, focused fine-tuning files that add a specific style, character, or concept to a base model without needing to retrain or redownload the entire checkpoint — genuinely the same "small adapter on a big frozen base" principle you've seen in other fine-tuning contexts.
Both are widely available on community model-sharing sites and work identically regardless of which interface you chose, since they're just additional .safetensors files loaded alongside your base checkpoint.
Want to Go Deeper?
If the ControlNet conditioning or LoRA fine-tuning concepts clicked and you want to dig into the theory behind diffusion models and generative architectures, Educative's Machine Learning path covers image generation algorithms in detail alongside the broader deep learning landscape — worth exploring if you're building custom fine-tuning or advanced workflows into a real production pipeline.
Common Mistakes People Make
- Installing on Python 3.11+ and fighting dependency errors that trace back to the version mismatch. Pin Python 3.10.6 exactly before troubleshooting anything else.
- Choosing Automatic1111 by habit when your actual goal is running Flux or SD 3.5. A1111's development has stalled specifically on newer architecture support — check what models you actually want to run before picking an interface.
- Attempting SDXL on 4GB VRAM without the right flags, or at all. SD 1.5 with
--lowvramis the realistic path on genuinely constrained hardware; SDXL wants more headroom. - Assuming switching interfaces later means losing your models. Checkpoints, LoRAs, and embeddings are interchangeable
.safetensorsfiles — only the frontend changes when you switch. - Skipping
--xformersor equivalent memory-efficient attention. This is close to a free performance and memory win on nearly every setup, regardless of GPU tier.
Where This Fits With the Rest of This Series
This article extends the edge AI and local deployment arc into visual generation — the same philosophy driving local LLM coverage applies directly to image generation weights. And if you're interested in the model compression techniques that make these weights possible, the quantization and GGUF format articles cover the mechanisms behind smaller, faster models — relevant to Stable Diffusion's own fp16 and fp8 checkpoint variants.
Wrapping This Up
Getting Stable Diffusion running locally in 2026 means picking ComfyUI as your default unless you specifically need A1111's mature extension ecosystem for SDXL or SD 1.5 work — the interface landscape has genuinely shifted since older tutorials were written, and Flux-family models specifically require the newer tooling. Pin Python 3.10.6, download SDXL base or Flux.1 schnell as your first checkpoint, and use --lowvram/--medvram flags to match your actual hardware.
Remember that your model files carry over freely between interfaces, so the initial choice isn't as permanent as it might feel, and that Python version mismatches remain the single most common installation failure across every guide covering this setup. FYI, this genuinely mirrors the local LLM tooling landscape from earlier in this series — different interfaces, different underlying models, but the exact same "own your weights, own your compute" philosophy driving why this ecosystem exists at all :)
Now go generate your first image with a genuinely simple prompt, then immediately try the same prompt through img2img on the result. That quick before-and-after comparison is honestly the fastest way to feel how much control local generation actually hands you compared to a locked-down cloud interface.