Running Multiple Local LLMs: Model Switching and Routing

September 29, 202614 min readSam Austin
Contents

Running multiple local LLMs with model switching and routing across one GPU
Running multiple local LLMs with model switching and routing across one GPU

Figure 1: One address in front, whichever model you actually need behind it

One model for coding, a smaller one for quick classification, maybe a vision model for images, all running on hardware that can't hold everything in VRAM at once. Ollama's "load one thing at a time and hope for the best" approach breaks down fast once you're juggling more than one workflow. Here's how to actually solve this.

I hit this wall running a coding assistant alongside a lightweight classifier for a side project. Manually starting and stopping llama-server processes got old within a day. Let's skip that pain for you.

Why This Gets Complicated Fast

The core constraint is simple: most consumer GPUs can't hold multiple large models in VRAM simultaneously. You need something that loads a model on demand, serves requests, and swaps it out for a different model when needed, all without your application code needing to know or care which model is currently resident.

Three real solutions exist for this today, and they solve genuinely different scopes of the problem:

ToolScopeBest For
llama.cpp router modeBuilt into llama-server itselfSimplest setup, no extra process
llama-swapExternal orchestrator in front of llama-serverFull control over per-model flags, crash isolation
LiteLLMMulti-provider proxy (local and cloud)Mixing local models with cloud APIs under one interface

If you haven't settled on hardware yet, our GPU guide for running local LLMs covers how much VRAM headroom each of these patterns actually needs.

Option 1: llama.cpp's Built-In Router Mode

For a long time, llama.cpp had a real limitation: one model per process, and switching meant a full restart. Recent builds fixed that directly with a router mode, bringing llama-server closer to what people expect from a modern local LLM runtime.

llama-server --router --models-dir /path/to/models

Once running, you manage models through simple HTTP calls:

curl -X POST http://localhost:8080/models/load -d '{"model": "qwen-coder"}'
curl -X POST http://localhost:8080/models/unload -d '{"model": "qwen-coder"}'

Important limitation to understand: even with multiple models defined, only one is resident in VRAM per worker at any given moment. Router mode supports multiple definitions, not simultaneous residency. Switching still means a full unload-and-reload cycle — it's just no longer a manual restart.

If every model you use sits under one llama.cpp router already, an extra proxy on top is often unnecessary. Router mode requires a reasonably recent llama-server build, so check your version supports it before assuming this path works. If you're still setting up llama.cpp from scratch, our llama.cpp tutorial covers the build first.

Option 2: llama-swap for Full Process Control

llama-swap is a lightweight Go binary that manages multiple local LLM server processes behind a single API endpoint. Unlike router mode, it's an external orchestrator sitting in front of one or more llama-server instances, giving you process-level isolation: a crash in one model's process doesn't take down the others.

Install it:

brew tap mostlygeek/llama-swap
brew install llama-swap

Configure with a single YAML file:

models:
  qwen-coder:
    cmd: llama-server -m /models/qwen-coder.gguf --port 9001
    ttl: 300
  smollm2:
    cmd: llama-server -m /models/smollm2.gguf --port 9002
    ttl: 300

Then run it:

./llama-swap --config config.yaml

llama-swap starts on port 8080 by default, and the model field in the request body determines which backend it activates. Client applications point to one stable URL and never need reconfiguring when you switch which model you're actually using. If VRAM can't fit both models simultaneously, llama-swap terminates the idle one before starting the newly requested one, handling the swap automatically.

The ttl value controls how long an idle model stays loaded before llama-swap terminates it to free memory. Tune this based on your usage pattern: short TTL if you're memory-constrained and switch often, longer TTL if reload latency bothers you more than memory pressure.

When llama-swap Is the Right Choice

llama-swap is most appropriate when you're already using llama-server directly and want to eliminate manual model-switching overhead while keeping full control over the flags used to run each model individually. It solves the specific gap between "Ollama is too simple for my needs" and "manually managing multiple llama-server instances has become genuinely disruptive."

It also ships an activity dashboard showing recent requests with model ID and performance metrics, plus a unified log view combining proxy and upstream logs — useful when you're debugging which model actually handled a given request.

One thing worth knowing: llama-swap doesn't download or manage model files. It only manages the processes that serve them. You're still responsible for having the GGUF files in place — our GGUF format guide covers what to pick.

Option 3: LiteLLM for Mixed Local and Cloud Routing

If your setup blends local models with cloud APIs, a different category of tool fits better. LiteLLM is an open-source proxy and client library that standardizes calls to over 100 LLMs using the OpenAI API format, and it works with Ollama as one of its supported backends alongside Anthropic, OpenAI, Bedrock, and others.

model_list:
  - model_name: local-coder
    litellm_params:
      model: ollama/qwen2.5-coder
      api_base: http://localhost:11434
  - model_name: cloud-fallback
    litellm_params:
      model: openai/gpt-4o
routing_strategy: latency-based-routing

LiteLLM provides several routing strategies out of the box, including least-busy routing, latency-based routing, and simple round-robin. This is the tool to reach for when you want a local model to handle most traffic but fall back to a cloud model for requests that need more capability than your local hardware can provide, all through one consistent API your application code talks to.

The mental model that helps here: the library is a client, the proxy is a platform. Running LiteLLM as a standalone proxy container turns it into a shared gateway your whole application stack talks to, rather than just a convenience import in one script.

Choosing Between These Three

Ask yourself what you're actually optimizing for:

  • Just need basic swap-on-demand with minimal setup? Start with llama.cpp's built-in router mode
  • Need per-model flag control, crash isolation, or a dashboard? Use llama-swap
  • Mixing local models with cloud fallback, or need sophisticated routing strategies? Use LiteLLM

IMO, most solo developers running a couple of local models should start with llama-swap. It hits the sweet spot between "too simple" (Ollama) and "too much infrastructure" (a full LiteLLM gateway deployment with Postgres and Redis behind it).

A Practical Decision Framework

Let me save you some research time with straightforward guidance:

  1. One or two models, casual use? Ollama alone is genuinely enough — don't add a proxy you won't benefit from
  2. Two or three models, same machine, one app? llama-swap with a tuned ttl, pointed at by a single stable endpoint
  3. Already living in llama-server? Router mode first; add a proxy only when you hit a wall it can't cover
  4. Any cloud model in the mix at all? LiteLLM, because the local-only tools don't cover that case
  5. Serving several users? Measure first — our local LLM benchmarking guide shows how much a swap actually costs you in latency
  6. Response too slow for chat? Smaller model for the frequent case, big model only on escalation

The mistake isn't picking the wrong row — it's deploying a gateway with a database behind it to swap between two models on one desk.

A Note on Semantic/Complexity Routing

Beyond simple model-name-based switching, more advanced setups route requests based on prompt complexity: simple queries go to a cheap, fast small model, complex ones escalate to a larger model. This pattern shows up in enterprise routing gateways using an inference-router approach, essentially a lightweight classifier deciding which backend handles each request before it's forwarded.

For local setups, you can build a rough version of this yourself: run a fast classifier — even a simple heuristic on prompt length or keyword matching — in front of your llama-swap or router-mode endpoint, and route accordingly. It's more setup work, but it means your tiny model handles the 80% of requests that don't need your biggest model's capacity.

Common Mistakes People Make

Setting TTL by guesswork

Recall the llama-swap section directly — watch your logs for how often you actually switch models before guessing at a value.

Ignoring context growth

Recall the hardware reality — leave some buffer beyond your largest model's footprint, since context length growth during a long conversation increases memory use.

Never testing cold-start latency

Recall the swap mechanics — know how long a swap actually takes on your hardware before you're surprised by it mid-conversation.

Messy model file paths

Recall the router section — a consistent --models-dir structure saves confusion once you're juggling more than three or four models.

Adding a proxy before you need one

Recall the choice section — an extra layer in front of a setup that already works is latency and a new failure mode, nothing more.

  • Designing Data-Intensive Applications by Martin Kleppmann — the systems thinking behind every proxy in this article: what happens when a request has to cross a process boundary and a model isn't resident yet.
  • Site Reliability Engineering edited by Beyer, Jones, Petoff, and Murphy — the operational discipline that turns a clever routing setup into something that stays up.
  • Building Microservices by Sam Newman — proxy patterns, single stable endpoints, and failure isolation, described in the general case before you apply them to inference.

Want to Go Deeper?

If you want structured practice on ML systems and serving, Educative's ML courses include hands-on labs that pair well with this kind of infrastructure work. The unlimited plan is useful when you're working through several deployment targets in one stretch.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

Can I run multiple local LLMs at the same time?

Usually not in VRAM simultaneously on consumer hardware. The practical answer is on-demand residency: load the model that is being asked for, serve it, then unload it when the next request needs something different. That is what router mode, llama-swap, and LiteLLM all automate.

What is llama.cpp router mode?

Router mode is built into recent llama-server builds. It lets you define several models and switch between them with HTTP calls instead of restarting the process. Note that it supports multiple definitions rather than simultaneous residency: only one model sits in VRAM per worker at a time.

What does llama-swap do?

llama-swap is a small Go binary that sits in front of one or more llama-server processes behind a single API endpoint. It gives you process-level isolation, per-model flags, an idle timeout that frees memory, and a dashboard showing which model handled each request.

When should I use LiteLLM instead?

When you are blending local models with cloud APIs, or need routing strategies such as least-busy or latency-based dispatch. LiteLLM standardizes calls to over 100 models behind the OpenAI API format, with Ollama as one supported backend alongside Anthropic, OpenAI, and Bedrock.

How long does it take to switch models?

A swap is a full unload-and-reload cycle, so latency depends on model size and your disk and VRAM speed. Measure it once on your own hardware before you rely on it mid-conversation, because a large model cold start is easily noticeable.

Does Ollama handle model routing?

Ollama loads one model at a time and is deliberately simple. If that default behaviour already matches your workflow it is fine, but it does not give you per-model process control, idle timeouts, or a routing layer in front of multiple backends.

What is complexity-based routing?

A pattern where a lightweight classifier inspects each prompt and sends simple requests to a small fast model while escalating hard ones to a larger model. Even a heuristic on prompt length or keywords captures much of the benefit and keeps your biggest model idle most of the time.

Wrapping This Up

Running multiple local LLMs stopped requiring manual process juggling once tools purpose-built for this problem matured. llama.cpp's router mode covers basic swap-on-demand, llama-swap adds process isolation and per-model control, and LiteLLM handles the broader case of blending local and cloud models under one API.

Will you need all three eventually? Maybe, if your setup grows complex enough. But start with whichever matches today's actual problem, not the one you imagine having in six months. Set up llama-swap with two models tonight, point your app at one stable endpoint, and watch the swap happen automatically the next time you switch tasks.

When that endpoint starts serving something real, our local LLM tools comparison and MLOps for beginners guide cover the surrounding infrastructure.