Blog
By Published 10 min read

How to Run an LLM Locally in 2026: Tools, Models, Hardware

A step-by-step guide to running an LLM locally in 2026: Ollama, LM Studio, llama.cpp and vLLM, which open models to pick, and the hardware you really need.

Legacies is a software and web studio from Romania, founded by Horia Stan and Alexandru Talnaci, and we run open models on our own machines every week for testing, private drafts and prototypes. This guide is the setup we would give a friend who wants to do the same.

To run an LLM locally in 2026, install a runner like Ollama or LM Studio, download an open-weight model that fits in your GPU memory or Mac unified memory, and chat with it or call it through a local API. A 16GB machine runs good small models. A 24-32GB GPU or a 64GB+ Mac runs models that feel close to last year's cloud assistants.

Quick answer

Why run a model locally at all

There are three honest reasons. Privacy, cost control and offline work.

Privacy is the big one. When the model runs on your machine, your prompts and documents never leave it. No vendor logs, no retention policy. For contracts, client code or anything under NDA, that matters.

Cost is the second. After you own the hardware, every token is free except electricity. You stop rationing. For the full math, we did it in our breakdown of local AI vs cloud API cost.

Offline is the third. A local model works on a plane or inside a network that blocks external AI services.

What you give up is raw capability. More on that below.

The four tools that matter in 2026

Pick one tool based on how you like to work.

ToolBest forInterfaceLicenseLocal API
OllamaDevelopers, servers, scriptingCLI plus desktop appMIT, open sourcePort 11434, OpenAI and Anthropic compatible
LM StudioExploring models, non-developersFull desktop app plus lms CLIProprietary, free for home and workPort 1234, OpenAI and Anthropic compatible
llama.cppMaximum control, odd hardwareCLI and llama-serverMIT, open sourceOpenAI compatible server
vLLMServing many users at once on NVIDIAPython serverApache 2.0OpenAI compatible server

Ollama is the default for most developers. It is MIT licensed, runs on macOS, Windows, Linux and Docker, and the latest release at the time of writing is v0.35.1. Local use is unlimited on every plan, according to Ollama's pricing page. The paid plans only cover their cloud models.

LM Studio is the friendliest. You browse models, see which ones fit your hardware, and chat in a clean interface. It runs GGUF and MLX models, supports MCP servers and document chat, and has been free for work use since July 2025. The app itself is closed source.

llama.cpp is the engine underneath much of this ecosystem, with CUDA, Metal, Vulkan, AMD HIP and CPU backends. Use it directly when you want every knob.

vLLM is for production serving, when many people hit the same GPU at once. Overkill for one person on a laptop.

If you are stuck between the first two, read our Ollama vs LM Studio comparison.

Which open models are worth running now

These are the families we would start with in October 2026, with sizes taken from the Ollama library and the official model cards.

ModelSizesDownload at default quantContextLicense
Gemma 4 E4B4.5B effectiveAbout 7-10GB128KApache 2.0
Gemma 4 12B12BAbout 8GB256KApache 2.0
Gemma 4 26B25.2B total, about 4B active (MoE)About 16-19GB256KApache 2.0
Gemma 4 31B30.7B denseAbout 19-20GB256KApache 2.0
Qwen3.6 27B27B denseAbout 18-19GB262K nativeApache 2.0
gpt-oss-20b20B14GB128KApache 2.0
gpt-oss-120b120B65GB128KApache 2.0

The Gemma 4 model card lists five sizes from E2B to 31B. The Qwen3.6-27B card describes a dense 27B model with a native 262,144 token context. gpt-oss uses MXFP4 weights, so the 20B fits in about 16GB of memory and the 120B fits on a single 80GB GPU.

Our short picks:

Bigger open models like GLM-5.3 and DeepSeek V4 need server-class memory. On Ollama you mostly reach them through cloud tags, which defeats the point.

Quantization in one paragraph

Quantization stores 16-bit weights in fewer bits so they fit in less memory. Q4 (about 4-5 bits per weight) is the usual default and keeps most of the quality. Q8 is near lossless but doubles the size. Below Q4, quality drops fast, especially for code. Start with the runner's default.

Hardware tiers: what runs where

The single number that matters is memory the model can live in. On a PC that is GPU VRAM. On a Mac it is unified memory. Speed then depends mostly on memory bandwidth.

Your machineComfortable model size at Q4Example models
CPU only, 16-32GB RAMUp to about 8B, slowlyGemma 4 E4B
NVIDIA 8GB (RTX 5060, 5050)Up to about 8BGemma 4 E4B, 12B with tight context
NVIDIA 12-16GB (RTX 5070, 5070 Ti, 5080, 5060 Ti 16GB)Up to about 20Bgpt-oss-20b, Gemma 4 12B
NVIDIA 24-32GB (used RTX 3090 or 4090, RTX 5090)Up to about 35BQwen3.6 27B, Gemma 4 31B
Mac with 64-128GB unified memoryUp to about 120Bgpt-oss-120b
DGX Spark or Mac Studio M5 Ultra120B to 200B classgpt-oss-120b with long context

VRAM figures come from NVIDIA's comparison page. The RTX 5090 has 32GB, the 5080 and 5070 Ti have 16GB, the 5070 has 12GB, and the 5060 Ti comes in 8GB or 16GB. DGX Spark has 128GB of unified memory and NVIDIA rates it for inference on models up to 200 billion parameters. Apple's Mac Studio specs list up to 128GB on M5 Max and up to 512GB on M5 Ultra.

Choosing a card is its own topic. We cover current prices and the VRAM-to-model table in detail in our guide to the best GPU for local AI.

Step by step: your first local model

The fastest path with Ollama. LM Studio does the same with buttons.

  1. Install the runnerDownload Ollama from ollama.com, or use the official install script on Linux. It starts a background service on port 11434.
  2. Pull a model that fitsPick one size from the tables above. On a 16GB machine start with gpt-oss:20b or gemma4:e4b. On a 24GB GPU try qwen3.6:27b.
  3. Chat in the terminalRun the model and type. The first load takes a few seconds while weights move into memory.
  4. Call it from codePoint any OpenAI client at the local endpoint. Your existing scripts work with a changed base URL.

The commands:

ollama pull qwen3.6:27b
ollama run qwen3.6:27b
ollama ps

And from code, using the OpenAI compatible endpoint that Ollama documents:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
reply = client.chat.completions.create(
    model="qwen3.6:27b",
    messages=[{"role": "user", "content": "Summarize this contract clause in plain English: ..."}],
)
print(reply.choices[0].message.content)

The API key is required by the client library but ignored by Ollama.

Realistic limits nobody puts in the thumbnail

They are not frontier models. A 27B or 31B model is very good at summarizing, drafting, classifying, extracting data and everyday coding help. It is noticeably weaker than the top cloud models at long multi-step reasoning and big refactors. For heavy agentic coding we still use cloud models, and we watch the bill closely. Our notes on Claude usage limits explain that trade-off.

Speed depends on bandwidth. Every generated token reads the active weights from memory. A rough ceiling is memory bandwidth divided by model size. That is why a 17GB dense model feels slower on a 273 GB/s DGX Spark than on a desktop GPU, and why MoE models feel fast everywhere.

Long context costs memory. A 256K context model will not give you 256K on a 24GB card. Raise the window only as far as memory allows.

One machine, few users. A desktop GPU serves one person well. Ten people at once need vLLM and more hardware.

You are the ops team. Updates, model choice, prompt formats and security are on you. Never expose port 11434 or 1234 to the internet without authentication in front of it.

When to do it yourself and when to bring in a team

Running a model on your own laptop is a weekend project. Do it yourself.

Putting a local or private model behind a real product is different. Authentication, rate limits, logging that does not leak personal data, fallbacks when the model is slow, evaluation so you know answers are actually right. That is where AI features usually break, and it is the part we build for clients. If you want an assistant on your site, our guide to adding an AI chatbot to a website is a good next read.

If you would rather have it built properly, our services start with a Web App from 3,499 lei (about EUR 665), with a fixed price agreed up front. Not sure where your current site stands? Start with a free website audit.

Frequently Asked Questions

Can I run an LLM locally without a GPU?

Yes, but keep expectations low. Ollama, LM Studio and llama.cpp all run on CPU alone. With 16-32GB of RAM you can run models up to about 8B parameters at a few tokens per second, which is fine for short tasks and testing but tiring for long chats.

How much RAM do I need to run a local LLM?

At 4-bit quantization, plan for about 0.6GB per billion parameters plus a few GB for context. That means about 8GB free for a small model, 16GB for gpt-oss-20b, around 24GB for a 27-31B model and roughly 80GB for gpt-oss-120b.

Is it legal to use open models like Gemma 4 or Qwen for commercial work?

Gemma 4, Qwen3.6 and gpt-oss are all published under the Apache 2.0 license, which allows commercial use. Always check the license on the specific model card you download, because some other open models use custom licenses with extra conditions.

Is a local LLM really private?

The model runs entirely on your machine and sends prompts nowhere. Privacy breaks when you add cloud sync, web search plugins or cloud model tags, so check you are running a local variant.

What is the best local LLM for coding in 2026?

For a 24-32GB GPU or a Mac with 48GB or more, Qwen3.6 27B is our first pick for coding, with Gemma 4 31B close behind. On 16GB, gpt-oss-20b is the best balance. None of them match the top cloud models on large multi-file changes.

Should I use Ollama or LM Studio to start?

Use LM Studio if you want a visual app and like to browse and compare models. Use Ollama if you are a developer, work in the terminal or plan to call the model from code or a server. Both are free for local use and both expose an OpenAI compatible API.

AILocal AILLMDeveloper Tools
Explore Legacies productsFree website auditTalk to the Legacies team