How to Run an LLM Locally in 2026: Tools, Models, Hardware
A step-by-step guide to running an LLM locally in 2026: Ollama, LM Studio, llama.cpp and vLLM, which open models to pick, and the hardware you really need.
Legacies is a software and web studio from Romania, founded by Horia Stan and Alexandru Talnaci, and we run open models on our own machines every week for testing, private drafts and prototypes. This guide is the setup we would give a friend who wants to do the same.
To run an LLM locally in 2026, install a runner like Ollama or LM Studio, download an open-weight model that fits in your GPU memory or Mac unified memory, and chat with it or call it through a local API. A 16GB machine runs good small models. A 24-32GB GPU or a 64GB+ Mac runs models that feel close to last year's cloud assistants.
Quick answer
- Easiest start: Ollama from the terminal or LM Studio with a full desktop interface. Both are free for local use.
- Good first models: Gemma 4 (E4B, 12B, 26B MoE or 31B), Qwen3.6 (27B or 35B) and gpt-oss-20b. All three families ship under Apache 2.0.
- Memory rule: at 4-bit quantization a model needs roughly 0.6GB per billion parameters, plus a few GB for context.
- Realistic limits: local models are private and free per token, but slower and less capable than the best cloud models on hard reasoning and long agent runs.
Why run a model locally at all
There are three honest reasons. Privacy, cost control and offline work.
Privacy is the big one. When the model runs on your machine, your prompts and documents never leave it. No vendor logs, no retention policy. For contracts, client code or anything under NDA, that matters.
Cost is the second. After you own the hardware, every token is free except electricity. You stop rationing. For the full math, we did it in our breakdown of local AI vs cloud API cost.
Offline is the third. A local model works on a plane or inside a network that blocks external AI services.
What you give up is raw capability. More on that below.
The four tools that matter in 2026
Pick one tool based on how you like to work.
| Tool | Best for | Interface | License | Local API |
|---|---|---|---|---|
| Ollama | Developers, servers, scripting | CLI plus desktop app | MIT, open source | Port 11434, OpenAI and Anthropic compatible |
| LM Studio | Exploring models, non-developers | Full desktop app plus lms CLI | Proprietary, free for home and work | Port 1234, OpenAI and Anthropic compatible |
| llama.cpp | Maximum control, odd hardware | CLI and llama-server | MIT, open source | OpenAI compatible server |
| vLLM | Serving many users at once on NVIDIA | Python server | Apache 2.0 | OpenAI compatible server |
Ollama is the default for most developers. It is MIT licensed, runs on macOS, Windows, Linux and Docker, and the latest release at the time of writing is v0.35.1. Local use is unlimited on every plan, according to Ollama's pricing page. The paid plans only cover their cloud models.
LM Studio is the friendliest. You browse models, see which ones fit your hardware, and chat in a clean interface. It runs GGUF and MLX models, supports MCP servers and document chat, and has been free for work use since July 2025. The app itself is closed source.
llama.cpp is the engine underneath much of this ecosystem, with CUDA, Metal, Vulkan, AMD HIP and CPU backends. Use it directly when you want every knob.
vLLM is for production serving, when many people hit the same GPU at once. Overkill for one person on a laptop.
If you are stuck between the first two, read our Ollama vs LM Studio comparison.
Which open models are worth running now
These are the families we would start with in October 2026, with sizes taken from the Ollama library and the official model cards.
| Model | Sizes | Download at default quant | Context | License |
|---|---|---|---|---|
| Gemma 4 E4B | 4.5B effective | About 7-10GB | 128K | Apache 2.0 |
| Gemma 4 12B | 12B | About 8GB | 256K | Apache 2.0 |
| Gemma 4 26B | 25.2B total, about 4B active (MoE) | About 16-19GB | 256K | Apache 2.0 |
| Gemma 4 31B | 30.7B dense | About 19-20GB | 256K | Apache 2.0 |
| Qwen3.6 27B | 27B dense | About 18-19GB | 262K native | Apache 2.0 |
| gpt-oss-20b | 20B | 14GB | 128K | Apache 2.0 |
| gpt-oss-120b | 120B | 65GB | 128K | Apache 2.0 |
The Gemma 4 model card lists five sizes from E2B to 31B. The Qwen3.6-27B card describes a dense 27B model with a native 262,144 token context. gpt-oss uses MXFP4 weights, so the 20B fits in about 16GB of memory and the 120B fits on a single 80GB GPU.
Our short picks:
- Laptop with 16GB: Gemma 4 E4B or gpt-oss-20b.
- 24-32GB GPU or 48GB+ Mac: Qwen3.6 27B for coding, Gemma 4 31B for general writing and vision.
- Want speed on modest hardware: Gemma 4 26B. Mixture-of-Experts models only activate a slice of their weights per token, so they run much faster than their total size suggests.
Bigger open models like GLM-5.3 and DeepSeek V4 need server-class memory. On Ollama you mostly reach them through cloud tags, which defeats the point.
Quantization in one paragraph
Quantization stores 16-bit weights in fewer bits so they fit in less memory. Q4 (about 4-5 bits per weight) is the usual default and keeps most of the quality. Q8 is near lossless but doubles the size. Below Q4, quality drops fast, especially for code. Start with the runner's default.
Hardware tiers: what runs where
The single number that matters is memory the model can live in. On a PC that is GPU VRAM. On a Mac it is unified memory. Speed then depends mostly on memory bandwidth.
| Your machine | Comfortable model size at Q4 | Example models |
|---|---|---|
| CPU only, 16-32GB RAM | Up to about 8B, slowly | Gemma 4 E4B |
| NVIDIA 8GB (RTX 5060, 5050) | Up to about 8B | Gemma 4 E4B, 12B with tight context |
| NVIDIA 12-16GB (RTX 5070, 5070 Ti, 5080, 5060 Ti 16GB) | Up to about 20B | gpt-oss-20b, Gemma 4 12B |
| NVIDIA 24-32GB (used RTX 3090 or 4090, RTX 5090) | Up to about 35B | Qwen3.6 27B, Gemma 4 31B |
| Mac with 64-128GB unified memory | Up to about 120B | gpt-oss-120b |
| DGX Spark or Mac Studio M5 Ultra | 120B to 200B class | gpt-oss-120b with long context |
VRAM figures come from NVIDIA's comparison page. The RTX 5090 has 32GB, the 5080 and 5070 Ti have 16GB, the 5070 has 12GB, and the 5060 Ti comes in 8GB or 16GB. DGX Spark has 128GB of unified memory and NVIDIA rates it for inference on models up to 200 billion parameters. Apple's Mac Studio specs list up to 128GB on M5 Max and up to 512GB on M5 Ultra.
Choosing a card is its own topic. We cover current prices and the VRAM-to-model table in detail in our guide to the best GPU for local AI.
Step by step: your first local model
The fastest path with Ollama. LM Studio does the same with buttons.
- Install the runnerDownload Ollama from ollama.com, or use the official install script on Linux. It starts a background service on port 11434.
- Pull a model that fitsPick one size from the tables above. On a 16GB machine start with gpt-oss:20b or gemma4:e4b. On a 24GB GPU try qwen3.6:27b.
- Chat in the terminalRun the model and type. The first load takes a few seconds while weights move into memory.
- Call it from codePoint any OpenAI client at the local endpoint. Your existing scripts work with a changed base URL.
The commands:
ollama pull qwen3.6:27b
ollama run qwen3.6:27b
ollama ps
And from code, using the OpenAI compatible endpoint that Ollama documents:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
reply = client.chat.completions.create(
model="qwen3.6:27b",
messages=[{"role": "user", "content": "Summarize this contract clause in plain English: ..."}],
)
print(reply.choices[0].message.content)
The API key is required by the client library but ignored by Ollama.
Realistic limits nobody puts in the thumbnail
They are not frontier models. A 27B or 31B model is very good at summarizing, drafting, classifying, extracting data and everyday coding help. It is noticeably weaker than the top cloud models at long multi-step reasoning and big refactors. For heavy agentic coding we still use cloud models, and we watch the bill closely. Our notes on Claude usage limits explain that trade-off.
Speed depends on bandwidth. Every generated token reads the active weights from memory. A rough ceiling is memory bandwidth divided by model size. That is why a 17GB dense model feels slower on a 273 GB/s DGX Spark than on a desktop GPU, and why MoE models feel fast everywhere.
Long context costs memory. A 256K context model will not give you 256K on a 24GB card. Raise the window only as far as memory allows.
One machine, few users. A desktop GPU serves one person well. Ten people at once need vLLM and more hardware.
You are the ops team. Updates, model choice, prompt formats and security are on you. Never expose port 11434 or 1234 to the internet without authentication in front of it.
When to do it yourself and when to bring in a team
Running a model on your own laptop is a weekend project. Do it yourself.
Putting a local or private model behind a real product is different. Authentication, rate limits, logging that does not leak personal data, fallbacks when the model is slow, evaluation so you know answers are actually right. That is where AI features usually break, and it is the part we build for clients. If you want an assistant on your site, our guide to adding an AI chatbot to a website is a good next read.
If you would rather have it built properly, our services start with a Web App from 3,499 lei (about EUR 665), with a fixed price agreed up front. Not sure where your current site stands? Start with a free website audit.
Frequently Asked Questions
Can I run an LLM locally without a GPU?
Yes, but keep expectations low. Ollama, LM Studio and llama.cpp all run on CPU alone. With 16-32GB of RAM you can run models up to about 8B parameters at a few tokens per second, which is fine for short tasks and testing but tiring for long chats.
How much RAM do I need to run a local LLM?
At 4-bit quantization, plan for about 0.6GB per billion parameters plus a few GB for context. That means about 8GB free for a small model, 16GB for gpt-oss-20b, around 24GB for a 27-31B model and roughly 80GB for gpt-oss-120b.
Is it legal to use open models like Gemma 4 or Qwen for commercial work?
Gemma 4, Qwen3.6 and gpt-oss are all published under the Apache 2.0 license, which allows commercial use. Always check the license on the specific model card you download, because some other open models use custom licenses with extra conditions.
Is a local LLM really private?
The model runs entirely on your machine and sends prompts nowhere. Privacy breaks when you add cloud sync, web search plugins or cloud model tags, so check you are running a local variant.
What is the best local LLM for coding in 2026?
For a 24-32GB GPU or a Mac with 48GB or more, Qwen3.6 27B is our first pick for coding, with Gemma 4 31B close behind. On 16GB, gpt-oss-20b is the best balance. None of them match the top cloud models on large multi-file changes.
Should I use Ollama or LM Studio to start?
Use LM Studio if you want a visual app and like to browse and compare models. Use Ollama if you are a developer, work in the terminal or plan to call the model from code or a server. Both are free for local use and both expose an OpenAI compatible API.