Local models: a complete overview
Open language models on your own server: seven families, memory for each size, local versus cloud, three examples with Ollama, mistakes and rules for production.
In short
Local models are open language models that run on your own server or computer instead of a provider’s cloud: Qwen, Llama, Mistral, Gemma, gpt-oss and others. Data never leaves your perimeter, there is no bill per request, and the model works without the internet. The price is hardware and quality: a model that fits on an ordinary server is weaker than the best cloud ones, so local models are chosen for clear, repeatable tasks — classification, extraction of data from documents, search by meaning, short answers from a knowledge base. Tools such as Ollama and llama.cpp start a model with one command and give it the same API as the cloud.
Local models at a glance
The main facts in one table — what is meant, what is needed and how it is run.
- What it is
- Language models with open weights that run on your own hardware
- Families
- Qwen, Llama, Mistral, Gemma, gpt-oss, DeepSeek, Phi
- Sizes
- From 1 to hundreds of billions of parameters; for a server usually 4–32 billion
- Quantisation
- Weights compressed to 4–8 bits: several times less memory, a small loss of quality
- How to run
- Ollama, llama.cpp, vLLM, LM Studio
- API
- OpenAI-compatible — code moves between cloud and local by changing the address
- Licences
- From Apache 2.0 and MIT to licences with conditions — read before commercial use
Open model families
Seven families most often run locally. New versions come out every few months, so the choice is made on your own examples, not by rankings.
| Family | Author | Licence | Strong side |
|---|---|---|---|
| Qwen | Alibaba | mostly Apache 2.0 | many sizes, many languages |
| Llama | Meta | own, with conditions | the largest ecosystem |
| Mistral | Mistral AI | Apache 2.0 for many models | compact and fast |
| Gemma | own, with conditions | small models, pictures as input | |
| gpt-oss | OpenAI | Apache 2.0 | reasoning and tools |
| DeepSeek | DeepSeek | MIT for many models | reasoning, code |
| Phi | Microsoft | MIT | very small models |
How much memory a model needs
Approximate memory for models compressed to 4 bits, plus a margin for the context. On a graphics card it is video memory; without one — ordinary memory, and the answer is several times slower.
| Model size | Memory | Suits |
|---|---|---|
| 1–4 billion | about 2–4 GB | classification, short extraction, phones |
| 7–9 billion | about 6–8 GB | answers from a knowledge base, summaries |
| 12–14 billion | about 10–12 GB | more complex texts and instructions |
| 27–32 billion | about 20–24 GB | close to cloud quality on many tasks |
| 70 billion and more | from about 40 GB | a dedicated server with graphics cards |
| Embedding models | about 0.5–2 GB | search by meaning, even without a graphics card |
Local model or cloud API
The choice is made per task: one project often uses both.
| Criterion | Local model | Cloud API |
|---|---|---|
| Where the data goes | nowhere | to the provider |
| Quality on hard tasks | lower | the best available |
| Cost | hardware, then almost free | per request |
| Large volumes | profitable | the bill grows with the volume |
| Speed | depends on your hardware | high and stable |
| Without the internet | works | does not |
| Maintenance | on you: updates, monitoring | on the provider |
What running a local model looks like: 3 examples
Starting a model with Ollama, one client for cloud and local models, and settings for a production server.
Start with Ollama
Two commands to a working model, and the same model over HTTP for applications.
# Download a model and ask it a question — all on your own server
ollama pull qwen3:8b
ollama run qwen3:8b "Summarise the delivery terms in two sentences"
# The same model over HTTP: an OpenAI-compatible API on port 11434
curl http://localhost:11434/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "qwen3:8b", "messages": [{"role": "user", "content": "Hello"}]}'
One client for both
The code does not know where the model is: the address and the name come from settings.
// One client for a cloud and a local model: only the address, the key and the model change
const LLM_URL = process.env.LLM_URL ?? 'http://localhost:11434/v1';
const LLM_MODEL = process.env.LLM_MODEL ?? 'qwen3:8b';
export async function complete(prompt: string): Promise<string> {
const res = await fetch(`${LLM_URL}/chat/completions`, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.LLM_KEY ?? 'local'}`,
},
body: JSON.stringify({
model: LLM_MODEL,
messages: [{ role: 'user', content: prompt }],
temperature: 0.2, // fewer surprises in business answers
}),
});
if (!res.ok) throw new Error(`LLM: HTTP ${res.status}`);
const data = (await res.json()) as { choices: { message: { content: string } }[] };
return data.choices[0].message.content;
}
Settings for a server
Closed to the outside, warm between requests and limited in memory so as not to affect the sites next to it.
# /etc/systemd/system/ollama.service.d/override.conf — Ollama on a production server
[Service]
# Listen only for applications on this server, not for the internet
Environment="OLLAMA_HOST=127.0.0.1:11434"
# Keep the model in memory between requests — no cold start
Environment="OLLAMA_KEEP_ALIVE=24h"
# How many requests one model serves at the same time
Environment="OLLAMA_NUM_PARALLEL=4"
Environment="OLLAMA_MAX_LOADED_MODELS=1"
# A ceiling: the model cannot take the memory of the whole server
MemoryMax=24G
Common mistakes with local models
-
Expecting cloud quality
A model that fits on one server will not replace the best cloud model on complex tasks.
-
Choosing by rankings
Public benchmarks say little about your documents and your language.
-
An API open to the internet
A model server without authentication becomes free computing for strangers.
-
No memory limit
A large model loaded next to the sites takes their memory, and they start to fail.
-
Ignoring the licence
Some open models have conditions for commercial use and large audiences.
-
One model for everything
A small model for sorting and a strong one for complex answers are cheaper and better than one compromise.
6 rules for local models in production
-
01
A test set first
30–50 real tasks with expected answers — every model is compared on them.
-
02
The smallest model that copes
Faster answers, less memory, more requests in parallel.
-
03
The same API as the cloud
The model can be swapped either way without rewriting the code.
-
04
Closed to the outside
Only applications on the server or the internal network can reach it.
-
05
Limits on memory and processor
The model shares the server with other work and must not take everything.
-
06
Versions are pinned
An exact model version, and a new one only after a run on the test set.
Questions about local models
What is a local language model?
An open model that runs on your own server or computer instead of a provider’s cloud.
Do I need a graphics card?
For fast answers from models of 7 billion parameters and more — yes; small models and embeddings work on the processor.
Are local models as good as Claude or GPT?
On narrow, clear tasks they come close; on complex reasoning and long texts the cloud is stronger.
Is it really free?
The model is free; the server, electricity and maintenance are not. It pays off on large volumes.
Ollama or vLLM?
Ollama — simple start and moderate load; vLLM — many parallel users on graphics cards.
Can a local model be fine-tuned?
Yes, open weights allow it; but for knowledge of your documents RAG is usually enough.
Can it run on a laptop?
Yes, models up to about 8 billion parameters work on a modern laptop with 16 GB of memory.
Online form
AI inside
your perimeter
I choose the model for the task: Claude or GPT where quality matters most, local models where data cannot leave your servers. Tell me about the task — I answer within one working day.