Ollama Model Sizes Explained: RAM, Speed, and Quality
open-context's --model flag works with any model you've already pulled into
Ollama,
and the right one depends mostly on how much memory your machine has, not on taste.
gpt-oss:20b, the default, needs roughly 16GB of RAM or VRAM; llama3:8b
runs on about half that with lower-quality output; qwen2.5:32b and
llama3:70b need around 32GB and 40GB+ respectively for a real jump in quality.
For most ChatGPT exports, the default is also the right answer — reach for a bigger model only
when you have the hardware and a large, dense conversation history to justify the extra wait.
Key takeaways
- open-context's default model,
gpt-oss:20b, is OpenAI's own open-weight model (Apache 2.0 license) and needs roughly 16GB of system memory to run. llama3:8bis the lightest documented option at about 5GB on disk and the fastest, at the cost of less nuancedpreferences.mdandmemory.mdoutput.qwen2.5:32bneeds about 19GB of VRAM at Q4 quantization, or roughly 32GB of system RAM if you're running it on CPU instead of a GPU.llama3:70bis the highest-quality documented option but needs 40-50GB of memory at typical quantization — realistic mainly on a dedicated GPU server.- Switching models never changes where your data goes: every option runs through the same local Ollama instance, so privacy is identical no matter what you pass to
--model.
Why model choice doesn't change what leaves your machine
It's worth separating two questions that get conflated: "which model is more private" and
"which model writes better output." For open-context there's no tradeoff on the first question
at all. Whether you run gpt-oss:20b or llama3:70b, the
OllamaPreferenceAnalyzer makes exactly one kind of network call — to the Ollama
server you specify with --ollama-host — and nothing else. We covered how that
pipeline works in
How open-context Uses Ollama to Analyze Your Chat History;
the short version is that picking a model only changes the quality and speed of the output, not
whether your conversation history ever leaves your own infrastructure.
How to actually switch models
The CLI accepts any model name you've pulled into Ollama via the --model flag:
npm start -- convert export.zip --model qwen2.5:32b
Running the Docker image instead, set OLLAMA_MODEL as an environment variable rather
than passing a flag:
docker run -p 3000:3000 \
-e OLLAMA_MODEL=llama3:70b \
-v opencontext-data:/root/.opencontext \
adityakarnam/open-context:latest
The HTTP server also exposes GET /api/ollama/models, which lists every model already
pulled on your Ollama host. Whichever way you set it, the model has to already exist locally —
open-context checks for it before running analysis and fails fast with a clear instruction
(ollama pull <model>) rather than silently substituting a different one.
RAM vs. VRAM: what actually decides whether a model runs
Model size on disk isn't the number that matters — the number that matters is how much memory the
model needs loaded at once, and that depends on quantization. gpt-oss:20b ships using
OpenAI's MXFP4 format, which compresses its mixture-of-experts weights down to about 4.25 bits per
parameter, which is how a 20-billion-parameter model fits in roughly 16GB
(TechRadar).
The same logic applies to the other models: a 4-bit quantized qwen2.5:32b needs about
19GB of VRAM on a GPU, or closer to 32GB of plain system RAM if it's running on CPU instead
(LocalLLM).
A GPU with enough VRAM is always faster than CPU-only inference on the same model — but Ollama
will fall back to system RAM and CPU if no GPU is available, just slower.
The four models open-context documents, side by side
| Model | Disk size | Memory needed | Speed | Best for |
|---|---|---|---|---|
gpt-oss:20b | 13GB | ~16GB RAM/VRAM | Medium | Default — best overall balance |
qwen2.5:32b | 20GB | ~19GB VRAM (Q4) or ~32GB RAM on CPU | Medium | Technical, code-heavy conversation history |
llama3:70b | 40GB | ~40-50GB at Q4_K_M | Slow | Maximum accuracy on a GPU server |
llama3:8b | 5GB | ~8GB RAM | Fast | Quick conversions on modest hardware |
These are the four models open-context's own README and CLI examples name directly, but the
--model flag isn't limited to this list — anything in the
Ollama model library
that you've pulled will work the same way.
When the bigger model is actually worth the wait
There's a detail that changes the calculus here: open-context caps how much of your conversation
history it ever sends to any model. Conversations are sorted chronologically and appended to the
analysis prompt until it hits roughly 100,000 tokens (about 400,000 characters) — after that, the
rest is truncated. That cap applies identically to llama3:8b and llama3:70b.
So a bigger model doesn't mean it reads more of your history; it means it tends to produce a more
coherent, specific memory.md and preferences.md from the same amount of text.
In practice, that makes the decision simpler than the table above might suggest. If your export is
a modest number of conversations, the default gpt-oss:20b already has enough text to
work with — upgrading to llama3:70b buys you a marginally better summary at a
significantly longer wait. If your history is large and heavy on technical detail — long
architecture discussions, code reviews, infrastructure debates — qwen2.5:32b is the
documented sweet spot for that kind of content, assuming you have the VRAM for it. Reach for
llama3:70b specifically when you have a GPU server sitting idle and want the best
possible synthesis of your conversation history, not because your export is unusually large.
If you don't have the hardware for any of this
You don't need a GPU at all to use open-context. Running --skip-preferences skips
Ollama entirely and produces a statistics-based preferences.md/memory.md
from conversation counts, date ranges, and topic keywords — instant, and useful if you'd rather
write your own preferences by hand. And if you have access to a more powerful machine elsewhere,
you don't need Ollama running on the same box as open-context: point --ollama-host at
a remote server (--ollama-host http://gpu-server:11434) and the analysis runs there
while open-context itself stays on your laptop. See
How to Migrate Your ChatGPT History to Claude
for the full conversion walkthrough if you haven't run open-context yet.
Convert your ChatGPT export and pick the model that fits your machine.
Get started with open-context →FAQ
Will a bigger Ollama model actually produce a better preferences.md and memory.md?
Usually, yes, for synthesis quality — but not for volume. open-context caps the conversation text it sends to any model at roughly 100,000 tokens (about 400,000 characters) regardless of which model you pick, so a bigger model doesn't read more of your history. It just tends to write a more coherent, nuanced summary of the same capped text.
Can I run gpt-oss:20b without a dedicated GPU?
Yes. gpt-oss:20b is designed to run on consumer hardware with around 16GB of system memory, using CPU-only inference if no GPU is available. It will be noticeably slower than running on a GPU with enough VRAM, but it will still complete.
What happens if I pass --model with a model I haven't pulled yet?
open-context checks Ollama's model list before running analysis and fails fast with "Model '<name>' not found. Run: ollama pull <model>" instead of silently falling back to a different model.
Does changing the --model flag affect privacy or what data leaves my machine?
No. Every model option runs through the same local Ollama instance. Regardless of which model you pick, the only network call open-context's analysis step makes is to that Ollama server — never to a cloud API.
Can I point open-context at a remote Ollama server with a bigger GPU?
Yes. Use --ollama-host http://your-gpu-server:11434 with the CLI, or set OLLAMA_HOST when running the Docker image, to run analysis against a more powerful machine than the one running open-context itself.