open-context open-context
October 5, 2026 · 7 min read

Ollama Model Sizes Explained: RAM, Speed, and Quality

open-context's --model flag works with any model you've already pulled into Ollama, and the right one depends mostly on how much memory your machine has, not on taste. gpt-oss:20b, the default, needs roughly 16GB of RAM or VRAM; llama3:8b runs on about half that with lower-quality output; qwen2.5:32b and llama3:70b need around 32GB and 40GB+ respectively for a real jump in quality. For most ChatGPT exports, the default is also the right answer — reach for a bigger model only when you have the hardware and a large, dense conversation history to justify the extra wait.

Key takeaways

  • open-context's default model, gpt-oss:20b, is OpenAI's own open-weight model (Apache 2.0 license) and needs roughly 16GB of system memory to run.
  • llama3:8b is the lightest documented option at about 5GB on disk and the fastest, at the cost of less nuanced preferences.md and memory.md output.
  • qwen2.5:32b needs about 19GB of VRAM at Q4 quantization, or roughly 32GB of system RAM if you're running it on CPU instead of a GPU.
  • llama3:70b is the highest-quality documented option but needs 40-50GB of memory at typical quantization — realistic mainly on a dedicated GPU server.
  • Switching models never changes where your data goes: every option runs through the same local Ollama instance, so privacy is identical no matter what you pass to --model.

Why model choice doesn't change what leaves your machine

It's worth separating two questions that get conflated: "which model is more private" and "which model writes better output." For open-context there's no tradeoff on the first question at all. Whether you run gpt-oss:20b or llama3:70b, the OllamaPreferenceAnalyzer makes exactly one kind of network call — to the Ollama server you specify with --ollama-host — and nothing else. We covered how that pipeline works in How open-context Uses Ollama to Analyze Your Chat History; the short version is that picking a model only changes the quality and speed of the output, not whether your conversation history ever leaves your own infrastructure.

How to actually switch models

The CLI accepts any model name you've pulled into Ollama via the --model flag:

npm start -- convert export.zip --model qwen2.5:32b

Running the Docker image instead, set OLLAMA_MODEL as an environment variable rather than passing a flag:

docker run -p 3000:3000 \
  -e OLLAMA_MODEL=llama3:70b \
  -v opencontext-data:/root/.opencontext \
  adityakarnam/open-context:latest

The HTTP server also exposes GET /api/ollama/models, which lists every model already pulled on your Ollama host. Whichever way you set it, the model has to already exist locally — open-context checks for it before running analysis and fails fast with a clear instruction (ollama pull <model>) rather than silently substituting a different one.

RAM vs. VRAM: what actually decides whether a model runs

Model size on disk isn't the number that matters — the number that matters is how much memory the model needs loaded at once, and that depends on quantization. gpt-oss:20b ships using OpenAI's MXFP4 format, which compresses its mixture-of-experts weights down to about 4.25 bits per parameter, which is how a 20-billion-parameter model fits in roughly 16GB (TechRadar). The same logic applies to the other models: a 4-bit quantized qwen2.5:32b needs about 19GB of VRAM on a GPU, or closer to 32GB of plain system RAM if it's running on CPU instead (LocalLLM). A GPU with enough VRAM is always faster than CPU-only inference on the same model — but Ollama will fall back to system RAM and CPU if no GPU is available, just slower.

The four models open-context documents, side by side

ModelDisk sizeMemory neededSpeedBest for
gpt-oss:20b13GB~16GB RAM/VRAMMediumDefault — best overall balance
qwen2.5:32b20GB~19GB VRAM (Q4) or ~32GB RAM on CPUMediumTechnical, code-heavy conversation history
llama3:70b40GB~40-50GB at Q4_K_MSlowMaximum accuracy on a GPU server
llama3:8b5GB~8GB RAMFastQuick conversions on modest hardware

These are the four models open-context's own README and CLI examples name directly, but the --model flag isn't limited to this list — anything in the Ollama model library that you've pulled will work the same way.

When the bigger model is actually worth the wait

There's a detail that changes the calculus here: open-context caps how much of your conversation history it ever sends to any model. Conversations are sorted chronologically and appended to the analysis prompt until it hits roughly 100,000 tokens (about 400,000 characters) — after that, the rest is truncated. That cap applies identically to llama3:8b and llama3:70b. So a bigger model doesn't mean it reads more of your history; it means it tends to produce a more coherent, specific memory.md and preferences.md from the same amount of text.

In practice, that makes the decision simpler than the table above might suggest. If your export is a modest number of conversations, the default gpt-oss:20b already has enough text to work with — upgrading to llama3:70b buys you a marginally better summary at a significantly longer wait. If your history is large and heavy on technical detail — long architecture discussions, code reviews, infrastructure debates — qwen2.5:32b is the documented sweet spot for that kind of content, assuming you have the VRAM for it. Reach for llama3:70b specifically when you have a GPU server sitting idle and want the best possible synthesis of your conversation history, not because your export is unusually large.

If you don't have the hardware for any of this

You don't need a GPU at all to use open-context. Running --skip-preferences skips Ollama entirely and produces a statistics-based preferences.md/memory.md from conversation counts, date ranges, and topic keywords — instant, and useful if you'd rather write your own preferences by hand. And if you have access to a more powerful machine elsewhere, you don't need Ollama running on the same box as open-context: point --ollama-host at a remote server (--ollama-host http://gpu-server:11434) and the analysis runs there while open-context itself stays on your laptop. See How to Migrate Your ChatGPT History to Claude for the full conversion walkthrough if you haven't run open-context yet.

Convert your ChatGPT export and pick the model that fits your machine.

Get started with open-context →

FAQ

Will a bigger Ollama model actually produce a better preferences.md and memory.md?

Usually, yes, for synthesis quality — but not for volume. open-context caps the conversation text it sends to any model at roughly 100,000 tokens (about 400,000 characters) regardless of which model you pick, so a bigger model doesn't read more of your history. It just tends to write a more coherent, nuanced summary of the same capped text.

Can I run gpt-oss:20b without a dedicated GPU?

Yes. gpt-oss:20b is designed to run on consumer hardware with around 16GB of system memory, using CPU-only inference if no GPU is available. It will be noticeably slower than running on a GPU with enough VRAM, but it will still complete.

What happens if I pass --model with a model I haven't pulled yet?

open-context checks Ollama's model list before running analysis and fails fast with "Model '<name>' not found. Run: ollama pull <model>" instead of silently falling back to a different model.

Does changing the --model flag affect privacy or what data leaves my machine?

No. Every model option runs through the same local Ollama instance. Regardless of which model you pick, the only network call open-context's analysis step makes is to that Ollama server — never to a cloud API.

Can I point open-context at a remote Ollama server with a bigger GPU?

Yes. Use --ollama-host http://your-gpu-server:11434 with the CLI, or set OLLAMA_HOST when running the Docker image, to run analysis against a more powerful machine than the one running open-context itself.