How open-context Uses Ollama to Analyze Your Chat History
open-context uses a local Ollama
model — gpt-oss:20b by default — to read your exported chat history and write two
Claude-ready documents: a preferences.md file describing how you like to communicate,
and a memory.md file describing who you are and what you're working on. The analysis
runs entirely on your own machine; the only network call it makes is to the Ollama instance you're
already running, and skipping AI analysis with --skip-preferences produces a lighter,
rule-based version of both files with no model involved at all.
Key takeaways
- open-context's
OllamaPreferenceAnalyzersends your exported conversations to a local Ollama server — never to a cloud API. - It runs two separate prompts against the same history: one writes first-person communication preferences, the other writes third-person factual memory.
- The default model is
gpt-oss:20b, OpenAI's own open-weight model, small enough to run on machines with around 16 GB of memory. - Conversations are sorted chronologically and capped at roughly 100,000 tokens (about 400,000 characters) before being sent to the model, so very large exports get truncated rather than failing.
- No Ollama running?
--skip-preferencesproduces a statistics-based fallback — topic keywords, message counts, date ranges — with no model call at all.
Two documents, two different prompts
When you run open-context against a ChatGPT export with AI
analysis enabled, it doesn't ask one model call to do everything. The
OllamaPreferenceAnalyzer class in src/analyzers/ollama-preferences.ts makes
two separate requests against the same conversation history, because the two outputs need to read
completely differently.
preferences.md is written in first person — "I prefer clear, direct explanations..." —
because it's meant to be pasted directly into Claude's Settings → Preferences field, where Claude
reads it as your own voice describing how you want to be talked to. memory.md is written
in third person, split into three fixed sections (Work context, Personal context, Top of mind),
because it's meant for Claude's Memory feature, which expects factual statements about you rather
than direct instructions.
How the analysis pipeline actually works
Under the hood, both prompts go through the same four-step pipeline:
- Check Ollama is reachable.
checkOllamaAvailable()callsollama.list()to confirm the server is running and the requested model has already been pulled. A connection refusal surfaces as "Ollama is not running. Start it with:ollama serve"; a missing model surfaces as "Run:ollama pull <model>". - Sort and format conversations. Conversations are sorted oldest-first by creation date, then each one is formatted as a block with its title, date, and every message labeled by role (
USER:,ASSISTANT:). - Truncate to fit the context window. Formatted conversations are appended to a running total until it would exceed roughly 100,000 tokens (approximated as 400,000 characters). Once the limit is hit, the loop stops and logs a warning naming how many conversations made it in.
- Generate. The assembled text is dropped into either the preferences prompt or the memory prompt and sent to
ollama.generate()withstream: false, so the CLI waits for one complete response rather than streaming tokens.
Why the default model is gpt-oss:20b
open-context defaults to gpt-oss:20b, one of two open-weight models
OpenAI released under the Apache 2.0 license
alongside the larger gpt-oss:120b. The 20B version is specifically sized for local
and edge use — it can run on hardware with around 16 GB of memory — which makes it a reasonable
default for a tool meant to run on a laptop rather than a GPU server. You aren't locked into it,
though; the --model flag accepts anything you've already pulled into Ollama.
| Model | Size | Speed | Best for |
|---|---|---|---|
gpt-oss:20b | 13GB | Medium | Best overall results (default) |
qwen2.5:32b | 20GB | Medium | Technical content |
llama3:70b | 40GB | Slow | Maximum accuracy |
llama3:8b | 5GB | Fast | Quick conversions |
What happens without Ollama
AI analysis is opt-in, not required. Pass --skip-preferences, or simply don't have
Ollama running, and open-context falls back to generateBasicPreferences() and
generateBasicMemory() — pure statistics, no model call. These functions count your
conversations and messages, compute the date range your export covers, and pull frequent words out
of conversation titles as a rough topic list, then drop all of that into the same
preferences.md/memory.md shape the AI path produces. It's less specific
than what a model writes, but it's instant and it never touches a network at all — useful if you
just want the raw conversation markdown quickly and plan to refine preferences by hand.
Pointing it at your own Ollama setup
Two flags control where and what model the analysis runs against:
npm start -- convert export.zip --model qwen2.5:32b
npm start -- convert export.zip --ollama-host http://192.168.1.100:11434
--ollama-host matters most if Ollama runs on a different machine — a desktop with a
GPU, say, while you run open-context from a laptop. The Docker image handles the common case of
"Ollama on your host, open-context in a container" automatically, reaching it at
host.docker.internal:11434 without any extra configuration. If you've already gone
through the ChatGPT export and conversion steps,
these two flags are the only difference between the default run and a customized one.
Privacy: the one network call that matters
The reason open-context routes through Ollama instead of a hosted API is straightforward: your
chat history is some of the most personal data you generate, and running the model locally means
it never leaves your machine. This isn't unique to open-context — it's the whole reason
Ollama
exists as a category. Because inference runs entirely on your own hardware, your prompts stay on
the device instead of sitting in a cloud provider's logs, which is also why self-hosted LLMs have
become a common recommendation for privacy- and compliance-sensitive work. For open-context
specifically, that means the CLI, the web UI, and the MCP server share one rule: no conversation
content leaves your machine unless you've pointed --ollama-host at infrastructure you
control.
See what preferences.md and memory.md look like for your own history.
Get started with open-context →FAQ
Does open-context send my conversations to OpenAI or Anthropic for analysis?
No. open-context only sends conversation text to a locally running Ollama server, which you control. It never makes a call to OpenAI's, Anthropic's, or any other vendor's API during analysis.
What Ollama model does open-context use by default?
gpt-oss:20b, OpenAI's own open-weight model released under the Apache 2.0 license. It's sized to run on machines with around 16 GB of memory, and you can swap in a different model with the --model flag.
What happens if I don't have Ollama installed?
open-context falls back to a statistics-based analysis with no model call at all: conversation counts, date ranges, and topic keywords pulled from your conversation titles, formatted into the same preferences.md and memory.md files.
How much of my conversation history actually gets analyzed?
Conversations are sorted chronologically and added to the analysis prompt until it hits roughly 100,000 tokens (about 400,000 characters). If your export is larger than that, the oldest conversations that still fit are kept and the rest are truncated, with a warning logged to the console.
Can I use a bigger or different model for better results?
Yes. The --model flag accepts any model you've pulled into Ollama — qwen2.5:32b and llama3:70b for higher quality at the cost of speed, or llama3:8b for faster, lighter-weight runs.