open-context open-context
September 15, 2026 · 6 min read

How open-context Uses Ollama to Analyze Your Chat History

open-context uses a local Ollama model — gpt-oss:20b by default — to read your exported chat history and write two Claude-ready documents: a preferences.md file describing how you like to communicate, and a memory.md file describing who you are and what you're working on. The analysis runs entirely on your own machine; the only network call it makes is to the Ollama instance you're already running, and skipping AI analysis with --skip-preferences produces a lighter, rule-based version of both files with no model involved at all.

Key takeaways

  • open-context's OllamaPreferenceAnalyzer sends your exported conversations to a local Ollama server — never to a cloud API.
  • It runs two separate prompts against the same history: one writes first-person communication preferences, the other writes third-person factual memory.
  • The default model is gpt-oss:20b, OpenAI's own open-weight model, small enough to run on machines with around 16 GB of memory.
  • Conversations are sorted chronologically and capped at roughly 100,000 tokens (about 400,000 characters) before being sent to the model, so very large exports get truncated rather than failing.
  • No Ollama running? --skip-preferences produces a statistics-based fallback — topic keywords, message counts, date ranges — with no model call at all.

Two documents, two different prompts

When you run open-context against a ChatGPT export with AI analysis enabled, it doesn't ask one model call to do everything. The OllamaPreferenceAnalyzer class in src/analyzers/ollama-preferences.ts makes two separate requests against the same conversation history, because the two outputs need to read completely differently.

preferences.md is written in first person — "I prefer clear, direct explanations..." — because it's meant to be pasted directly into Claude's Settings → Preferences field, where Claude reads it as your own voice describing how you want to be talked to. memory.md is written in third person, split into three fixed sections (Work context, Personal context, Top of mind), because it's meant for Claude's Memory feature, which expects factual statements about you rather than direct instructions.

How the analysis pipeline actually works

Under the hood, both prompts go through the same four-step pipeline:

  1. Check Ollama is reachable. checkOllamaAvailable() calls ollama.list() to confirm the server is running and the requested model has already been pulled. A connection refusal surfaces as "Ollama is not running. Start it with: ollama serve"; a missing model surfaces as "Run: ollama pull <model>".
  2. Sort and format conversations. Conversations are sorted oldest-first by creation date, then each one is formatted as a block with its title, date, and every message labeled by role (USER:, ASSISTANT:).
  3. Truncate to fit the context window. Formatted conversations are appended to a running total until it would exceed roughly 100,000 tokens (approximated as 400,000 characters). Once the limit is hit, the loop stops and logs a warning naming how many conversations made it in.
  4. Generate. The assembled text is dropped into either the preferences prompt or the memory prompt and sent to ollama.generate() with stream: false, so the CLI waits for one complete response rather than streaming tokens.

Why the default model is gpt-oss:20b

open-context defaults to gpt-oss:20b, one of two open-weight models OpenAI released under the Apache 2.0 license alongside the larger gpt-oss:120b. The 20B version is specifically sized for local and edge use — it can run on hardware with around 16 GB of memory — which makes it a reasonable default for a tool meant to run on a laptop rather than a GPU server. You aren't locked into it, though; the --model flag accepts anything you've already pulled into Ollama.

ModelSizeSpeedBest for
gpt-oss:20b13GBMediumBest overall results (default)
qwen2.5:32b20GBMediumTechnical content
llama3:70b40GBSlowMaximum accuracy
llama3:8b5GBFastQuick conversions

What happens without Ollama

AI analysis is opt-in, not required. Pass --skip-preferences, or simply don't have Ollama running, and open-context falls back to generateBasicPreferences() and generateBasicMemory() — pure statistics, no model call. These functions count your conversations and messages, compute the date range your export covers, and pull frequent words out of conversation titles as a rough topic list, then drop all of that into the same preferences.md/memory.md shape the AI path produces. It's less specific than what a model writes, but it's instant and it never touches a network at all — useful if you just want the raw conversation markdown quickly and plan to refine preferences by hand.

Pointing it at your own Ollama setup

Two flags control where and what model the analysis runs against:

npm start -- convert export.zip --model qwen2.5:32b
npm start -- convert export.zip --ollama-host http://192.168.1.100:11434

--ollama-host matters most if Ollama runs on a different machine — a desktop with a GPU, say, while you run open-context from a laptop. The Docker image handles the common case of "Ollama on your host, open-context in a container" automatically, reaching it at host.docker.internal:11434 without any extra configuration. If you've already gone through the ChatGPT export and conversion steps, these two flags are the only difference between the default run and a customized one.

Privacy: the one network call that matters

The reason open-context routes through Ollama instead of a hosted API is straightforward: your chat history is some of the most personal data you generate, and running the model locally means it never leaves your machine. This isn't unique to open-context — it's the whole reason Ollama exists as a category. Because inference runs entirely on your own hardware, your prompts stay on the device instead of sitting in a cloud provider's logs, which is also why self-hosted LLMs have become a common recommendation for privacy- and compliance-sensitive work. For open-context specifically, that means the CLI, the web UI, and the MCP server share one rule: no conversation content leaves your machine unless you've pointed --ollama-host at infrastructure you control.

See what preferences.md and memory.md look like for your own history.

Get started with open-context →

FAQ

Does open-context send my conversations to OpenAI or Anthropic for analysis?

No. open-context only sends conversation text to a locally running Ollama server, which you control. It never makes a call to OpenAI's, Anthropic's, or any other vendor's API during analysis.

What Ollama model does open-context use by default?

gpt-oss:20b, OpenAI's own open-weight model released under the Apache 2.0 license. It's sized to run on machines with around 16 GB of memory, and you can swap in a different model with the --model flag.

What happens if I don't have Ollama installed?

open-context falls back to a statistics-based analysis with no model call at all: conversation counts, date ranges, and topic keywords pulled from your conversation titles, formatted into the same preferences.md and memory.md files.

How much of my conversation history actually gets analyzed?

Conversations are sorted chronologically and added to the analysis prompt until it hits roughly 100,000 tokens (about 400,000 characters). If your export is larger than that, the oldest conversations that still fit are kept and the rest are truncated, with a warning logged to the console.

Can I use a bigger or different model for better results?

Yes. The --model flag accepts any model you've pulled into Ollama — qwen2.5:32b and llama3:70b for higher quality at the cost of speed, or llama3:8b for faster, lighter-weight runs.