open-context open-context
October 3, 2026 · 8 min read

How to Write a New AI Provider Parser in open-context

Adding a new AI provider to open-context means writing one parser file that turns that provider's raw export into the project's normalized conversation schema — nothing else in the pipeline changes. The existing ChatGPT parser in src/parsers/chatgpt.ts and src/parsers/normalizer.ts is the template: parse the raw JSON, walk it into a flat list of messages, then map each message's role and content into the shape every downstream step already expects. A working first version is realistic in an afternoon once you have a real export to test against.

Key takeaways

  • open-context's conversion pipeline has three fixed stages — parse, normalize, format — and only the parse stage is provider-specific; the formatter, Ollama analyzer, and MCP store never see provider-specific fields.
  • The NormalizedConversation / NormalizedMessage types in src/parsers/types.ts are the actual contract a new parser has to satisfy — there's no abstract base class to extend, just that shape.
  • ChatGPT's export isn't a flat transcript — it's a tree of message nodes with parent/child pointers, because edited or regenerated messages create branches. The ChatGPT parser walks that tree with BFS; a flat-array provider can skip that step entirely.
  • New parsers get fixtures under tests/fixtures/<provider>/ built from synthetic data, never a real export — open-context's own test suite never ships anyone's actual conversations.
  • Once a parser's own tests pass, wiring it into the CLI's convert command is the only integration point left; everything downstream is already provider-agnostic.

Why the pipeline is split this way

open-context's conversion pipeline is three stages: parse the provider's raw export, normalize it into a shared schema, then format that normalized data into markdown (and, optionally, run it through the Ollama preference/memory analyzer). The CLI (src/index.ts) wires these stages together in order inside convertExport(): extract the zip, parse conversations, normalize them, write markdown, then analyze. Each stage only depends on the output type of the one before it — the formatter and analyzer are written entirely against NormalizedConversation[], and have no idea whether that data originally came from ChatGPT, Gemini, or anything else. That's the whole reason adding a provider is a one-file change instead of a project-wide one.

The schema a new parser has to produce

Everything downstream is built against two interfaces in src/parsers/types.ts:

interface NormalizedConversation {
  id: string;
  title: string;
  created: Date;
  updated: Date;
  messages: NormalizedMessage[];
}

interface NormalizedMessage {
  role: 'user' | 'assistant' | 'system';
  content: string;
  images?: string[];
  timestamp: Date;
  metadata?: { model?: string; attachments?: string[] };
}

There's no Parser interface or abstract class to implement — ChatGPTParser is just a plain class with the methods it happens to need (parseConversations, walkMessageTree, extractTextContent, extractImages). A new provider's parser doesn't have to match that method-by-method; it only has to end up producing conversations and messages in the shape above. ConversationNormalizer.normalize() is the function that actually builds a NormalizedConversation, and it's reasonable for a new provider to follow the same split: one class that reads the raw export, one function that maps it onto the schema.

What the ChatGPT parser already solves

Reading chatgpt.ts before writing a new parser is worth the time, because it already handles the two hardest problems a chat export tends to have:

  1. Non-linear history. A ChatGPT export's mapping field is a dictionary of message nodes, each with a parent and a list of children — not an ordered array. Regenerating or editing a message creates a new branch instead of overwriting the old one. walkMessageTree() finds the node with no parent and walks the tree breadth-first, collecting every message it reaches. A provider that exports a flat, already-ordered array of turns doesn't need any of this — most won't.
  2. Content that isn't a plain string. A message's text can live in content.text or inside a content.parts array that mixes strings with image objects (asset_pointer). extractTextContent() and extractImages() split those apart so the normalizer only ever has to deal with a string and a string[].

normalizer.ts then does the smaller, provider-specific mapping: collapsing ChatGPT's tool role into assistant, converting Unix-epoch timestamps (create_time * 1000) into Date objects, and pulling model_slug and attachment names into metadata. This is usually the shortest part of a new parser — most of the work is understanding the source format, not the normalization itself.

Step by step: adding a provider

1. Get a real export and read it

Request or generate a sample export from the provider (Google Takeout for Gemini, for example) and open the raw JSON before writing any code. Decide up front whether messages arrive as a flat array (easy) or some other tree/graph shape (match the ChatGPT pattern).

2. Write src/parsers/<provider>.ts

Parse the raw file, and produce whatever intermediate list of messages makes sense for that format — it doesn't need to look like ChatGPTParser internally.

3. Map into the normalized schema

Convert roles into 'user' | 'assistant' | 'system', join or extract the message text into a single content string, and convert the provider's timestamp format into a Date. This is the step most likely to lose information if rushed — check what the provider calls a "system" turn versus a tool call before collapsing roles.

4. Add fixtures and tests under tests/fixtures/<provider>/

Use small, synthetic conversations — a normal back-and-forth, an empty message, and anything the format does that ChatGPT's doesn't (a different attachment shape, multi-turn branching, etc). Real exported conversations, yours or anyone else's, never belong in the repo.

5. Register the parser in the CLI

convertExport() in src/index.ts currently calls ChatGPTParser directly; a second provider means branching on a flag or detected format before choosing which parser and normalizer to run, the same way the convert command already branches on --skip-preferences.

Nothing past step 5 needs to change. MarkdownFormatter, the Ollama-backed preference and memory analyzer, and the MCP context store all consume NormalizedConversation[] and have no provider-specific code paths to touch.

Why Gemini is the open case

CLAUDE.md and the README both already list Google Gemini as "planned" rather than implemented — open-context currently ships exactly one parser. A Gemini parser is a good first exercise precisely because Google Takeout's export format is structurally different enough to actually test the schema: it's organized as an activity log rather than ChatGPT's per-conversation tree, so the interesting work is in step 1 and step 3 above, not in copying walkMessageTree().

The portability problem a parser like this is solving

None of this would be necessary if chat platforms exported to a shared format. They don't, even though data portability has been a legal right in the EU since GDPR Article 20 took effect in 2018 — it entitles you to your own data in a "structured, commonly used, machine-readable format," but it doesn't require that format to match any other platform's import tool (The IP Press, on data portability and AI). ChatGPT's own export is a good example of the gap: it's machine-readable JSON, but its tree-shaped mapping structure is specific enough to OpenAI's product that almost nothing else can read it directly without a translation step — which is exactly what chatgpt.ts and normalizer.ts are. Every new parser in this project is a small, concrete fix to that gap for one more provider.

Already have a provider in mind? Start from the existing ChatGPT parser.

Explore open-context on GitHub →

FAQ

Does open-context already support importing from Gemini?

Not yet. The CLI, web UI, and Docker image currently parse ChatGPT's conversations.json export only; Gemini (via Google Takeout) is listed as planned in both the README and CLAUDE.md.

Do I need to modify the markdown formatter or MCP server to add a provider?

No. MarkdownFormatter, the Ollama preference/memory analyzer, and the MCP context store are all written against the normalized NormalizedConversation type and have no provider-specific logic to update.

Is there an abstract parser class or interface I have to implement?

No. ChatGPTParser is a plain class, not an implementation of a shared interface. The actual contract a new parser has to satisfy is the output shape defined by NormalizedConversation and NormalizedMessage in src/parsers/types.ts.

Why does the ChatGPT parser walk a tree with BFS instead of just reading an array of messages?

Because ChatGPT's export isn't a linear transcript — editing or regenerating a message creates a new branch inside the mapping object instead of replacing the old one. The parser has to walk parent/child pointers to collect every message, starting from the node with no parent.

Where should test fixtures for a new provider live?

Under tests/fixtures/<provider>/, built from small synthetic conversations you write yourself — never from a real exported chat history, yours or someone else's.