How to Write a New AI Provider Parser in open-context
Adding a new AI provider to open-context means writing one parser file that turns that
provider's raw export into the project's normalized conversation schema — nothing else in
the pipeline changes. The existing ChatGPT parser in src/parsers/chatgpt.ts and
src/parsers/normalizer.ts is the template: parse the raw JSON, walk it into a
flat list of messages, then map each message's role and content into the shape every
downstream step already expects. A working first version is realistic in an afternoon once
you have a real export to test against.
Key takeaways
- open-context's conversion pipeline has three fixed stages — parse, normalize, format — and only the parse stage is provider-specific; the formatter, Ollama analyzer, and MCP store never see provider-specific fields.
- The
NormalizedConversation/NormalizedMessagetypes insrc/parsers/types.tsare the actual contract a new parser has to satisfy — there's no abstract base class to extend, just that shape. - ChatGPT's export isn't a flat transcript — it's a tree of message nodes with parent/child pointers, because edited or regenerated messages create branches. The ChatGPT parser walks that tree with BFS; a flat-array provider can skip that step entirely.
- New parsers get fixtures under
tests/fixtures/<provider>/built from synthetic data, never a real export — open-context's own test suite never ships anyone's actual conversations. - Once a parser's own tests pass, wiring it into the CLI's
convertcommand is the only integration point left; everything downstream is already provider-agnostic.
Why the pipeline is split this way
open-context's conversion pipeline is three stages: parse the provider's raw export,
normalize it into a shared schema, then format that normalized data into markdown (and,
optionally, run it through the Ollama preference/memory analyzer). The CLI
(src/index.ts) wires these stages together in order inside
convertExport(): extract the zip, parse conversations, normalize them, write
markdown, then analyze. Each stage only depends on the output type of the one before it —
the formatter and analyzer are written entirely against NormalizedConversation[],
and have no idea whether that data originally came from ChatGPT, Gemini, or anything else.
That's the whole reason adding a provider is a one-file change instead of a project-wide one.
The schema a new parser has to produce
Everything downstream is built against two interfaces in src/parsers/types.ts:
interface NormalizedConversation {
id: string;
title: string;
created: Date;
updated: Date;
messages: NormalizedMessage[];
}
interface NormalizedMessage {
role: 'user' | 'assistant' | 'system';
content: string;
images?: string[];
timestamp: Date;
metadata?: { model?: string; attachments?: string[] };
}
There's no Parser interface or abstract class to implement —
ChatGPTParser is just a plain class with the methods it happens to need
(parseConversations, walkMessageTree, extractTextContent,
extractImages). A new provider's parser doesn't have to match that
method-by-method; it only has to end up producing conversations and messages in the shape
above. ConversationNormalizer.normalize() is the function that actually builds a
NormalizedConversation, and it's reasonable for a new provider to follow the same
split: one class that reads the raw export, one function that maps it onto the schema.
What the ChatGPT parser already solves
Reading chatgpt.ts before writing a new parser is worth the time, because it
already handles the two hardest problems a chat export tends to have:
- Non-linear history. A ChatGPT export's
mappingfield is a dictionary of message nodes, each with aparentand a list ofchildren— not an ordered array. Regenerating or editing a message creates a new branch instead of overwriting the old one.walkMessageTree()finds the node with no parent and walks the tree breadth-first, collecting every message it reaches. A provider that exports a flat, already-ordered array of turns doesn't need any of this — most won't. - Content that isn't a plain string. A message's text can live in
content.textor inside acontent.partsarray that mixes strings with image objects (asset_pointer).extractTextContent()andextractImages()split those apart so the normalizer only ever has to deal with astringand astring[].
normalizer.ts then does the smaller, provider-specific mapping: collapsing
ChatGPT's tool role into assistant, converting Unix-epoch
timestamps (create_time * 1000) into Date objects, and pulling
model_slug and attachment names into metadata. This is usually the
shortest part of a new parser — most of the work is understanding the source format, not the
normalization itself.
Step by step: adding a provider
1. Get a real export and read it
Request or generate a sample export from the provider (Google Takeout for Gemini, for example) and open the raw JSON before writing any code. Decide up front whether messages arrive as a flat array (easy) or some other tree/graph shape (match the ChatGPT pattern).
2. Write src/parsers/<provider>.ts
Parse the raw file, and produce whatever intermediate list of messages makes sense for that
format — it doesn't need to look like ChatGPTParser internally.
3. Map into the normalized schema
Convert roles into 'user' | 'assistant' | 'system', join or extract the message
text into a single content string, and convert the provider's timestamp format
into a Date. This is the step most likely to lose information if rushed — check
what the provider calls a "system" turn versus a tool call before collapsing roles.
4. Add fixtures and tests under tests/fixtures/<provider>/
Use small, synthetic conversations — a normal back-and-forth, an empty message, and anything the format does that ChatGPT's doesn't (a different attachment shape, multi-turn branching, etc). Real exported conversations, yours or anyone else's, never belong in the repo.
5. Register the parser in the CLI
convertExport() in src/index.ts currently calls
ChatGPTParser directly; a second provider means branching on a flag or detected
format before choosing which parser and normalizer to run, the same way the
convert command already branches on --skip-preferences.
Nothing past step 5 needs to change. MarkdownFormatter, the
Ollama-backed preference and memory analyzer,
and the MCP context store all consume NormalizedConversation[] and have no
provider-specific code paths to touch.
Why Gemini is the open case
CLAUDE.md and the README both already list Google Gemini as "planned" rather than
implemented — open-context currently ships exactly one parser. A Gemini parser is a good
first exercise precisely because Google Takeout's export format is structurally different
enough to actually test the schema: it's organized as an activity log rather than ChatGPT's
per-conversation tree, so the interesting work is in step 1 and step 3 above, not in copying
walkMessageTree().
The portability problem a parser like this is solving
None of this would be necessary if chat platforms exported to a shared format. They don't,
even though data portability has been a legal right in the EU since GDPR Article 20 took
effect in 2018 — it entitles you to your own data in a "structured, commonly used,
machine-readable format," but it doesn't require that format to match any other platform's
import tool (The IP Press, on data portability and AI).
ChatGPT's own export is a good example of the gap: it's machine-readable JSON, but its
tree-shaped mapping structure is specific enough to OpenAI's product that
almost nothing else can read it directly without a translation step — which is exactly what
chatgpt.ts and normalizer.ts are. Every new parser in this project
is a small, concrete fix to that gap for one more provider.
Already have a provider in mind? Start from the existing ChatGPT parser.
Explore open-context on GitHub →FAQ
Does open-context already support importing from Gemini?
Not yet. The CLI, web UI, and Docker image currently parse ChatGPT's conversations.json export only; Gemini (via Google Takeout) is listed as planned in both the README and CLAUDE.md.
Do I need to modify the markdown formatter or MCP server to add a provider?
No. MarkdownFormatter, the Ollama preference/memory analyzer, and the MCP context store are all written against the normalized NormalizedConversation type and have no provider-specific logic to update.
Is there an abstract parser class or interface I have to implement?
No. ChatGPTParser is a plain class, not an implementation of a shared interface. The actual contract a new parser has to satisfy is the output shape defined by NormalizedConversation and NormalizedMessage in src/parsers/types.ts.
Why does the ChatGPT parser walk a tree with BFS instead of just reading an array of messages?
Because ChatGPT's export isn't a linear transcript — editing or regenerating a message creates a new branch inside the mapping object instead of replacing the old one. The parser has to walk parent/child pointers to collect every message, starting from the node with no parent.
Where should test fixtures for a new provider live?
Under tests/fixtures/<provider>/, built from small synthetic conversations you write yourself — never from a real exported chat history, yours or someone else's.