686 lines
32 KiB
Markdown
686 lines
32 KiB
Markdown
# AI Chat Exporter
|
||
|
||
A personal backup tool for ChatGPT and Claude conversation history. Exports your chats to Markdown files and syncs them to [Joplin](https://joplinapp.org/) as notes. Each conversation becomes a single `.md` file with YAML frontmatter, organised into folders that map directly to Joplin notebooks.
|
||
|
||
Supports incremental sync — only new or updated conversations are exported on each run. Every run is resumable: if interrupted, re-running picks up exactly where it left off.
|
||
|
||
---
|
||
|
||
## ⚠️ Terms of Service Warning
|
||
|
||
**Read this before using this tool.**
|
||
|
||
This tool works by accessing **unofficial, undocumented internal web API endpoints** used by the ChatGPT and Claude web apps. These endpoints are not publicly supported by OpenAI or Anthropic and are subject to change or removal without notice.
|
||
|
||
**Use of this tool may conflict with their Terms of Service:**
|
||
- OpenAI: https://openai.com/policies/terms-of-use
|
||
- Anthropic: https://www.anthropic.com/legal/consumer-terms
|
||
|
||
**By using this tool, you accept that:**
|
||
- You are using it entirely at your own risk
|
||
- Your account could potentially be suspended for automated or scripted access
|
||
- The internal APIs this tool relies on may break at any time without notice
|
||
- This tool is for **personal archival use only** — not commercial use
|
||
|
||
This tool is designed for a single user backing up their own conversations. Do not use it to scrape data at scale or for any commercial purpose.
|
||
|
||
---
|
||
|
||
## Installation
|
||
|
||
### Linux / macOS
|
||
|
||
```bash
|
||
git clone <repo-url>
|
||
cd AIChatExporter
|
||
./ai-chat-exporter doctor
|
||
```
|
||
|
||
That's the whole install. The `ai-chat-exporter` wrapper creates `.venv` and
|
||
installs dependencies on first run, and reinstalls whenever `pyproject.toml`
|
||
changes — so there is no `python3 -m venv` / `source .venv/bin/activate` to
|
||
remember, on this machine or the next one you clone onto. Every command in this
|
||
README works the same way:
|
||
|
||
```bash
|
||
./ai-chat-exporter export --provider all
|
||
./ai-chat-exporter sync
|
||
```
|
||
|
||
The wrapper does **not** change directory: `.env`, `cache/` and `exports/` all
|
||
resolve against your current directory, which is what lets one checkout archive
|
||
different machines into different places. Run it from the repo. (It warns if you
|
||
don't, because a different working directory means a different `cache/manifest.json`
|
||
— which would re-export everything and orphan your existing Joplin notes.)
|
||
|
||
Prefer the traditional route, or want the test dependencies? That still works:
|
||
|
||
```bash
|
||
python3 -m venv .venv && source .venv/bin/activate && pip install -e ".[dev]"
|
||
```
|
||
|
||
### Windows
|
||
|
||
No admin access required. Run these in **Command Prompt** (`cmd.exe`) — it's the simplest option on Windows because it doesn't have PowerShell's script execution policy restrictions.
|
||
|
||
```bat
|
||
git clone <repo-url>
|
||
cd AIChatExporter
|
||
ai-chat-exporter doctor
|
||
```
|
||
|
||
`ai-chat-exporter.cmd` does the same bootstrap as the POSIX wrapper — it creates
|
||
`.venv` and installs dependencies on first run. No `python -m venv`, no
|
||
`.venv\Scripts\activate`.
|
||
|
||
**How you invoke it differs between the two Windows shells:**
|
||
|
||
| Shell | Command |
|
||
|---|---|
|
||
| Command Prompt (`cmd.exe`) | `ai-chat-exporter export --provider all` |
|
||
| PowerShell | `.\ai-chat-exporter.cmd export --provider all` |
|
||
|
||
`cmd.exe` searches the current directory before `PATH` and resolves the bare name
|
||
through `PATHEXT`, so it finds `ai-chat-exporter.cmd` with no prefix and no
|
||
extension. PowerShell deliberately does *not* search the current directory, and
|
||
`.\ai-chat-exporter` there would resolve to the extensionless POSIX script, which
|
||
PowerShell cannot execute — so name the `.cmd` explicitly. Command Prompt is the
|
||
simpler of the two here.
|
||
|
||
All `ai-chat-exporter` commands work identically in Command Prompt.
|
||
|
||
**Using PowerShell instead?** If you prefer PowerShell, you may need to allow script execution first (one-time, current user only):
|
||
|
||
```powershell
|
||
Set-ExecutionPolicy RemoteSigned -Scope CurrentUser
|
||
```
|
||
|
||
Then activate the venv and run commands the same way.
|
||
|
||
**Prerequisites:**
|
||
- Python 3.11 or later — install from [python.org](https://www.python.org/downloads/windows/). During installation, tick **"Add Python to PATH"**.
|
||
- Git — install from [git-scm.com](https://git-scm.com/) if not already present.
|
||
|
||
**Notes:**
|
||
- The cache manifest and logs are stored in `cache\` inside the install directory — the same as on Linux.
|
||
- File permission hardening (`chmod 600`) is silently ignored on Windows — not a concern for single-user desktop use.
|
||
- Joplin Web Clipper runs on `localhost:41184` on all platforms; no configuration changes needed.
|
||
|
||
---
|
||
|
||
## First Run: Run Doctor
|
||
|
||
Before anything else, validate your setup:
|
||
|
||
```bash
|
||
ai-chat-exporter doctor
|
||
```
|
||
|
||
This checks token presence, format, token health (via the `/api/auth/session` `error` field), directory permissions, disk space, and live API connectivity. Fix any failures before proceeding.
|
||
|
||
### Checking for API drift
|
||
|
||
The exporter reads ChatGPT's and Claude's undocumented internal web APIs, which can change shape without notice. The worst failure for a backup tool is *silent* — a response change that makes the exporter skip or mis-parse content without erroring. Run the canary to catch that early:
|
||
|
||
```bash
|
||
ai-chat-exporter canary
|
||
```
|
||
|
||
It probes one conversation per provider and checks only the fields the parser depends on (a renamed retrieval-tool author that would bypass the content-collapse, a new content type, drifted message fields, ignored attachments). `OK` means the live schema still matches; `WARN`/`ERROR` tells you exactly what changed. Worth running before a large export or on a schedule.
|
||
|
||
---
|
||
|
||
## Getting Your Session Tokens
|
||
|
||
Session tokens are how your browser stays logged in. This tool uses them to access your chat history on your behalf.
|
||
|
||
Tokens are entered manually — copy them from your browser's DevTools and run the wizard:
|
||
|
||
```bash
|
||
ai-chat-exporter auth
|
||
```
|
||
|
||
The wizard detects your OS, shows the correct DevTools shortcut, and writes the values to `.env` without echoing them to the terminal.
|
||
|
||
> **Why no auto-extraction from the browser?** Modern Chromium browsers (Chrome 127+, current Brave) encrypt cookies on Windows with **App-Bound Encryption**: the key is bound to the browser through a SYSTEM-level service, so reading cookies off disk requires Administrator rights and a PsExec-style SYSTEM impersonation that antivirus flags as credential theft — and it still fails on Brave specifically. There is no reliable, non-invasive way to pull these cookies automatically, so the manual DevTools flow below is the supported path.
|
||
|
||
### Token Lifetimes
|
||
|
||
| Provider | Cookie Name | Lifetime | Expiry Detection |
|
||
|----------|-------------|----------|-----------------|
|
||
| ChatGPT | `__Secure-next-auth.session-token.0` + `.1` | refresh ~weekly | `error` field of `/api/auth/session` — `doctor` reports "ChatGPT token active". The token is an encrypted JWE, so its `exp` is **not** readable client-side, and the `expires` field is a misleading rolling window; the `error` (`RefreshAccessTokenError` when dead) is the honest signal. |
|
||
| Claude | `sessionKey` | ~30 days | Opaque token — only detectable via 401 response |
|
||
|
||
### Finding Tokens in Chrome DevTools
|
||
|
||
1. Open the provider's website and make sure you're logged in
|
||
2. Press **F12** (Windows/Linux) or **Cmd+Option+I** (macOS) to open DevTools
|
||
3. Click the **Application** tab
|
||
4. In the left panel, expand **Cookies** and click the site URL
|
||
5. Find the cookie by name and copy its **Value**
|
||
|
||
**ChatGPT:** go to `https://chatgpt.com` → find **two** cookies:
|
||
- `__Secure-next-auth.session-token.0` — copy Value (starts with `eyJ`) → `CHATGPT_SESSION_TOKEN`
|
||
- `__Secure-next-auth.session-token.1` — copy Value → `CHATGPT_SESSION_TOKEN_1`
|
||
|
||
ChatGPT splits large session tokens across two cookies to stay under the browser's 4KB cookie limit. Both are required.
|
||
|
||
**Claude:** go to `https://claude.ai` → find `sessionKey` → copy Value
|
||
|
||
### When Tokens Expire
|
||
|
||
When a token expires you'll see a `401 Unauthorized` error. To refresh:
|
||
- Re-run the `auth` wizard: `ai-chat-exporter auth`
|
||
- Or manually update the value in your `.env` file
|
||
|
||
---
|
||
|
||
## The `auth` Command
|
||
|
||
The easiest way to configure tokens is the interactive wizard:
|
||
|
||
```bash
|
||
ai-chat-exporter auth
|
||
```
|
||
|
||
This walks you through finding your token, validates it, shows the expiry date (ChatGPT only), and offers to write it to your `.env` automatically. Tokens are never echoed to the terminal.
|
||
|
||
---
|
||
|
||
## `.env` Setup
|
||
|
||
Copy `.env.example` to `.env` and fill in your values:
|
||
|
||
```bash
|
||
cp .env.example .env
|
||
```
|
||
|
||
### Provider tokens
|
||
|
||
| Variable | Description |
|
||
|----------|-------------|
|
||
| `CHATGPT_SESSION_TOKEN` | ChatGPT session token chunk `.0` (starts with `eyJ…`) |
|
||
| `CHATGPT_SESSION_TOKEN_1` | ChatGPT session token chunk `.1` (the remainder) |
|
||
| `CHATGPT_PROJECT_IDS` | Comma-separated ChatGPT project IDs (see below) |
|
||
| `CLAUDE_SESSION_KEY` | Your Claude session key |
|
||
|
||
### Output
|
||
|
||
| Variable | Default | Description |
|
||
|----------|---------|-------------|
|
||
| `EXPORT_DIR` | `./exports` | Where to write exported Markdown files |
|
||
| `OUTPUT_STRUCTURE` | `provider/project/year` | Folder structure (see below) |
|
||
| `EXPORTER_HIDDEN_CONTENT` | `placeholder` | What to do with content invisible in the provider's web UI (file-retrieval tool dumps, Custom Instructions): `placeholder` collapses to a one-line note with size, `full` keeps everything, `omit` drops it. Dumps can be 90% of a conversation's bytes. |
|
||
| `EXPORTER_DOWNLOAD_MEDIA` | `images` | Download conversation assets into a `media/` folder beside each export and inline them: `images` (images only), `all` (also audio/voice clips), `off` (text placeholders only). Downloaded media is uploaded to Joplin as resources on the next `joplin` run. |
|
||
| `MAX_CONVERSATIONS_PER_RUN` | unlimited | Cap downloads per export run (per provider). Runs are resumable, so re-running continues where the cap stopped — useful to spread a large first export across several sessions. |
|
||
| `REQUEST_DELAY` | `1.0` | Seconds between consecutive API requests, with small random jitter, so traffic stays human-paced. Set `0` to disable. |
|
||
|
||
### Joplin
|
||
|
||
| Variable | Default | Description |
|
||
|----------|---------|-------------|
|
||
| `JOPLIN_API_TOKEN` | — | Authorization token from Joplin Web Clipper settings |
|
||
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
|
||
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
|
||
|
||
### Cache & logging
|
||
|
||
| Variable | Default | Description |
|
||
|----------|---------|-------------|
|
||
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
|
||
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
|
||
|
||
---
|
||
|
||
## ChatGPT Projects
|
||
|
||
ChatGPT project conversations are stored separately from your main conversation list and require extra configuration.
|
||
|
||
### Finding your project IDs
|
||
|
||
The quickest way is to let the exporter find them:
|
||
|
||
```
|
||
ai-chat-exporter projects
|
||
```
|
||
|
||
It lists every project your conversations actually belong to, marks which are
|
||
missing from `.env`, and prints a paste-ready `CHATGPT_PROJECT_IDS=` line
|
||
(`--write` updates `.env` for you). If your ChatGPT account does not include
|
||
the project on conversation summaries, add `--deep` and it reads each
|
||
conversation's detail instead — slower, one request per conversation, but
|
||
complete.
|
||
|
||
Why it matters: project attribution is resolved from each conversation's own
|
||
`gizmo_id`, so exports file correctly whether or not a project is configured.
|
||
But the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that
|
||
live only inside a project never appear in the default conversation list, so an
|
||
unlisted project's chats are never fetched at all. Every export run also names
|
||
any unconfigured project it encounters.
|
||
|
||
To find them by hand instead:
|
||
|
||
1. Open ChatGPT and click a Project in the left sidebar
|
||
2. Look at the browser URL — it will look like:
|
||
`https://chatgpt.com/g/g-p-68c2b2b3037c8191890036fb4ae3ed9f-my-project/project`
|
||
3. Copy the `g-p-…` part (everything up to but not including the slug after the second `-`)
|
||
|
||
Add all your project IDs to `.env` as a comma-separated list:
|
||
|
||
```
|
||
CHATGPT_PROJECT_IDS=g-p-68c2b2b3037c8191890036fb4ae3ed9f,g-p-anotherprojectid
|
||
```
|
||
|
||
The `auth` wizard can also guide you through this step interactively.
|
||
|
||
---
|
||
|
||
## Claude Code Sessions
|
||
|
||
The `claude-code` provider archives your local [Claude Code](https://claude.com/claude-code) agent transcripts — no tokens, no API, no ToS exposure. Sessions are read from `~/.claude/projects/`.
|
||
|
||
```bash
|
||
ai-chat-exporter export --provider claude-code
|
||
ai-chat-exporter joplin --provider claude-code
|
||
```
|
||
|
||
Exports are prose-only by default: your prompts and Claude's write-ups are kept, tool activity is grouped into one-line placeholders (`> 🔧 Tool output — 14 calls: Read ×9, Bash ×2 (86KB) — omitted`), and internal reasoning is dropped (counted in the run summary). Set `EXPORTER_HIDDEN_CONTENT=full` to keep everything. Sessions become notes under their own top-level **`AI-ClaudeCode`** Joplin notebook, in a sub-notebook per launch folder. The provider is included in `--provider all` whenever a sessions directory exists.
|
||
|
||
**Subagents.** Claude Code stores subagent (Task-tool) transcripts as separate files under `<session>/subagents/`; each is folded into its parent session inline, as a collapsible `<details>` block labeled with the subagent's type and description (its own tool traffic is collapsed like the main dialogue).
|
||
|
||
**Repo tags.** Sessions launched from a workspace root all share one folder-named notebook, so titles carry the repos each session touched — `Resume StartWRT project work [start-technologies]` — for at-a-glance scanning and search. A file's repo is the git repository it lives in (nearest ancestor with a `.git`), resolved from the tool paths in the transcript, so work is tagged wherever it happened — even across workspaces — and config/one-off files are ignored (they aren't repos). Repos are frequency-ordered and capped at 3. To never tag specific repos, set `CLAUDE_CODE_REPO_TAG_IGNORE` (comma-separated). Note this reads your current git layout, so a repo you later delete or move drops from the tag on re-export.
|
||
|
||
**Multiple locations.** By default the provider scans `~/.claude/projects/` plus `$CLAUDE_CONFIG_DIR/projects` when `CLAUDE_CONFIG_DIR` is set. To scan additional roots (e.g. other machines' sessions copied onto this box), set `CLAUDE_CODE_DIR` to a `:`-separated list:
|
||
|
||
```bash
|
||
CLAUDE_CODE_DIR="$HOME/.claude/projects:/mnt/backup/laptop/.claude/projects"
|
||
```
|
||
|
||
Sessions from all roots are merged by folder (no per-machine label); if the same session UUID appears in two roots, the newer copy wins. Note the exporter only sees this machine's disk and cannot recover sessions Claude Code has already pruned — run it regularly.
|
||
|
||
---
|
||
|
||
## Codex Sessions
|
||
|
||
The `codex` provider archives your local [Codex CLI](https://chatgpt.com/codex) agent transcripts — same deal as Claude Code: no tokens, no API, no ToS exposure. Rollout files are read from `~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl`.
|
||
|
||
```bash
|
||
ai-chat-exporter export --provider codex
|
||
ai-chat-exporter joplin --provider codex
|
||
```
|
||
|
||
Exports are prose-only by default, with the same placeholder format as Claude Code, into their own top-level **`AI-Codex`** Joplin notebook. Repo tags work the same way (`CODEX_REPO_TAG_IGNORE` to suppress), and additional roots can be scanned with `CODEX_DIR` (`:`-separated); `CODEX_HOME`'s `sessions/` is picked up automatically when that variable is set.
|
||
|
||
Three things differ from Claude Code, all forced by how Codex stores its data:
|
||
|
||
**Reasoning cannot be exported.** Codex encrypts it at rest — every reasoning record carries `encrypted_content` with no plaintext summary in any layer of the file. It is always dropped and counted; `EXPORTER_HIDDEN_CONTENT=full` cannot bring it back.
|
||
|
||
**Incomplete tool calls are reported.** Codex writes each session twice in one file: its own typed items (what actually ran) and the raw model-facing wire format (everything attempted). The provider reads the typed layer — it is already decoded, and it omits harness plumbing that would otherwise need stripping — but cross-checks the raw layer for calls that never produced a result, so placeholders read `3 calls: exec_command ×3 (+2 did not complete)`. Those are commands that failed to launch, that you aborted, or that were still running when the turn ended.
|
||
|
||
**Cloud tasks are out of scope.** `codex cloud` tasks run server-side and are reachable at `chatgpt.com/backend-api/api/codex/tasks`, but local CLI sessions are never uploaded there, so the cloud API is not an alternative source for these transcripts and this provider stays entirely offline. If you start using `codex cloud exec`, those transcripts would be cloud-only and would need separate work.
|
||
|
||
---
|
||
|
||
## Scheduling a Daily Run
|
||
|
||
`sync` chains `export` then `joplin` in one invocation, which is what a scheduler
|
||
wants — one command, and a meaningful exit code so a failed run is visible
|
||
instead of silent.
|
||
|
||
```bash
|
||
./ai-chat-exporter sync # every configured provider
|
||
./ai-chat-exporter sync --provider codex # just one
|
||
```
|
||
|
||
### Linux (systemd user timer)
|
||
|
||
```bash
|
||
./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
|
||
./scheduling/install-systemd-timer.sh --uninstall
|
||
```
|
||
|
||
Defaults to 09:00 daily (`--time 21:30` to change). `Persistent=true` means a
|
||
machine that was off at the scheduled time runs the archive at next boot rather
|
||
than skipping the day. To archive while logged out, `loginctl enable-linger $USER`.
|
||
|
||
Check on it with `systemctl --user list-timers aichat-sync.timer` and
|
||
`journalctl --user -u aichat-sync.service -n 50`.
|
||
|
||
### Windows (Task Scheduler)
|
||
|
||
```powershell
|
||
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
|
||
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
|
||
```
|
||
|
||
Per-user task, no admin rights needed. `-StartWhenAvailable` is the counterpart
|
||
of systemd's `Persistent=true`.
|
||
|
||
### What to know before you rely on it
|
||
|
||
**List only the providers that work unattended on that machine.** The local
|
||
providers (`claude-code`, `codex`) need no credentials and always work. The web
|
||
providers depend on a session token that expires and can only be refreshed by
|
||
hand via DevTools — so on a machine where that token is stale, scheduling them
|
||
means a failed run every single day, which is a good way to learn to ignore
|
||
failures you actually want to see. That is why `--provider` is repeatable in both
|
||
installers: schedule the coding machine for `claude-code` + `codex`, the browser
|
||
machine for `chatgpt` + `claude`.
|
||
|
||
**Both installers pass `--joplin-optional`.** If Joplin desktop isn't running,
|
||
the sync warns instead of failing: the export has already captured the local
|
||
transcripts (the part that can disappear), and the notes are rebuilt from the
|
||
cache on the next run that finds Joplin up.
|
||
|
||
**One provider failing does not skip the rest.** Both installers loop over the
|
||
providers in a single action rather than one action each, because systemd
|
||
`oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only
|
||
the last action's result. Every provider is attempted; the run still exits
|
||
non-zero if any failed.
|
||
|
||
**Acknowledge the ToS notice once, interactively.** It's stored in the cache
|
||
manifest per machine. Until then a scheduled run exits 1 with an explanation
|
||
rather than hanging on a prompt no one can answer.
|
||
|
||
---
|
||
|
||
## Output Structure
|
||
|
||
All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks.
|
||
|
||
### Default: `provider/project/year`
|
||
|
||
```
|
||
exports/
|
||
├── chatgpt/
|
||
│ ├── no-project/
|
||
│ │ └── 2024/
|
||
│ │ └── 2024-03-15_my-conversation_abc12345.md
|
||
│ └── learning-python/
|
||
│ └── 2024/
|
||
│ └── 2024-03-15_async-tutorial_def67890.md
|
||
└── claude/
|
||
├── no-project/
|
||
│ └── 2024/
|
||
│ └── 2024-06-01_docker-explained_ghi11111.md
|
||
└── startos-packaging/
|
||
└── 2024/
|
||
└── 2024-06-10_manifest-setup_jkl22222.md
|
||
```
|
||
|
||
### Joplin Notebook Mapping
|
||
|
||
Each provider+project combination maps to a flat Joplin notebook created automatically by the `joplin` command:
|
||
|
||
| Export folder | Joplin notebook |
|
||
|---------------|-----------------|
|
||
| `exports/chatgpt/learning-python/` | `ChatGPT - Learning Python` |
|
||
| `exports/claude/startos-packaging/` | `Claude - Startos Packaging` |
|
||
| `exports/chatgpt/no-project/` | `ChatGPT - No Project` |
|
||
| `exports/claude/no-project/` | `Claude - No Project` |
|
||
|
||
### Other `OUTPUT_STRUCTURE` options
|
||
|
||
| Value | Result |
|
||
|-------|--------|
|
||
| `provider/project/year` (default) | `exports/claude/my-project/2024/file.md` |
|
||
| `provider/project` | `exports/claude/my-project/file.md` |
|
||
| `provider/year` | `exports/claude/2024/file.md` (projects ignored) |
|
||
|
||
### Filename format
|
||
|
||
`YYYY-MM-DD_{title-slug}_{id[:8]}.md` — e.g. `2024-06-10_manifest-setup_jkl22222.md`
|
||
|
||
---
|
||
|
||
## CLI Reference
|
||
|
||
### Global flags
|
||
|
||
```
|
||
--verbose / -v DEBUG output to console
|
||
--quiet / -q WARNING and above only
|
||
--debug DEBUG + full tracebacks + redacted API response bodies
|
||
--no-log-file Disable file logging
|
||
--version Print version and exit
|
||
```
|
||
|
||
### `auth` — Interactive token setup
|
||
|
||
```bash
|
||
ai-chat-exporter auth # manual wizard (DevTools flow)
|
||
```
|
||
|
||
Guided wizard to find and save session tokens and ChatGPT project IDs. Detects OS and shows the correct DevTools shortcut, then writes the values to `.env`.
|
||
|
||
### `doctor` — Health check
|
||
|
||
```bash
|
||
ai-chat-exporter doctor
|
||
```
|
||
|
||
Checks: token presence, JWT validity and expiry, directory permissions, disk space, live API reachability. Exits with code 0 if all pass, 1 if any fail.
|
||
|
||
### `export` — Export conversations
|
||
|
||
```bash
|
||
# Export everything (new/updated only)
|
||
ai-chat-exporter export
|
||
|
||
# Single provider
|
||
ai-chat-exporter export --provider claude
|
||
|
||
# JSON output
|
||
ai-chat-exporter export --format json
|
||
|
||
# Both Markdown and JSON
|
||
ai-chat-exporter export --format both
|
||
|
||
# Only conversations updated since a date
|
||
ai-chat-exporter export --since 2024-06-01
|
||
|
||
# Only conversations in a specific project (case-insensitive substring)
|
||
ai-chat-exporter export --project "learning python"
|
||
|
||
# Only conversations outside any project
|
||
ai-chat-exporter export --project none
|
||
|
||
# Write to a custom directory
|
||
ai-chat-exporter export --output /path/to/my/notes
|
||
|
||
# Preview without writing anything
|
||
ai-chat-exporter export --dry-run
|
||
```
|
||
|
||
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--hidden-content [full|placeholder|omit]`, `--download-media [images|all|off]`, `--max-conversations N`, `--force`, `--dry-run`
|
||
|
||
**Re-rendering the whole archive after an upgrade.** New formatting or features (collapse policy, media downloads) only change conversations as they're re-exported. To re-render everything you already have, use `--force` — it re-exports every conversation even if unchanged, **without** `cache --clear`, so your Joplin note links are preserved (a later `joplin` run updates the existing notes instead of duplicating them).
|
||
|
||
A force re-render runs as a tracked "campaign": each run re-renders the least-recently-exported conversations, the "still to go" count shrinks each run, and finished providers do no further work. The simplest approach is one uncapped pass:
|
||
|
||
```bash
|
||
ai-chat-exporter export --force # re-renders everything in one paced pass
|
||
ai-chat-exporter joplin # update notes + upload media
|
||
```
|
||
|
||
Or spread the load with `--max-conversations`, re-running until it reports "Force re-render complete":
|
||
|
||
```bash
|
||
ai-chat-exporter export --force --max-conversations 50 # repeat until complete
|
||
```
|
||
|
||
### `list` — List conversations
|
||
|
||
```bash
|
||
# List all conversations for all providers
|
||
ai-chat-exporter list
|
||
|
||
# Single provider
|
||
ai-chat-exporter list --provider chatgpt
|
||
|
||
# Filter by project
|
||
ai-chat-exporter list --project "learning python"
|
||
|
||
# Only conversations outside any project
|
||
ai-chat-exporter list --project none
|
||
```
|
||
|
||
Fetches and displays all conversations without exporting them. Useful for verifying what the tool can see before running an export.
|
||
|
||
### `joplin` — Sync to Joplin
|
||
|
||
```bash
|
||
# Sync all pending conversations to Joplin
|
||
ai-chat-exporter joplin
|
||
|
||
# Preview what would be synced without sending anything
|
||
ai-chat-exporter joplin --dry-run
|
||
|
||
# Sync a single provider
|
||
ai-chat-exporter joplin --provider chatgpt
|
||
|
||
# Sync only conversations in a specific project
|
||
ai-chat-exporter joplin --project "learning python"
|
||
|
||
# Sync only conversations outside any project
|
||
ai-chat-exporter joplin --project none
|
||
```
|
||
|
||
Reads the local export cache and pushes each exported Markdown file to Joplin as a note. Notebooks are created automatically. Re-running is safe — notes are updated (not duplicated).
|
||
|
||
**Prerequisites:**
|
||
1. Run `export` first to generate the Markdown files
|
||
2. Open Joplin → Tools → Options → Web Clipper → enable the service
|
||
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
|
||
4. Joplin desktop must be open when you run this command
|
||
|
||
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
|
||
|
||
### `prune` — Delete stale export files
|
||
|
||
```bash
|
||
# Preview what would be deleted
|
||
ai-chat-exporter prune --dry-run
|
||
|
||
# Delete (asks for confirmation; -y skips the prompt)
|
||
ai-chat-exporter prune
|
||
```
|
||
|
||
Deletes export files no longer referenced by the cache manifest — leftovers
|
||
from old folder layouts or fixed bugs that would otherwise double-sync into
|
||
Joplin. Refuses to run when the manifest is empty (e.g. right after
|
||
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
|
||
command separately verifies that every manifest entry's file exists on disk.
|
||
|
||
### `cache` — Manage the sync manifest
|
||
|
||
```bash
|
||
# Show statistics
|
||
ai-chat-exporter cache --show
|
||
|
||
# Clear all cached entries (forces full re-export next run)
|
||
ai-chat-exporter cache --clear
|
||
|
||
# Clear a single provider
|
||
ai-chat-exporter cache --clear --provider claude
|
||
```
|
||
|
||
---
|
||
|
||
## How the Cache Works
|
||
|
||
The cache manifest lives at `cache/manifest.json` (inside the install directory) and records every exported conversation: its title, project, `updated_at` timestamp, output file path, and (after Joplin sync) the Joplin note ID.
|
||
|
||
On every `export` run:
|
||
1. Fetch the full conversation list from the provider
|
||
2. Compare each conversation's `updated_at` against the manifest
|
||
3. Export only conversations that are new or have been updated
|
||
4. Write each successfully exported conversation to the manifest **immediately** (not batched)
|
||
|
||
On every `joplin` run:
|
||
1. Read the manifest to find conversations not yet synced to Joplin, or re-exported since last sync
|
||
2. Push each pending Markdown file to Joplin (create or update)
|
||
3. Store the Joplin note ID in the manifest so subsequent runs update rather than duplicate
|
||
|
||
**This design makes every run inherently resumable.** If the tool is interrupted for any reason — rate limit, network drop, Ctrl+C, crash — simply re-run the same command. It will skip already-processed conversations and continue from where it stopped.
|
||
|
||
To force a full re-export: `ai-chat-exporter cache --clear` then re-run export.
|
||
|
||
---
|
||
|
||
## Troubleshooting
|
||
|
||
### `401 Unauthorized`
|
||
Your session token has expired.
|
||
- Run `ai-chat-exporter auth` to get a new token interactively
|
||
- Or manually copy a fresh cookie value into your `.env` file
|
||
|
||
Note: Claude's `sessionKey` is an opaque string — the only way to know it's expired is the 401 error. ChatGPT JWTs have an `exp` claim that the `doctor` command can decode and display.
|
||
|
||
### `429 Rate Limited`
|
||
The tool automatically pauses, saves progress, and exits with a clear message showing how many conversations were exported vs remaining. Just re-run the same export command to resume — the cache picks up exactly where it left off.
|
||
|
||
### Joplin: "JOPLIN_API_TOKEN is not set"
|
||
You need to configure the token before running the `joplin` command:
|
||
1. Open Joplin desktop
|
||
2. Go to Tools → Options → Web Clipper
|
||
3. Enable the Web Clipper service
|
||
4. Copy the Authorization token shown on that page
|
||
5. Add `JOPLIN_API_TOKEN=<token>` to your `.env` file
|
||
|
||
### Joplin: "Joplin is not responding"
|
||
Joplin desktop must be running when you run the `joplin` command. The Web Clipper service shuts down when Joplin is closed.
|
||
|
||
### Joplin: "Joplin rejected the API token (HTTP 401)"
|
||
The token in `JOPLIN_API_TOKEN` doesn't match what Joplin expects. Get a fresh token from Joplin → Tools → Options → Web Clipper → Authorization token.
|
||
|
||
### Joplin: note timed out
|
||
If you see a timeout error, Joplin took longer than `JOPLIN_REQUEST_TIMEOUT` seconds (default: 30) to respond. Possible causes:
|
||
- The conversation is very large and Joplin is slow to index it
|
||
- Joplin is busy syncing or loading a large library
|
||
- Joplin has frozen — try restarting it
|
||
|
||
To increase the timeout: add `JOPLIN_REQUEST_TIMEOUT=60` to your `.env`.
|
||
|
||
### ChatGPT project conversations not appearing
|
||
Make sure you've added the project IDs to `CHATGPT_PROJECT_IDS` in your `.env`. See [ChatGPT Projects](#chatgpt-projects) for how to find them. Project conversations are not included in the default conversation listing — they must be fetched separately.
|
||
|
||
### Schema warnings in logs (`Unexpected API response shape`)
|
||
The provider's internal API may have changed. Run with `--debug`, sanitize the output (remove any personal content), and check the project's GitHub Issues for known fixes.
|
||
|
||
### Non-text content warnings
|
||
Since v0.4.0, rich content is preserved as typed blocks in the export. ChatGPT voice transcripts render as text and audio assets as `📎 File attached` placeholders with size and duration metadata. Anything the extractor doesn't recognise renders as a visible `> ⚠️ Unsupported content` block naming the type and observed keys, *and* increments a counter in the post-export summary so you can tell whether real content is being silently skipped.
|
||
|
||
### Embedded images and media
|
||
Since v0.6.0, image attachments are downloaded by default into a `media/` folder beside each export and inlined as ``. Audio and other files download only with `EXPORTER_DOWNLOAD_MEDIA=all`. On the next `joplin` run these are uploaded as Joplin resources so they render inside the note. Assets that have expired on the provider's side (older AI-generated images often do) keep their text placeholder and are tallied as `media failed` in the run summary — they're never fatal to the export. Set `EXPORTER_DOWNLOAD_MEDIA=off` to skip downloads entirely.
|
||
|
||
### Tool output / hidden context collapsed
|
||
Since v0.6.0, content that was invisible in the ChatGPT web UI is collapsed by default to one-line placeholders like `> 🔧 Tool output — file_search (15.0 KB) — omitted`. This is ChatGPT's file-retrieval tool re-injecting your attached files on every run — it can be 90% of a conversation's bytes while containing none of the dialogue. Custom Instructions collapse the same way (`> ℹ️ Hidden context`). The post-export summary lists every collapsed message by origin with the total KB omitted. To keep everything, set `EXPORTER_HIDDEN_CONTENT=full` in `.env` or pass `--hidden-content full`.
|
||
|
||
### Empty export / all conversations skipped
|
||
No new or updated conversations since your last run. To verify: `ai-chat-exporter cache --show`. To force a full re-export: `ai-chat-exporter cache --clear`.
|
||
|
||
### Filing a bug report
|
||
1. Run with `--debug`: `ai-chat-exporter export --debug 2>&1 | tee debug.log`
|
||
2. Remove any personal conversation content from `debug.log`
|
||
3. Open a GitHub Issue with the sanitized log and the exact command you ran
|
||
|
||
---
|
||
|
||
## Future Work
|
||
|
||
See `FUTURE.md` for the full roadmap. Current priorities:
|
||
|
||
- **Watch/scheduled mode** on the way to a headless StartOS service
|
||
|
||
---
|
||
|
||
## Security Notes
|
||
|
||
- All exported data is stored **locally only** — nothing is sent anywhere except to your local Joplin instance
|
||
- Exported files and the cache manifest are created with `600` permissions (owner read/write only)
|
||
- `.env` is in `.gitignore` — **never commit it**
|
||
- Session tokens are never logged, printed, or included in error messages
|
||
- The Joplin API token is only ever sent to `localhost` — it never leaves your machine
|
||
- If you accidentally commit `.env`: immediately log out and back in to invalidate the token, then remove it from git history using [BFG Repo Cleaner](https://rtyley.github.io/bfg-repo-cleaner/) or `git filter-branch`
|