15 Commits
Author SHA1 Message Date
Jesse.MarkowitzandClaude Sonnet 4.6 19bfdaecbe fix: v0.2.1 — chunked ChatGPT cookies and Claude project path
- Support __Secure-next-auth.session-token.0/.1 split cookies; ChatGPT
  now issues tokens that exceed the 4KB per-cookie limit and must be
  sent as two named chunks or the auth endpoint returns no accessToken.
  Add CHATGPT_SESSION_TOKEN_1 env var; update auth wizard instructions.

- Fix Claude conversations exported to wrong directory when project name
  is present in the listing but absent from the detail endpoint response.
  Explicitly propagate "project" alongside _-prefixed annotation keys.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 22:32:14 -04:00
Jesse.MarkowitzandClaude Opus 4.6 4ccd918eb1 fix: list command shows Claude titles and fits 80-col terminals
Claude's list endpoint returns conversations with a `name` field rather
than `title`, so every Claude row was falling through to "Untitled".
Also set no_wrap + ellipsis overflow and tune column widths so the table
renders one row per conversation in Windows Command Prompt (80 cols).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-08 14:49:58 -04:00
Jesse.Markowitz a869e8c7ba fix for project files written to wrong directory 2026-03-30 15:25:18 -04:00
Jesse.Markowitz 340293ab94 fix for project files not extracted 2026-03-30 13:22:05 -04:00
Jesse.Markowitz 050cd49124 updated to run on Windows and add est capabilities 2026-03-30 11:08:05 -04:00
Jesse.MarkowitzandClaude Sonnet 4.6 304cf4fde4 feat: v0.2.0 — Joplin import, ChatGPT Projects, --project filter
Core features:
- Add `joplin` command: syncs exported Markdown to Joplin via local REST API
- Notebooks auto-created per provider+project (e.g. "ChatGPT - My Project")
- Idempotent: notes updated (not duplicated) on re-run; note ID tracked in manifest
- Add `--project` filter to `export` and `list` commands (substring or 'none')
- Add ChatGPT Projects support via CHATGPT_PROJECT_IDS env var

Config:
- Add JOPLIN_API_TOKEN, JOPLIN_API_URL, JOPLIN_REQUEST_TIMEOUT
- Version now read from importlib.metadata (single source of truth: pyproject.toml)
- Bump version to 0.2.0

Quality:
- Explicit Timeout handling in JoplinClient with actionable error messages
- token validation (validate_token) separate from connectivity (ping)
- Remove debug_auth.py, debug_claude.py, and untracked .har file
- Add *.har to .gitignore (may contain auth cookies/session tokens)
- Update README, CHANGELOG, FUTURE.md to reflect v0.2.0

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-01 06:04:03 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 23d7c17255 fix: use curl_cffi Chrome TLS impersonation for Claude provider
claude.ai has the same Cloudflare TLS fingerprinting protection as
chatgpt.com. Apply the same fix: curl_cffi impersonate=chrome120,
remove base class User-Agent to avoid JA3/UA mismatch.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-28 05:34:35 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 51c806c2c6 fix: set Claude sessionKey in cookie jar instead of raw header
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-28 05:33:04 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 f500c038cb fix: remove User-Agent override to prevent TLS/UA fingerprint mismatch
curl_cffi sets a User-Agent consistent with its JA3 TLS fingerprint.
BaseProvider's custom UA (Chrome/121) conflicted with the chrome120
TLS fingerprint, causing Cloudflare to flag the request as a bot.
Removing the UA from session headers lets curl_cffi manage its own.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-28 05:27:53 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 bb92ed2731 fix: update debug_auth.py to check accessToken presence by key
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-28 05:24:18 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 5c6dcafa34 fix: use curl_cffi Chrome TLS impersonation to bypass Cloudflare
chatgpt.com uses Cloudflare's TLS fingerprinting (JA3/JA4) which
blocks Python requests regardless of cookies. curl_cffi impersonates
Chrome's exact TLS handshake, making requests indistinguishable from
a real browser at the transport layer.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-28 05:20:52 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 d236fdb21a fix: set session cookie in cookie jar instead of manual header
Using self._session.cookies.set() ensures the cookie is sent correctly
by the requests session on all calls, including /api/auth/session.
Also add sec-fetch-* headers required by chatgpt.com.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 23:39:38 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 6a33de682a fix: implement two-step ChatGPT auth (session cookie → access token)
The __Secure-next-auth.session-token cannot be used directly as a Bearer
token. It must first be exchanged via GET /api/auth/session (with the token
sent as a Cookie) to obtain a short-lived accessToken. This accessToken is
then used as the Authorization: Bearer header for all backend-api calls.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 23:37:42 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 b41634d892 fix: load .env in doctor command; handle JWE tokens gracefully
Doctor was reading env vars before loading .env, so tokens set in .env
were invisible. ChatGPT now uses JWE (encrypted JWT) tokens which
PyJWT cannot decode without the server key — treat decode failure as
"token set, expiry unknown" rather than a FAIL.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 23:34:54 -05:00
Jesse.MarkowitzandClaude Sonnet 4.6 8e9ca36b57 docs: add README
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 23:14:37 -05:00
20 changed files with 2803 additions and 126 deletions
+27 -5
View File
@@ -6,9 +6,19 @@
# --- ChatGPT --- # --- ChatGPT ---
# How to get: open chatgpt.com in Chrome → F12 → Application tab # How to get: open chatgpt.com in Chrome → F12 → Application tab
# → Cookies → https://chatgpt.com → find "__Secure-next-auth.session-token" → copy Value # → Cookies → https://chatgpt.com → find the two cookie chunks:
# Token type: JWT (starts with "eyJ"). Typically valid for ~7 days. # __Secure-next-auth.session-token.0 (starts with "eyJ") → CHATGPT_SESSION_TOKEN
# __Secure-next-auth.session-token.1 (the remainder) → CHATGPT_SESSION_TOKEN_1
# Token type: JWE. Typically valid for ~7 days.
CHATGPT_SESSION_TOKEN= CHATGPT_SESSION_TOKEN=
CHATGPT_SESSION_TOKEN_1=
# ChatGPT Projects (optional): comma-separated list of project gizmo IDs.
# Project conversations are NOT included in the default /conversations listing.
# How to find: open chatgpt.com → click a Project → look at the browser URL:
# https://chatgpt.com/g/g-p-<ID>-<slug>/project → copy "g-p-<ID>"
# Example: CHATGPT_PROJECT_IDS=g-p-68c2b2b3037c8191890036fb4ae3ed9f,g-p-anotherproject
CHATGPT_PROJECT_IDS=
# --- Claude --- # --- Claude ---
# How to get: open claude.ai in Chrome → F12 → Application tab # How to get: open claude.ai in Chrome → F12 → Application tab
@@ -26,10 +36,22 @@ EXPORT_DIR=./exports
# provider/year → exports/claude/2024/file.md (ignores projects) # provider/year → exports/claude/2024/file.md (ignores projects)
OUTPUT_STRUCTURE=provider/project/year OUTPUT_STRUCTURE=provider/project/year
# --- Joplin ---
# Automate importing exported conversations into Joplin as notes.
# Requires Joplin desktop running with the Web Clipper service enabled.
# How to get the token:
# Joplin → Tools → Options → Web Clipper → copy "Authorization token"
JOPLIN_API_TOKEN=
# API URL (default port is 41184; change only if you've customised it)
JOPLIN_API_URL=http://localhost:41184
# Request timeout in seconds (default: 30). Increase if Joplin times out on
# large conversations. Example: JOPLIN_REQUEST_TIMEOUT=60
# JOPLIN_REQUEST_TIMEOUT=30
# --- Cache --- # --- Cache ---
# Where the sync manifest and logs are stored (default: ~/.ai-chat-exporter) # Where the sync manifest is stored (default: ./cache, inside the install directory)
CACHE_DIR=~/.ai-chat-exporter CACHE_DIR=./cache
# --- Logging --- # --- Logging ---
# Log file path. Set to "none" to disable file logging. # Log file path. Set to "none" to disable file logging.
LOG_FILE=~/.ai-chat-exporter/logs/exporter.log LOG_FILE=./cache/logs/exporter.log
+7
View File
@@ -25,10 +25,14 @@ exports/
!CHANGELOG.md !CHANGELOG.md
# Cache and logs # Cache and logs
cache/
.ai-chat-exporter/ .ai-chat-exporter/
logs/ logs/
*.log *.log
# Test tracking
test-plan.csv
# Editor / OS # Editor / OS
.DS_Store .DS_Store
.idea/ .idea/
@@ -36,3 +40,6 @@ logs/
*.swp *.swp
*.swo *.swo
Thumbs.db Thumbs.db
# HTTP traffic captures — may contain auth cookies and session tokens
*.har
+10
View File
@@ -3,6 +3,16 @@
All notable changes to this project will be documented here. All notable changes to this project will be documented here.
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
## [0.2.0] - Unreleased
### Added
- Joplin import automation: `joplin` command syncs exported Markdown files to Joplin as notes
- Notebooks created automatically per provider+project (`ChatGPT - My Project`, etc.)
- Re-running is safe: notes are updated, not duplicated (Joplin note ID stored in manifest)
- `JOPLIN_API_TOKEN`, `JOPLIN_API_URL`, `JOPLIN_REQUEST_TIMEOUT` config variables
- Configurable request timeout with clear error messages and actionable hints on timeout
- `--project` filter on `export` and `list` commands (case-insensitive substring or `none`)
- ChatGPT Projects support via `CHATGPT_PROJECT_IDS` env var
## [0.1.0] - Unreleased ## [0.1.0] - Unreleased
### Added ### Added
- Initial implementation: ChatGPT and Claude export via internal web APIs - Initial implementation: ChatGPT and Claude export via internal web APIs
+116 -25
View File
@@ -1,9 +1,17 @@
# Planned Future Work # Planned Future Work
These items are explicitly out of scope for v0.1.0 but have been designed for. Items completed in each release are moved to the changelog. Items here are
The codebase is structured to make each of these additions straightforward. designed for but not yet implemented. The codebase is structured to make each
of these additions straightforward.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
---
## Export `--force` Flag (v0.2.x)
## Export --force Flag (v0.1.x)
Add `--force` to the `export` command to re-export already-cached conversations Add `--force` to the `export` command to re-export already-cached conversations
without permanently clearing the entire manifest. Useful for re-generating files without permanently clearing the entire manifest. Useful for re-generating files
after changing the Markdown template or output structure. after changing the Markdown template or output structure.
@@ -13,30 +21,27 @@ returns all conversations regardless of cache state when force is True.
Current workaround: `python -m src.main cache --clear` then re-run export. Current workaround: `python -m src.main cache --clear` then re-run export.
## Joplin Integration (v0.2.0) ## Joplin `--force` Flag (v0.2.x)
Automate importing exported Markdown files into Joplin as new notes.
Joplin exposes a local REST API (requires Joplin desktop running with Web Clipper enabled).
Approach: after export, iterate exported files and POST each to Similarly, add `--force` to the `joplin` command to re-sync all cached
`http://localhost:41184/notes` with the appropriate notebook ID. conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
The output folder structure maps directly to Joplin notebooks: Implementation: in `get_joplin_pending()`, return all entries that have a
- exports/chatgpt/my-project/ → Joplin notebook "ChatGPT - My Project" `file_path` when `force=True`, ignoring `joplin_synced_at`.
- exports/claude/my-project/ → Joplin notebook "Claude - My Project"
- exports/chatgpt/no-project/ → Joplin notebook "ChatGPT - No Project"
- exports/claude/no-project/ → Joplin notebook "Claude - No Project"
Prerequisites: ## Per-Conversation Cache Reset (v0.2.x)
- Joplin desktop must be running with Web Clipper enabled
- `JOPLIN_API_TOKEN` env var (get from Joplin → Tools → Web Clipper Options)
- The Joplin import script will need to create notebooks if they don't exist,
then POST each note into the correct notebook
Note: The default OUTPUT_STRUCTURE of provider/project/year is assumed when Add `cache --reset --conversation <id>` to force re-export or re-sync of a
implementing the import script. If the user has changed OUTPUT_STRUCTURE, single conversation without clearing the entire provider cache.
the import script will need updating accordingly.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
---
## Official API Fallback (v0.3.0)
## Official API Migration (v0.3.0)
If the unofficial internal web API approach breaks, migrate to official export If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback: file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data - ChatGPT: parse `conversations.json` from Settings → Export Data
@@ -44,14 +49,17 @@ file parsing as a fallback:
The `BaseProvider` abstract class is intentionally designed so that a The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface `FileProvider` subclass can implement the same interface
(list_conversations, get_conversation, normalize_conversation) (`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code. without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the `src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON. export command to accept a pre-downloaded export ZIP or JSON.
---
## Rich Content Support (v0.4.0) ## Rich Content Support (v0.4.0)
Currently only text content is exported. Future versions should handle: Currently only text content is exported. Future versions should handle:
### Claude ### Claude
@@ -68,5 +76,88 @@ Currently only text content is exported. Future versions should handle:
Implementation note: the normalized message schema already includes a Implementation note: the normalized message schema already includes a
`content_type` field placeholder. When this work begins, extend the schema `content_type` field placeholder. When this work begins, extend the schema
rather than replacing it. In v0.1.0, log a WARNING whenever non-text content rather than replacing it. Non-text content already logs a WARNING when
is encountered so users know what was skipped. encountered so users can see what was skipped.
---
## Scheduled / Watch Mode (v0.5.0)
Add a `watch` command (or cron integration helper) to run exports automatically
on a schedule:
```bash
python -m src.main watch --interval 6h # poll every 6 hours
```
This would run `export` + `joplin` in sequence, then sleep. Alternatively,
provide a `cron` command that prints the correct crontab line for the user's
setup.
Implementation: simple loop with `time.sleep()`, or emit a crontab entry
string that calls the export and joplin commands in sequence. A `--once`
flag would do a single run then exit (useful for cron itself).
---
## Obsidian Vault Output (v0.5.0)
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
---
## Joplin Nested Notebooks (future)
Currently notebooks are flat: `ChatGPT - My Project`. Joplin supports nested
notebooks via `parent_id`. A future option (`JOPLIN_NESTED_NOTEBOOKS=true`)
could create a two-level hierarchy:
```
ChatGPT/
My Project/
No Project/
Claude/
Budget Tracker/
```
Implementation: `get_or_create_notebook` would first find/create the provider
notebook, then find/create the project notebook as a child.
---
## Token Expiry Notifications (future)
Proactively warn when a token is close to expiry (within 48h for ChatGPT),
rather than only surfacing the warning at startup. Options:
- Add an `expiry` subcommand that prints token status and exits non-zero if
any token is expired or expiring soon (useful in scripts/cron)
- Send a desktop notification via `notify-send` (Linux) or `osascript` (macOS)
when a token is within 24h of expiry
---
## Search Command (future)
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
+454
View File
@@ -0,0 +1,454 @@
# AI Chat Exporter
A personal backup tool for ChatGPT and Claude conversation history. Exports your chats to Markdown files and syncs them to [Joplin](https://joplinapp.org/) as notes. Each conversation becomes a single `.md` file with YAML frontmatter, organised into folders that map directly to Joplin notebooks.
Supports incremental sync — only new or updated conversations are exported on each run. Every run is resumable: if interrupted, re-running picks up exactly where it left off.
---
## ⚠️ Terms of Service Warning
**Read this before using this tool.**
This tool works by accessing **unofficial, undocumented internal web API endpoints** used by the ChatGPT and Claude web apps. These endpoints are not publicly supported by OpenAI or Anthropic and are subject to change or removal without notice.
**Use of this tool may conflict with their Terms of Service:**
- OpenAI: https://openai.com/policies/terms-of-use
- Anthropic: https://www.anthropic.com/legal/consumer-terms
**By using this tool, you accept that:**
- You are using it entirely at your own risk
- Your account could potentially be suspended for automated or scripted access
- The internal APIs this tool relies on may break at any time without notice
- This tool is for **personal archival use only** — not commercial use
This tool is designed for a single user backing up their own conversations. Do not use it to scrape data at scale or for any commercial purpose.
---
## Installation
### Linux / macOS
```bash
git clone <repo-url>
cd ai-chat-exporter
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
```
### Windows
No admin access required. Run these in **Command Prompt** (`cmd.exe`) — it's the simplest option on Windows because it doesn't have PowerShell's script execution policy restrictions.
```bat
git clone <repo-url>
cd ai-chat-exporter
python -m venv .venv
.venv\Scripts\activate
pip install -e ".[dev]"
```
All `ai-chat-exporter` commands work identically in Command Prompt.
**Using PowerShell instead?** If you prefer PowerShell, you may need to allow script execution first (one-time, current user only):
```powershell
Set-ExecutionPolicy RemoteSigned -Scope CurrentUser
```
Then activate the venv and run commands the same way.
**Prerequisites:**
- Python 3.11 or later — install from [python.org](https://www.python.org/downloads/windows/). During installation, tick **"Add Python to PATH"**.
- Git — install from [git-scm.com](https://git-scm.com/) if not already present.
**Notes:**
- The cache manifest and logs are stored in `cache\` inside the install directory — the same as on Linux.
- File permission hardening (`chmod 600`) is silently ignored on Windows — not a concern for single-user desktop use.
- Joplin Web Clipper runs on `localhost:41184` on all platforms; no configuration changes needed.
---
## First Run: Run Doctor
Before anything else, validate your setup:
```bash
ai-chat-exporter doctor
```
This checks token presence, format, expiry, directory permissions, disk space, and live API connectivity. Fix any failures before proceeding.
---
## Getting Your Session Tokens
Session tokens are how your browser stays logged in. This tool uses them to access your chat history on your behalf.
### Token Lifetimes
| Provider | Cookie Name | Lifetime | Expiry Detection |
|----------|-------------|----------|-----------------|
| ChatGPT | `__Secure-next-auth.session-token` | ~7 days | JWT `exp` claim (decoded automatically) |
| Claude | `sessionKey` | ~30 days | Only detectable via 401 response |
### Finding Tokens in Chrome DevTools
1. Open the provider's website and make sure you're logged in
2. Press **F12** (Windows/Linux) or **Cmd+Option+I** (macOS) to open DevTools
3. Click the **Application** tab
4. In the left panel, expand **Cookies** and click the site URL
5. Find the cookie by name and copy its **Value**
**ChatGPT:** go to `https://chatgpt.com` → find `__Secure-next-auth.session-token` → copy Value (starts with `eyJ`)
**Claude:** go to `https://claude.ai` → find `sessionKey` → copy Value
### When Tokens Expire
When a token expires you'll see a `401 Unauthorized` error. To refresh:
- Re-run the `auth` wizard: `ai-chat-exporter auth`
- Or manually update the value in your `.env` file
---
## The `auth` Command
The easiest way to configure tokens is the interactive wizard:
```bash
ai-chat-exporter auth
```
This walks you through finding your token, validates it, shows the expiry date (ChatGPT only), and offers to write it to your `.env` automatically. Tokens are never echoed to the terminal.
---
## `.env` Setup
Copy `.env.example` to `.env` and fill in your values:
```bash
cp .env.example .env
```
### Provider tokens
| Variable | Description |
|----------|-------------|
| `CHATGPT_SESSION_TOKEN` | Your ChatGPT JWT session token (`eyJ…`) |
| `CHATGPT_PROJECT_IDS` | Comma-separated ChatGPT project IDs (see below) |
| `CLAUDE_SESSION_KEY` | Your Claude session key |
### Output
| Variable | Default | Description |
|----------|---------|-------------|
| `EXPORT_DIR` | `./exports` | Where to write exported Markdown files |
| `OUTPUT_STRUCTURE` | `provider/project/year` | Folder structure (see below) |
### Joplin
| Variable | Default | Description |
|----------|---------|-------------|
| `JOPLIN_API_TOKEN` | — | Authorization token from Joplin Web Clipper settings |
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
### Cache & logging
| Variable | Default | Description |
|----------|---------|-------------|
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
---
## ChatGPT Projects
ChatGPT project conversations are stored separately from your main conversation list and require extra configuration.
### Finding your project IDs
1. Open ChatGPT and click a Project in the left sidebar
2. Look at the browser URL — it will look like:
`https://chatgpt.com/g/g-p-68c2b2b3037c8191890036fb4ae3ed9f-my-project/project`
3. Copy the `g-p-…` part (everything up to but not including the slug after the second `-`)
Add all your project IDs to `.env` as a comma-separated list:
```
CHATGPT_PROJECT_IDS=g-p-68c2b2b3037c8191890036fb4ae3ed9f,g-p-anotherprojectid
```
The `auth` wizard can also guide you through this step interactively.
---
## Output Structure
All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks.
### Default: `provider/project/year`
```
exports/
├── chatgpt/
│ ├── no-project/
│ │ └── 2024/
│ │ └── 2024-03-15_my-conversation_abc12345.md
│ └── learning-python/
│ └── 2024/
│ └── 2024-03-15_async-tutorial_def67890.md
└── claude/
├── no-project/
│ └── 2024/
│ └── 2024-06-01_docker-explained_ghi11111.md
└── startos-packaging/
└── 2024/
└── 2024-06-10_manifest-setup_jkl22222.md
```
### Joplin Notebook Mapping
Each provider+project combination maps to a flat Joplin notebook created automatically by the `joplin` command:
| Export folder | Joplin notebook |
|---------------|-----------------|
| `exports/chatgpt/learning-python/` | `ChatGPT - Learning Python` |
| `exports/claude/startos-packaging/` | `Claude - Startos Packaging` |
| `exports/chatgpt/no-project/` | `ChatGPT - No Project` |
| `exports/claude/no-project/` | `Claude - No Project` |
### Other `OUTPUT_STRUCTURE` options
| Value | Result |
|-------|--------|
| `provider/project/year` (default) | `exports/claude/my-project/2024/file.md` |
| `provider/project` | `exports/claude/my-project/file.md` |
| `provider/year` | `exports/claude/2024/file.md` (projects ignored) |
### Filename format
`YYYY-MM-DD_{title-slug}_{id[:8]}.md` — e.g. `2024-06-10_manifest-setup_jkl22222.md`
---
## CLI Reference
### Global flags
```
--verbose / -v DEBUG output to console
--quiet / -q WARNING and above only
--debug DEBUG + full tracebacks + redacted API response bodies
--no-log-file Disable file logging
--version Print version and exit
```
### `auth` — Interactive token setup
```bash
ai-chat-exporter auth
```
Guided wizard to find and save session tokens and ChatGPT project IDs. Detects OS and shows the correct DevTools shortcut.
### `doctor` — Health check
```bash
ai-chat-exporter doctor
```
Checks: token presence, JWT validity and expiry, directory permissions, disk space, live API reachability. Exits with code 0 if all pass, 1 if any fail.
### `export` — Export conversations
```bash
# Export everything (new/updated only)
ai-chat-exporter export
# Single provider
ai-chat-exporter export --provider claude
# JSON output
ai-chat-exporter export --format json
# Both Markdown and JSON
ai-chat-exporter export --format both
# Only conversations updated since a date
ai-chat-exporter export --since 2024-06-01
# Only conversations in a specific project (case-insensitive substring)
ai-chat-exporter export --project "learning python"
# Only conversations outside any project
ai-chat-exporter export --project none
# Write to a custom directory
ai-chat-exporter export --output /path/to/my/notes
# Preview without writing anything
ai-chat-exporter export --dry-run
```
Options: `--provider [chatgpt|claude|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--dry-run`
### `list` — List conversations
```bash
# List all conversations for all providers
ai-chat-exporter list
# Single provider
ai-chat-exporter list --provider chatgpt
# Filter by project
ai-chat-exporter list --project "learning python"
# Only conversations outside any project
ai-chat-exporter list --project none
```
Fetches and displays all conversations without exporting them. Useful for verifying what the tool can see before running an export.
### `joplin` — Sync to Joplin
```bash
# Sync all pending conversations to Joplin
ai-chat-exporter joplin
# Preview what would be synced without sending anything
ai-chat-exporter joplin --dry-run
# Sync a single provider
ai-chat-exporter joplin --provider chatgpt
# Sync only conversations in a specific project
ai-chat-exporter joplin --project "learning python"
# Sync only conversations outside any project
ai-chat-exporter joplin --project none
```
Reads the local export cache and pushes each exported Markdown file to Joplin as a note. Notebooks are created automatically. Re-running is safe — notes are updated (not duplicated).
**Prerequisites:**
1. Run `export` first to generate the Markdown files
2. Open Joplin → Tools → Options → Web Clipper → enable the service
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
4. Joplin desktop must be open when you run this command
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
### `cache` — Manage the sync manifest
```bash
# Show statistics
ai-chat-exporter cache --show
# Clear all cached entries (forces full re-export next run)
ai-chat-exporter cache --clear
# Clear a single provider
ai-chat-exporter cache --clear --provider claude
```
---
## How the Cache Works
The cache manifest lives at `cache/manifest.json` (inside the install directory) and records every exported conversation: its title, project, `updated_at` timestamp, output file path, and (after Joplin sync) the Joplin note ID.
On every `export` run:
1. Fetch the full conversation list from the provider
2. Compare each conversation's `updated_at` against the manifest
3. Export only conversations that are new or have been updated
4. Write each successfully exported conversation to the manifest **immediately** (not batched)
On every `joplin` run:
1. Read the manifest to find conversations not yet synced to Joplin, or re-exported since last sync
2. Push each pending Markdown file to Joplin (create or update)
3. Store the Joplin note ID in the manifest so subsequent runs update rather than duplicate
**This design makes every run inherently resumable.** If the tool is interrupted for any reason — rate limit, network drop, Ctrl+C, crash — simply re-run the same command. It will skip already-processed conversations and continue from where it stopped.
To force a full re-export: `ai-chat-exporter cache --clear` then re-run export.
---
## Troubleshooting
### `401 Unauthorized`
Your session token has expired.
- Run `ai-chat-exporter auth` to get a new token interactively
- Or manually copy a fresh cookie value into your `.env` file
Note: Claude's `sessionKey` is an opaque string — the only way to know it's expired is the 401 error. ChatGPT JWTs have an `exp` claim that the `doctor` command can decode and display.
### `429 Rate Limited`
The tool automatically pauses, saves progress, and exits with a clear message showing how many conversations were exported vs remaining. Just re-run the same export command to resume — the cache picks up exactly where it left off.
### Joplin: "JOPLIN_API_TOKEN is not set"
You need to configure the token before running the `joplin` command:
1. Open Joplin desktop
2. Go to Tools → Options → Web Clipper
3. Enable the Web Clipper service
4. Copy the Authorization token shown on that page
5. Add `JOPLIN_API_TOKEN=<token>` to your `.env` file
### Joplin: "Joplin is not responding"
Joplin desktop must be running when you run the `joplin` command. The Web Clipper service shuts down when Joplin is closed.
### Joplin: "Joplin rejected the API token (HTTP 401)"
The token in `JOPLIN_API_TOKEN` doesn't match what Joplin expects. Get a fresh token from Joplin → Tools → Options → Web Clipper → Authorization token.
### Joplin: note timed out
If you see a timeout error, Joplin took longer than `JOPLIN_REQUEST_TIMEOUT` seconds (default: 30) to respond. Possible causes:
- The conversation is very large and Joplin is slow to index it
- Joplin is busy syncing or loading a large library
- Joplin has frozen — try restarting it
To increase the timeout: add `JOPLIN_REQUEST_TIMEOUT=60` to your `.env`.
### ChatGPT project conversations not appearing
Make sure you've added the project IDs to `CHATGPT_PROJECT_IDS` in your `.env`. See [ChatGPT Projects](#chatgpt-projects) for how to find them. Project conversations are not included in the default conversation listing — they must be fetched separately.
### Schema warnings in logs (`Unexpected API response shape`)
The provider's internal API may have changed. Run with `--debug`, sanitize the output (remove any personal content), and check the project's GitHub Issues for known fixes.
### Non-text content warnings
Images, code interpreter outputs, DALL-E generations, and Claude artifacts are not exported in v0.2.0. A WARNING is logged for each skipped item. See `FUTURE.md` for the roadmap.
### Empty export / all conversations skipped
No new or updated conversations since your last run. To verify: `ai-chat-exporter cache --show`. To force a full re-export: `ai-chat-exporter cache --clear`.
### Filing a bug report
1. Run with `--debug`: `ai-chat-exporter export --debug 2>&1 | tee debug.log`
2. Remove any personal conversation content from `debug.log`
3. Open a GitHub Issue with the sanitized log and the exact command you ran
---
## Future Work
See `FUTURE.md` for planned features:
- **v0.2.x** — `export --force` flag; `joplin --force` flag; per-conversation cache reset
- **v0.3.0** — Official API fallback: parse export ZIP files from ChatGPT/Claude settings
- **v0.4.0** — Rich content: images, artifacts, code interpreter output, extended thinking
- **v0.5.0** — Watch/scheduled mode; Obsidian vault output
---
## Security Notes
- All exported data is stored **locally only** — nothing is sent anywhere except to your local Joplin instance
- Exported files and the cache manifest are created with `600` permissions (owner read/write only)
- `.env` is in `.gitignore`**never commit it**
- Session tokens are never logged, printed, or included in error messages
- The Joplin API token is only ever sent to `localhost` — it never leaves your machine
- If you accidentally commit `.env`: immediately log out and back in to invalidate the token, then remove it from git history using [BFG Repo Cleaner](https://rtyley.github.io/bfg-repo-cleaner/) or `git filter-branch`
+2 -1
View File
@@ -4,11 +4,12 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "ai-chat-exporter" name = "ai-chat-exporter"
version = "0.1.0" version = "0.2.1"
description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin" description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin"
requires-python = ">=3.11" requires-python = ">=3.11"
dependencies = [ dependencies = [
"requests==2.31.0", "requests==2.31.0",
"curl_cffi==0.14.0",
"click==8.1.7", "click==8.1.7",
"python-dotenv==1.0.1", "python-dotenv==1.0.1",
"rich==13.7.1", "rich==13.7.1",
+3
View File
@@ -1,14 +1,17 @@
# Editable Git install with no remote (ai-chat-exporter==0.1.0) # Editable Git install with no remote (ai-chat-exporter==0.1.0)
-e /home/jesse/services/ai-chatexport -e /home/jesse/services/ai-chatexport
certifi==2026.2.25 certifi==2026.2.25
cffi==2.0.0
charset-normalizer==3.4.4 charset-normalizer==3.4.4
click==8.1.7 click==8.1.7
curl_cffi==0.14.0
idna==3.11 idna==3.11
iniconfig==2.3.0 iniconfig==2.3.0
markdown-it-py==4.0.0 markdown-it-py==4.0.0
mdurl==0.1.2 mdurl==0.1.2
packaging==26.0 packaging==26.0
pluggy==1.6.0 pluggy==1.6.0
pycparser==3.0
Pygments==2.19.2 Pygments==2.19.2
PyJWT==2.8.0 PyJWT==2.8.0
pytest==8.1.1 pytest==8.1.1
+64 -5
View File
@@ -1,4 +1,4 @@
"""Local cache manifest for tracking exported conversations.""" """Local cache manifest for tracking exported and Joplin-synced conversations."""
import json import json
import logging import logging
@@ -18,11 +18,17 @@ class CacheError(Exception):
class Cache: class Cache:
"""Manages the local JSON manifest of exported conversations. """Manages the local JSON manifest of exported and Joplin-synced conversations.
The manifest is the single source of truth for what has been exported. The manifest is the single source of truth for what has been exported and
Every run compares the provider's full conversation list against this synced. Every export run compares the provider's full conversation list
manifest to determine what is new or updated. against this manifest to determine what is new or updated. The Joplin sync
run reads it to find conversations not yet pushed to Joplin (or re-exported
since the last sync).
Each entry tracks:
title, project, updated_at, exported_at, file_path,
joplin_note_id (after first sync), joplin_synced_at (after first sync)
File security: File security:
- Permissions: 600 (owner read/write only) - Permissions: 600 (owner read/write only)
@@ -150,6 +156,59 @@ class Cache:
"""Return all cached entries for a provider (for --cache --show).""" """Return all cached entries for a provider (for --cache --show)."""
return dict(self._data.get(provider, {})) return dict(self._data.get(provider, {}))
def mark_joplin_synced(self, provider: str, conv_id: str, note_id: str) -> None:
"""Record a successful Joplin sync for a conversation.
Adds ``joplin_note_id`` and ``joplin_synced_at`` to the manifest entry
and writes atomically to disk.
"""
entry = self._data.get(provider, {}).get(conv_id)
if entry is None:
logger.warning(
"[cache] mark_joplin_synced: no cache entry for %s/%s", provider, conv_id[:8]
)
return
entry["joplin_note_id"] = note_id
entry["joplin_synced_at"] = datetime.now(tz=timezone.utc).isoformat()
self._save()
def get_joplin_pending(self, provider: str) -> list[tuple[str, dict]]:
"""Return (conv_id, entry) pairs that need to be synced to Joplin.
A conversation is pending when:
- It has never been synced (no ``joplin_note_id``), OR
- It was re-exported after the last Joplin sync
(``exported_at`` > ``joplin_synced_at``).
Returns:
List of (conv_id, entry_dict) tuples, where entry_dict includes
``file_path``, ``title``, ``project``, and optionally ``joplin_note_id``.
"""
pending = []
for conv_id, entry in self._data.get(provider, {}).items():
if not isinstance(entry, dict):
continue
if not entry.get("file_path"):
continue
note_id = entry.get("joplin_note_id")
if not note_id:
pending.append((conv_id, entry))
continue
# Re-sync if the file was re-exported after the last Joplin sync
exported_at = entry.get("exported_at", "")
synced_at = entry.get("joplin_synced_at", "")
if exported_at and synced_at:
try:
from src.utils import _parse_dt
if _parse_dt(exported_at) > _parse_dt(synced_at):
pending.append((conv_id, entry))
except Exception:
pass
return pending
def last_run(self) -> str | None: def last_run(self) -> str | None:
"""Return the ISO8601 timestamp of the last export run, or None.""" """Return the ISO8601 timestamp of the last export run, or None."""
return self._data.get("last_run") return self._data.get("last_run")
+46 -6
View File
@@ -28,6 +28,7 @@ class ConfigError(Exception):
@dataclass @dataclass
class Config: class Config:
chatgpt_session_token: str | None chatgpt_session_token: str | None
chatgpt_session_token_1: str | None
claude_session_key: str | None claude_session_key: str | None
export_dir: Path export_dir: Path
output_structure: str output_structure: str
@@ -35,6 +36,13 @@ class Config:
log_file: str log_file: str
# Decoded ChatGPT JWT expiry (None if token absent or not a JWT) # Decoded ChatGPT JWT expiry (None if token absent or not a JWT)
chatgpt_token_expiry: datetime | None = field(default=None, repr=False) chatgpt_token_expiry: datetime | None = field(default=None, repr=False)
# ChatGPT Project gizmo IDs (g-p-xxx) — project conversations are not
# included in the default /conversations listing; they must be fetched
# separately via /backend-api/gizmos/{id}/conversations.
chatgpt_project_ids: list[str] = field(default_factory=list)
# Joplin local REST API settings (Web Clipper service)
joplin_api_token: str | None = None
joplin_api_url: str = "http://localhost:41184"
def load_config() -> Config: def load_config() -> Config:
@@ -48,11 +56,30 @@ def load_config() -> Config:
load_dotenv(override=False) load_dotenv(override=False)
chatgpt_token = os.getenv("CHATGPT_SESSION_TOKEN", "").strip() or None chatgpt_token = os.getenv("CHATGPT_SESSION_TOKEN", "").strip() or None
chatgpt_token_1 = os.getenv("CHATGPT_SESSION_TOKEN_1", "").strip() or None
claude_key = os.getenv("CLAUDE_SESSION_KEY", "").strip() or None claude_key = os.getenv("CLAUDE_SESSION_KEY", "").strip() or None
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser() export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
output_structure = os.getenv("OUTPUT_STRUCTURE", "provider/project/year").strip() output_structure = os.getenv("OUTPUT_STRUCTURE", "provider/project/year").strip()
cache_dir = Path(os.getenv("CACHE_DIR", "~/.ai-chat-exporter")).expanduser() cache_dir = Path(os.getenv("CACHE_DIR", "./cache")).expanduser()
log_file = os.getenv("LOG_FILE", "~/.ai-chat-exporter/logs/exporter.log").strip() log_file = os.getenv("LOG_FILE", "./cache/logs/exporter.log").strip()
# Joplin
joplin_token = os.getenv("JOPLIN_API_TOKEN", "").strip() or None
joplin_url = os.getenv("JOPLIN_API_URL", "http://localhost:41184").strip()
# Parse CHATGPT_PROJECT_IDS — comma-separated list of gizmo IDs (g-p-xxx)
_project_ids_raw = os.getenv("CHATGPT_PROJECT_IDS", "").strip()
chatgpt_project_ids = [
pid.strip()
for pid in _project_ids_raw.split(",")
if pid.strip() and pid.strip().startswith("g-p-")
] if _project_ids_raw else []
if _project_ids_raw and not chatgpt_project_ids:
logger.warning(
"CHATGPT_PROJECT_IDS is set but contains no valid project IDs. "
"Each ID should start with 'g-p-' (e.g. g-p-68c2b2b3037c8191890036fb4ae3ed9f). "
"Find your project ID in the browser URL when viewing a project."
)
errors: list[str] = [] errors: list[str] = []
@@ -76,7 +103,7 @@ def load_config() -> Config:
if not chatgpt_token and not claude_key: if not chatgpt_token and not claude_key:
logger.warning( logger.warning(
"Neither CHATGPT_SESSION_TOKEN nor CLAUDE_SESSION_KEY is set. " "Neither CHATGPT_SESSION_TOKEN nor CLAUDE_SESSION_KEY is set. "
"Run 'python -m src.main auth' to configure credentials." "Run 'ai-chat-exporter auth' to configure credentials."
) )
# Create and validate output directory # Create and validate output directory
@@ -102,12 +129,16 @@ def load_config() -> Config:
config = Config( config = Config(
chatgpt_session_token=chatgpt_token, chatgpt_session_token=chatgpt_token,
chatgpt_session_token_1=chatgpt_token_1,
claude_session_key=claude_key, claude_session_key=claude_key,
export_dir=export_dir, export_dir=export_dir,
output_structure=output_structure, output_structure=output_structure,
cache_dir=cache_dir, cache_dir=cache_dir,
log_file=log_file, log_file=log_file,
chatgpt_token_expiry=chatgpt_expiry, chatgpt_token_expiry=chatgpt_expiry,
chatgpt_project_ids=chatgpt_project_ids,
joplin_api_token=joplin_token,
joplin_api_url=joplin_url,
) )
_log_startup_summary(config) _log_startup_summary(config)
@@ -125,8 +156,12 @@ def _validate_chatgpt_token(token: str) -> datetime | None:
try: try:
payload = jwt.decode(token, options={"verify_signature": False}) payload = jwt.decode(token, options={"verify_signature": False})
except jwt.DecodeError as e: except jwt.DecodeError:
logger.warning("CHATGPT_SESSION_TOKEN could not be decoded as JWT: %s", e) # JWE (encrypted JWT, alg=dir) — cannot be decoded without the server key.
# This is normal for ChatGPT's current token format. Token is valid; expiry unknown.
logger.debug(
"CHATGPT_SESSION_TOKEN is an encrypted JWE token — expiry cannot be decoded client-side."
)
return None return None
exp = payload.get("exp") exp = payload.get("exp")
@@ -141,7 +176,7 @@ def _validate_chatgpt_token(token: str) -> datetime | None:
if delta.total_seconds() < 0: if delta.total_seconds() < 0:
logger.warning( logger.warning(
"CHATGPT_SESSION_TOKEN expired at %s. " "CHATGPT_SESSION_TOKEN expired at %s. "
"Run 'python -m src.main auth' to refresh it.", "Run 'ai-chat-exporter auth' to refresh it.",
expiry.strftime("%Y-%m-%d %H:%M UTC"), expiry.strftime("%Y-%m-%d %H:%M UTC"),
) )
elif delta.total_seconds() < 86400: elif delta.total_seconds() < 86400:
@@ -178,16 +213,21 @@ def _log_startup_summary(cfg: Config) -> None:
"""Log a single INFO line summarising the active configuration.""" """Log a single INFO line summarising the active configuration."""
chatgpt_status = format_token_status(cfg.chatgpt_session_token, cfg.chatgpt_token_expiry) chatgpt_status = format_token_status(cfg.chatgpt_session_token, cfg.chatgpt_token_expiry)
claude_status = format_token_status(cfg.claude_session_key) claude_status = format_token_status(cfg.claude_session_key)
joplin_status = "configured" if cfg.joplin_api_token else "not configured"
logger.info( logger.info(
"Config loaded | " "Config loaded | "
"ChatGPT: %s | " "ChatGPT: %s | "
"Claude: %s | " "Claude: %s | "
"chatgpt_projects: %d | "
"Joplin: %s | "
"export_dir=%s | " "export_dir=%s | "
"structure=%s | " "structure=%s | "
"cache_dir=%s", "cache_dir=%s",
chatgpt_status, chatgpt_status,
claude_status, claude_status,
len(cfg.chatgpt_project_ids),
joplin_status,
cfg.export_dir, cfg.export_dir,
cfg.output_structure, cfg.output_structure,
cfg.cache_dir, cfg.cache_dir,
+303
View File
@@ -0,0 +1,303 @@
"""Joplin Data API client for importing notes into Joplin desktop."""
import logging
import os
from typing import Any
import requests
logger = logging.getLogger(__name__)
# HTTP timeout for regular API calls (seconds). Notes can be large Markdown
# files so we allow more time than a typical JSON API call.
# Override with JOPLIN_REQUEST_TIMEOUT env var if you have very large conversations.
_REQUEST_TIMEOUT: int = int(os.getenv("JOPLIN_REQUEST_TIMEOUT", "30"))
class JoplinError(Exception):
"""Raised when the Joplin API returns an error or is unreachable."""
class JoplinClient:
"""HTTP client for the Joplin local REST API (Web Clipper service).
Requires Joplin desktop to be running with the Web Clipper service enabled.
Get your API token from: Joplin → Tools → Options → Web Clipper.
Args:
base_url: Joplin API base URL (default: http://localhost:41184).
token: API authorization token from Joplin Web Clipper settings.
"""
def __init__(self, base_url: str, token: str) -> None:
self._base_url = base_url.rstrip("/")
self._token = token
# In-memory cache of notebook title → ID to avoid repeated GET /folders
self._notebook_cache: dict[str, str] = {}
self._notebooks_loaded = False
logger.debug("[joplin] Client initialised with base_url=%s", self._base_url)
# ------------------------------------------------------------------
# Connectivity
# ------------------------------------------------------------------
def ping(self) -> bool:
"""Return True if the Joplin API is reachable and responding.
Note: /ping does not require authentication. A successful ping only
confirms Joplin is running — not that the token is valid. Call
``validate_token()`` to confirm authentication separately.
Raises:
JoplinError: If the API returns an unexpected non-connection error.
"""
url = f"{self._base_url}/ping"
logger.debug("[joplin] GET %s", url)
try:
resp = requests.get(url, timeout=5)
resp.raise_for_status()
ok = "JoplinClipperServer" in resp.text
logger.debug("[joplin] ping → %s (body: %r)", "OK" if ok else "unexpected response", resp.text[:80])
return ok
except requests.exceptions.ConnectionError:
logger.debug("[joplin] ping → connection refused at %s", url)
return False
except requests.exceptions.Timeout:
logger.debug("[joplin] ping → timed out after 5s at %s", url)
return False
except requests.exceptions.RequestException as e:
raise JoplinError(f"Joplin ping failed: {e}") from e
def validate_token(self) -> None:
"""Verify the API token is accepted by Joplin.
Does a minimal authenticated call (GET /folders?limit=1) and raises
``JoplinError`` if authentication fails.
Raises:
JoplinError: If the token is rejected (401) or Joplin is unreachable.
"""
logger.debug("[joplin] Validating API token…")
self._get("/folders", params={"limit": 1, "fields": "id"})
logger.debug("[joplin] Token validated OK")
# ------------------------------------------------------------------
# Notebooks (folders)
# ------------------------------------------------------------------
def list_notebooks(self) -> list[dict]:
"""Return all Joplin notebooks (folders), handling pagination.
Returns:
List of folder dicts with at least ``id`` and ``title`` keys.
"""
results: list[dict] = []
page = 1
while True:
logger.debug("[joplin] GET /folders page=%d", page)
resp = self._get("/folders", params={"page": page, "fields": "id,title"})
items = resp.get("items", [])
results.extend(items)
logger.debug("[joplin] /folders page=%d%d items, has_more=%s", page, len(items), resp.get("has_more"))
if not resp.get("has_more"):
break
page += 1
return results
def get_or_create_notebook(self, title: str) -> str:
"""Return the Joplin folder ID for ``title``, creating it if needed.
Args:
title: Notebook display name (e.g. "ChatGPT - My Project").
Returns:
Joplin folder ID string.
"""
if not self._notebooks_loaded:
self._load_notebook_cache()
if title in self._notebook_cache:
folder_id = self._notebook_cache[title]
logger.debug("[joplin] Notebook cache hit: %r%s", title, folder_id)
return folder_id
# Not found — create it
logger.info("[joplin] Creating notebook: %r", title)
resp = self._post("/folders", {"title": title})
folder_id = resp["id"]
self._notebook_cache[title] = folder_id
logger.debug("[joplin] Notebook created: %r%s", title, folder_id)
return folder_id
# ------------------------------------------------------------------
# Notes
# ------------------------------------------------------------------
def create_note(self, title: str, body: str, parent_id: str) -> str:
"""Create a new note in the specified notebook.
Args:
title: Note title.
body: Note body (Markdown).
parent_id: Notebook (folder) ID.
Returns:
ID of the created note.
"""
logger.debug(
"[joplin] Creating note: %r in notebook %s (%d chars)",
title, parent_id, len(body),
)
resp = self._post("/notes", {"title": title, "body": body, "parent_id": parent_id})
note_id = resp["id"]
logger.info("[joplin] Note created: %r%s", title, note_id)
return note_id
def update_note(self, note_id: str, title: str, body: str) -> None:
"""Update the title and body of an existing note.
Args:
note_id: Joplin note ID.
title: New note title.
body: New note body (Markdown).
"""
logger.debug(
"[joplin] Updating note %s: %r (%d chars)",
note_id, title, len(body),
)
self._put(f"/notes/{note_id}", {"title": title, "body": body})
logger.info("[joplin] Note updated: %r (%s)", title, note_id)
# ------------------------------------------------------------------
# HTTP helpers
# ------------------------------------------------------------------
def _get(self, path: str, params: dict | None = None) -> dict[str, Any]:
url = f"{self._base_url}{path}"
query = {"token": self._token, **(params or {})}
logger.debug("[joplin] GET %s params=%s", path, {k: v for k, v in (params or {}).items()})
try:
resp = requests.get(url, params=query, timeout=_REQUEST_TIMEOUT)
logger.debug("[joplin] GET %s → HTTP %d", path, resp.status_code)
resp.raise_for_status()
return resp.json()
except requests.exceptions.ConnectionError as e:
raise JoplinError(
"Cannot connect to Joplin. Is Joplin desktop running with Web Clipper enabled?"
) from e
except requests.exceptions.Timeout as e:
raise JoplinError(_timeout_message("GET", path)) from e
except requests.exceptions.HTTPError as e:
raise JoplinError(_http_error_message("GET", path, e)) from e
except requests.exceptions.RequestException as e:
raise JoplinError(f"Joplin GET {path} failed: {e}") from e
def _post(self, path: str, data: dict) -> dict[str, Any]:
url = f"{self._base_url}{path}"
logger.debug("[joplin] POST %s", path)
try:
resp = requests.post(url, params={"token": self._token}, json=data, timeout=_REQUEST_TIMEOUT)
logger.debug("[joplin] POST %s → HTTP %d", path, resp.status_code)
resp.raise_for_status()
return resp.json()
except requests.exceptions.ConnectionError as e:
raise JoplinError(
"Cannot connect to Joplin. Is Joplin desktop running with Web Clipper enabled?"
) from e
except requests.exceptions.Timeout as e:
raise JoplinError(_timeout_message("POST", path)) from e
except requests.exceptions.HTTPError as e:
raise JoplinError(_http_error_message("POST", path, e)) from e
except requests.exceptions.RequestException as e:
raise JoplinError(f"Joplin POST {path} failed: {e}") from e
def _put(self, path: str, data: dict) -> dict[str, Any]:
url = f"{self._base_url}{path}"
logger.debug("[joplin] PUT %s", path)
try:
resp = requests.put(url, params={"token": self._token}, json=data, timeout=_REQUEST_TIMEOUT)
logger.debug("[joplin] PUT %s → HTTP %d", path, resp.status_code)
resp.raise_for_status()
return resp.json()
except requests.exceptions.ConnectionError as e:
raise JoplinError(
"Cannot connect to Joplin. Is Joplin desktop running with Web Clipper enabled?"
) from e
except requests.exceptions.Timeout as e:
raise JoplinError(_timeout_message("PUT", path)) from e
except requests.exceptions.HTTPError as e:
raise JoplinError(_http_error_message("PUT", path, e)) from e
except requests.exceptions.RequestException as e:
raise JoplinError(f"Joplin PUT {path} failed: {e}") from e
def _load_notebook_cache(self) -> None:
logger.debug("[joplin] Loading notebook list from Joplin…")
notebooks = self.list_notebooks()
self._notebook_cache = {nb["title"]: nb["id"] for nb in notebooks}
self._notebooks_loaded = True
logger.debug("[joplin] Notebook cache loaded: %d notebooks", len(self._notebook_cache))
for title, folder_id in self._notebook_cache.items():
logger.debug("[joplin] %r%s", title, folder_id)
# ------------------------------------------------------------------
# Error message helper
# ------------------------------------------------------------------
def _timeout_message(method: str, path: str) -> str:
"""Build a clear timeout error message with actionable suggestions."""
return (
f"Joplin {method} {path} timed out after {_REQUEST_TIMEOUT}s. "
"Possible causes:\n"
" • The note body is very large and Joplin is slow to process it.\n"
" • Joplin is busy (syncing, indexing, or loading a large library).\n"
" • Joplin has frozen — try restarting it.\n"
f"If this happens repeatedly, increase JOPLIN_REQUEST_TIMEOUT in your .env "
f"(currently {_REQUEST_TIMEOUT}s)."
)
def _http_error_message(method: str, path: str, e: requests.exceptions.HTTPError) -> str:
"""Build a human-friendly error message from an HTTP error, with auth hint on 401."""
resp = e.response
status = resp.status_code if resp is not None else "?"
if status == 401:
return (
f"Joplin rejected the API token (HTTP 401 on {method} {path}). "
"Check that JOPLIN_API_TOKEN is correct: "
"Joplin → Tools → Options → Web Clipper → Authorization token."
)
if status == 404:
return f"Joplin resource not found (HTTP 404 on {method} {path}). The note may have been deleted in Joplin."
body_snippet = ""
if resp is not None:
try:
body_snippet = f"{resp.text[:120]}"
except Exception:
pass
return f"Joplin {method} {path} failed: HTTP {status}{body_snippet}"
# ------------------------------------------------------------------
# Notebook naming helper
# ------------------------------------------------------------------
_PROVIDER_DISPLAY = {
"chatgpt": "ChatGPT",
"claude": "Claude",
}
def notebook_title(provider: str, project: str | None) -> str:
"""Derive a flat Joplin notebook title from provider and project name.
Examples:
notebook_title("chatgpt", "no-project") → "ChatGPT - No Project"
notebook_title("claude", "budget-tracker") → "Claude - Budget Tracker"
notebook_title("chatgpt", None) → "ChatGPT - No Project"
"""
prov_display = _PROVIDER_DISPLAY.get(provider, provider.capitalize())
proj = (project or "no-project").replace("-", " ").title()
return f"{prov_display} - {proj}"
+445 -28
View File
@@ -1,5 +1,6 @@
"""CLI entry point for ai-chat-exporter.""" """CLI entry point for ai-chat-exporter."""
import importlib.metadata
import logging import logging
import platform import platform
import shutil import shutil
@@ -19,6 +20,7 @@ from src.providers.base import ProviderError
console = Console() console = Console()
err_console = Console(stderr=True) err_console = Console(stderr=True)
logger = logging.getLogger(__name__)
TOS_NOTICE = """\ TOS_NOTICE = """\
⚠️ IMPORTANT — TERMS OF SERVICE NOTICE ⚠️ IMPORTANT — TERMS OF SERVICE NOTICE
@@ -45,7 +47,10 @@ Type 'yes' to acknowledge and continue, or Ctrl+C to exit: \
@click.group() @click.group()
@click.version_option(version="0.1.0", prog_name="ai-chat-exporter") @click.version_option(
version=importlib.metadata.version("ai-chat-exporter"),
prog_name="ai-chat-exporter",
)
@click.option("--verbose", "-v", is_flag=True, help="Enable DEBUG output to console.") @click.option("--verbose", "-v", is_flag=True, help="Enable DEBUG output to console.")
@click.option("--quiet", "-q", is_flag=True, help="Show WARNING and above only.") @click.option("--quiet", "-q", is_flag=True, help="Show WARNING and above only.")
@click.option("--debug", is_flag=True, help="DEBUG + full tracebacks + redacted API bodies.") @click.option("--debug", is_flag=True, help="DEBUG + full tracebacks + redacted API bodies.")
@@ -65,7 +70,7 @@ def cli(ctx: click.Context, verbose: bool, quiet: bool, debug: bool, no_log_file
# Determine log file path from env (setup_logging handles "none") # Determine log file path from env (setup_logging handles "none")
import os import os
log_file = os.getenv("LOG_FILE", "~/.ai-chat-exporter/logs/exporter.log") log_file = os.getenv("LOG_FILE", "./cache/logs/exporter.log")
setup_logging(level=level, log_file=log_file, no_log_file=no_log_file) setup_logging(level=level, log_file=log_file, no_log_file=no_log_file)
@@ -74,7 +79,7 @@ def cli(ctx: click.Context, verbose: bool, quiet: bool, debug: bool, no_log_file
# Initialise cache (needed for ToS gate on every command) # Initialise cache (needed for ToS gate on every command)
import os import os
cache_dir = Path(os.getenv("CACHE_DIR", "~/.ai-chat-exporter")).expanduser() cache_dir = Path(os.getenv("CACHE_DIR", "./cache")).expanduser()
try: try:
cache = Cache(cache_dir) cache = Cache(cache_dir)
except CacheError as e: except CacheError as e:
@@ -135,7 +140,7 @@ def auth(ctx: click.Context) -> None:
if configure_claude: if configure_claude:
_auth_claude(os_name) _auth_claude(os_name)
console.print("\n[green]Done! Run 'python -m src.main doctor' to verify your setup.[/green]") console.print("\n[green]Done! Run 'ai-chat-exporter doctor' to verify your setup.[/green]")
def _auth_chatgpt(os_name: str) -> None: def _auth_chatgpt(os_name: str) -> None:
@@ -148,15 +153,19 @@ def _auth_chatgpt(os_name: str) -> None:
else: else:
console.print("2. Press [bold]F12[/bold] to open DevTools → Application tab.") console.print("2. Press [bold]F12[/bold] to open DevTools → Application tab.")
console.print("3. Expand [bold]Cookies[/bold] → [bold]https://chatgpt.com[/bold]") console.print("3. Expand [bold]Cookies[/bold] → [bold]https://chatgpt.com[/bold]")
console.print("4. Find [bold]__Secure-next-auth.session-token[/bold] → copy the Value.") console.print("4. ChatGPT splits the session token across two cookies:")
console.print(" (Token starts with 'eyJ...' — it is a long JWT string)") console.print(" [bold]__Secure-next-auth.session-token.0[/bold] (starts with 'eyJ')")
console.print("5. Paste it below (input is hidden).\n") console.print(" [bold]__Secure-next-auth.session-token.1[/bold] (the remainder)")
console.print(" Copy each Value in turn and paste below.")
console.print(" (If you only see one cookie without a .0/.1 suffix, paste it for .0 and leave .1 blank.)\n")
token = click.prompt("ChatGPT session token", hide_input=True, default="", show_default=False).strip() token = click.prompt("ChatGPT session token (.0)", hide_input=True, default="", show_default=False).strip()
if not token: if not token:
console.print("[yellow]Skipped ChatGPT token.[/yellow]") console.print("[yellow]Skipped ChatGPT token.[/yellow]")
return return
token_1 = click.prompt("ChatGPT session token (.1, leave blank if absent)", hide_input=True, default="", show_default=False).strip() or None
# Validate # Validate
if not token.startswith("eyJ"): if not token.startswith("eyJ"):
console.print("[yellow]Warning: token doesn't look like a JWT (expected 'eyJ...').[/yellow]") console.print("[yellow]Warning: token doesn't look like a JWT (expected 'eyJ...').[/yellow]")
@@ -173,7 +182,61 @@ def _auth_chatgpt(os_name: str) -> None:
except Exception: except Exception:
console.print("[yellow]Could not decode token expiry.[/yellow]") console.print("[yellow]Could not decode token expiry.[/yellow]")
# Live validation — exchange session token for an access token
_valid = False
_error: str | None = None
with console.status("[dim]Validating token with ChatGPT API…[/dim]"):
try:
from src.providers.chatgpt import ChatGPTProvider
_prov = ChatGPTProvider(session_token=token, session_token_1=token_1)
_prov._fetch_access_token()
_valid = True
except ProviderError as e:
_error = str(e.original)
except Exception as e:
_error = str(e)
if _valid:
console.print("[green]✓ Token verified — connected to ChatGPT API.[/green]")
else:
console.print(f"[red]✗ Token validation failed: {_error}[/red]")
_write_token_to_env("CHATGPT_SESSION_TOKEN", token) _write_token_to_env("CHATGPT_SESSION_TOKEN", token)
if token_1:
_write_token_to_env("CHATGPT_SESSION_TOKEN_1", token_1)
# --- ChatGPT Projects ---
console.print("\n[bold]ChatGPT Projects (optional)[/bold]")
console.print(
"Project conversations are stored separately and are not included in the\n"
"default conversation listing. To export them, you need each project's ID.\n"
)
console.print("How to find a project ID:")
console.print(" 1. Open ChatGPT and click into a Project in the left sidebar.")
console.print(" 2. Look at the browser URL — it will look like:")
console.print(" [dim]https://chatgpt.com/g/[bold]g-p-68c2b2b3037c8191890036fb4ae3ed9f[/bold]-my-project/project[/dim]")
console.print(" 3. Copy the part starting with [bold]g-p-[/bold] up to (but not including) the slug.")
console.print(" Enter multiple IDs separated by commas. Leave blank to skip.\n")
project_ids_raw = click.prompt(
"ChatGPT project IDs (comma-separated, e.g. g-p-xxx,g-p-yyy)",
default="",
show_default=False,
).strip()
if project_ids_raw:
ids = [pid.strip() for pid in project_ids_raw.split(",") if pid.strip()]
valid = [pid for pid in ids if pid.startswith("g-p-")]
invalid = [pid for pid in ids if not pid.startswith("g-p-")]
if invalid:
console.print(f"[yellow]Warning: skipping IDs that don't start with 'g-p-': {invalid}[/yellow]")
if valid:
_write_token_to_env("CHATGPT_PROJECT_IDS", ",".join(valid))
console.print(f"[green]Saved {len(valid)} project ID(s).[/green]")
else:
console.print("[yellow]No valid project IDs — skipping.[/yellow]")
else:
console.print("[dim]Skipped project IDs.[/dim]")
def _auth_claude(os_name: str) -> None: def _auth_claude(os_name: str) -> None:
@@ -193,7 +256,25 @@ def _auth_claude(os_name: str) -> None:
console.print("[yellow]Skipped Claude token.[/yellow]") console.print("[yellow]Skipped Claude token.[/yellow]")
return return
console.print("[green]Claude session key saved.[/green]") # Live validation — fetch org ID (the first call any Claude operation makes)
_valid = False
_error: str | None = None
with console.status("[dim]Validating token with Claude API…[/dim]"):
try:
from src.providers.claude import ClaudeProvider
_prov = ClaudeProvider(session_key=key)
_prov._get_org_id()
_valid = True
except ProviderError as e:
_error = str(e.original)
except Exception as e:
_error = str(e)
if _valid:
console.print("[green]✓ Token verified — connected to Claude API.[/green]")
else:
console.print(f"[red]✗ Token validation failed: {_error}[/red]")
_write_token_to_env("CLAUDE_SESSION_KEY", key) _write_token_to_env("CLAUDE_SESSION_KEY", key)
@@ -256,6 +337,10 @@ def _run_doctor_checks() -> list[dict]:
import os import os
import jwt as pyjwt import jwt as pyjwt
from datetime import timezone from datetime import timezone
from dotenv import load_dotenv
# Load .env so doctor works without the user having to export vars manually
load_dotenv(override=False)
checks = [] checks = []
@@ -288,8 +373,10 @@ def _run_doctor_checks() -> list[dict]:
add("ChatGPT token expiry warning", False, "Expires in < 24h — refresh soon") add("ChatGPT token expiry warning", False, "Expires in < 24h — refresh soon")
else: else:
add("ChatGPT token expiry", False, "JWT has no 'exp' claim") add("ChatGPT token expiry", False, "JWT has no 'exp' claim")
except Exception as e: except pyjwt.exceptions.DecodeError:
add("ChatGPT token decode", False, str(e)) # JWE (encrypted JWT) — cannot decode without the server key.
# This is normal for ChatGPT's current token format. Token is present and valid.
add("ChatGPT token expiry", True, "Encrypted token (JWE) — expiry not decodable client-side")
# Claude key # Claude key
if claude_key: if claude_key:
@@ -297,7 +384,7 @@ def _run_doctor_checks() -> list[dict]:
# Directories # Directories
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser() export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
cache_dir = Path(os.getenv("CACHE_DIR", "~/.ai-chat-exporter")).expanduser() cache_dir = Path(os.getenv("CACHE_DIR", "./cache")).expanduser()
for label, dirpath in [("Export dir writable", export_dir), ("Cache dir writable", cache_dir)]: for label, dirpath in [("Export dir writable", export_dir), ("Cache dir writable", cache_dir)]:
try: try:
@@ -321,7 +408,8 @@ def _run_doctor_checks() -> list[dict]:
if chatgpt_token: if chatgpt_token:
try: try:
from src.providers.chatgpt import ChatGPTProvider from src.providers.chatgpt import ChatGPTProvider
p = ChatGPTProvider(chatgpt_token) chatgpt_token_1 = os.getenv("CHATGPT_SESSION_TOKEN_1", "").strip() or None
p = ChatGPTProvider(chatgpt_token, session_token_1=chatgpt_token_1)
results = p.list_conversations(offset=0, limit=1) results = p.list_conversations(offset=0, limit=1)
add("ChatGPT API reachable", True, f"Got {len(results)} result(s)") add("ChatGPT API reachable", True, f"Got {len(results)} result(s)")
except ProviderError as e: except ProviderError as e:
@@ -389,6 +477,15 @@ def _print_doctor_table(checks: list[dict]) -> None:
default=None, default=None,
help="Only export conversations updated after this date (YYYY-MM-DD).", help="Only export conversations updated after this date (YYYY-MM-DD).",
) )
@click.option(
"--project",
"project_filter",
default=None,
help=(
"Only export conversations in a matching project (case-insensitive substring). "
"Use 'none' for conversations outside any project."
),
)
@click.option("--dry-run", is_flag=True, help="Show what would be exported without writing anything.") @click.option("--dry-run", is_flag=True, help="Show what would be exported without writing anything.")
@click.pass_context @click.pass_context
def export( def export(
@@ -397,6 +494,7 @@ def export(
fmt: str, fmt: str,
output_dir: str | None, output_dir: str | None,
since: str | None, since: str | None,
project_filter: str | None,
dry_run: bool, dry_run: bool,
) -> None: ) -> None:
"""Export new and updated conversations to Markdown or JSON. """Export new and updated conversations to Markdown or JSON.
@@ -442,7 +540,7 @@ def export(
providers_to_run = _resolve_providers(provider, cfg) providers_to_run = _resolve_providers(provider, cfg)
if not providers_to_run: if not providers_to_run:
err_console.print( err_console.print(
"[red]No providers configured. Run 'python -m src.main auth' to set up tokens.[/red]" "[red]No providers configured. Run 'ai-chat-exporter auth' to set up tokens.[/red]"
) )
sys.exit(1) sys.exit(1)
@@ -468,6 +566,12 @@ def export(
summary[prov_name]["failed"] += len(all_convs) if "all_convs" in dir() else 0 summary[prov_name]["failed"] += len(all_convs) if "all_convs" in dir() else 0
continue continue
if project_filter is not None:
all_convs = _filter_by_project(all_convs, project_filter)
console.print(
f" [dim]--project filter '{project_filter}': {len(all_convs)} matching conversations.[/dim]"
)
to_export = cache.get_new_or_updated(prov_name, all_convs) to_export = cache.get_new_or_updated(prov_name, all_convs)
skipped = len(all_convs) - len(to_export) skipped = len(all_convs) - len(to_export)
summary[prov_name]["skipped"] = skipped summary[prov_name]["skipped"] = skipped
@@ -497,6 +601,16 @@ def export(
conv_id = raw_conv.get("id") or raw_conv.get("uuid", "unknown") conv_id = raw_conv.get("id") or raw_conv.get("uuid", "unknown")
try: try:
full_raw = prov_instance.get_conversation(conv_id) full_raw = prov_instance.get_conversation(conv_id)
# Propagate metadata from the listing summary into the full
# detail so normalize_conversation can use it.
# - Keys starting with "_" are provider annotations
# (e.g. _project_name injected by ChatGPT project fetching).
# - "project" is included explicitly because Claude's detail
# endpoint omits it even though the listing returns it.
_PROPAGATE_KEYS = {"project"}
for key, val in raw_conv.items():
if (key.startswith("_") or key in _PROPAGATE_KEYS) and key not in full_raw:
full_raw[key] = val
normalized = prov_instance.normalize_conversation(full_raw) normalized = prov_instance.normalize_conversation(full_raw)
exported_path: Path | None = None exported_path: Path | None = None
@@ -516,13 +630,11 @@ def export(
progress.advance(task) progress.advance(task)
except ProviderError as e: except ProviderError as e:
logger = logging.getLogger(__name__)
logger.error("Failed to export conversation %s: %s", conv_id[:8], e) logger.error("Failed to export conversation %s: %s", conv_id[:8], e)
summary[prov_name]["failed"] += 1 summary[prov_name]["failed"] += 1
progress.advance(task) progress.advance(task)
continue continue
except OSError as e: except OSError as e:
logger = logging.getLogger(__name__)
logger.error("File write failed for conversation %s: %s", conv_id[:8], e) logger.error("File write failed for conversation %s: %s", conv_id[:8], e)
summary[prov_name]["failed"] += 1 summary[prov_name]["failed"] += 1
progress.advance(task) progress.advance(task)
@@ -554,7 +666,22 @@ def _resolve_providers(provider: str, cfg) -> list[tuple[str, object]]:
from src.providers.claude import ClaudeProvider from src.providers.claude import ClaudeProvider
if provider in ("chatgpt", "all"): if provider in ("chatgpt", "all"):
try_add("chatgpt", cfg.chatgpt_session_token, ChatGPTProvider) if cfg.chatgpt_session_token:
try:
result.append((
"chatgpt",
ChatGPTProvider(
session_token=cfg.chatgpt_session_token,
session_token_1=cfg.chatgpt_session_token_1,
project_ids=cfg.chatgpt_project_ids,
),
))
except ProviderError as e:
logging.getLogger(__name__).warning(
"[chatgpt] Could not initialise provider: %s", e
)
elif provider == "chatgpt" or provider == "all":
logging.getLogger(__name__).warning("[chatgpt] Skipping — token not configured.")
if provider in ("claude", "all"): if provider in ("claude", "all"):
try_add("claude", cfg.claude_session_key, ClaudeProvider) try_add("claude", cfg.claude_session_key, ClaudeProvider)
@@ -590,6 +717,44 @@ def _print_dry_run_table(prov_name, to_export, prov_instance, export_base, struc
console.print(f" [dim]{skipped} conversations already cached (would be skipped).[/dim]") console.print(f" [dim]{skipped} conversations already cached (would be skipped).[/dim]")
def _raw_project_name(conv: dict) -> str | None:
"""Extract the project name from a raw conversation summary dict.
Handles both ChatGPT (annotated _project_name) and Claude (project dict).
"""
# ChatGPT: annotated during fetch_all_conversations
if "_project_name" in conv:
return conv["_project_name"] or None
# Claude: project is a dict with a 'name' key, or a plain string
project = conv.get("project")
if isinstance(project, dict):
return project.get("name") or None
if isinstance(project, str):
return project or None
return None
def _filter_by_project(convs: list[dict], project_filter: str) -> list[dict]:
"""Filter conversations by project name.
project_filter='none' → keep only conversations with no project.
Otherwise → case-insensitive substring match on the project name.
"""
want_none = project_filter.lower() == "none"
needle = project_filter.lower()
result = []
for conv in convs:
name = _raw_project_name(conv)
if want_none:
if name is None:
result.append(conv)
else:
if name and needle in name.lower():
result.append(conv)
return result
def _print_export_summary(summary: dict[str, dict[str, int]]) -> None: def _print_export_summary(summary: dict[str, dict[str, int]]) -> None:
table = Table(title="Export Summary") table = Table(title="Export Summary")
table.add_column("Provider", style="bold") table.add_column("Provider", style="bold")
@@ -620,8 +785,17 @@ def _print_export_summary(summary: dict[str, dict[str, int]]) -> None:
default="all", default="all",
show_default=True, show_default=True,
) )
@click.option(
"--project",
"project_filter",
default=None,
help=(
"Only list conversations in a matching project (case-insensitive substring). "
"Use 'none' for conversations outside any project."
),
)
@click.pass_context @click.pass_context
def list_conversations(ctx: click.Context, provider: str) -> None: def list_conversations(ctx: click.Context, provider: str, project_filter: str | None) -> None:
"""List conversations without exporting them.""" """List conversations without exporting them."""
debug = ctx.obj.get("debug", False) debug = ctx.obj.get("debug", False)
cfg = _load_config_or_exit(debug) cfg = _load_config_or_exit(debug)
@@ -635,20 +809,29 @@ def list_conversations(ctx: click.Context, provider: str) -> None:
_handle_provider_error(e, debug) _handle_provider_error(e, debug)
continue continue
table = Table() if project_filter is not None:
table.add_column("Title") all_convs = _filter_by_project(all_convs, project_filter)
table.add_column("Project")
table.add_column("Updated") # no_wrap + overflow="ellipsis" prevents Rich from wrapping cells to
table.add_column("ID") # multiple lines on narrow terminals (e.g. Windows Command Prompt),
# which can otherwise make the output look garbled. Widths are tuned
# to fit within an 80-column terminal.
# Total width budget for 80-column terminals:
# borders (5) + padding (4 cols * 2) = 13 chars of overhead
# remaining 67 chars split: 34 title + 15 project + 10 date + 8 id
table = Table(show_lines=False, expand=False, padding=(0, 1))
table.add_column("Title", no_wrap=True, overflow="ellipsis", max_width=34)
table.add_column("Project", no_wrap=True, overflow="ellipsis", max_width=15)
table.add_column("Updated", no_wrap=True, min_width=10)
table.add_column("ID", no_wrap=True, min_width=8)
for conv in all_convs: for conv in all_convs:
title = conv.get("title") or "Untitled" # ChatGPT uses "title"; Claude uses "name".
project = conv.get("project_title") or "" title = conv.get("title") or conv.get("name") or "Untitled"
if isinstance(conv.get("project"), dict): project = _raw_project_name(conv) or ""
project = conv["project"].get("name", "")
updated = (conv.get("updated_at") or conv.get("update_time") or "")[:10] updated = (conv.get("updated_at") or conv.get("update_time") or "")[:10]
conv_id = (conv.get("id") or conv.get("uuid") or "")[:8] conv_id = (conv.get("id") or conv.get("uuid") or "")[:8]
table.add_row(title[:60], project[:30], updated, conv_id) table.add_row(title, project, updated, conv_id)
console.print(table) console.print(table)
console.print(f"Total: {len(all_convs)} conversations") console.print(f"Total: {len(all_convs)} conversations")
@@ -694,6 +877,240 @@ def cache(ctx: click.Context, show: bool, clear: bool, provider: str) -> None:
console.print("Specify --show or --clear. Use --help for options.") console.print("Specify --show or --clear. Use --help for options.")
# ──────────────────────────────────────────────────────────────────────────────
# joplin command
# ──────────────────────────────────────────────────────────────────────────────
@cli.command()
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "all"], case_sensitive=False),
default="all",
show_default=True,
help="Which provider's conversations to sync to Joplin.",
)
@click.option(
"--project",
"project_filter",
default=None,
help=(
"Only sync conversations in a matching project (case-insensitive substring). "
"Use 'none' for conversations outside any project."
),
)
@click.option("--dry-run", is_flag=True, help="Show what would be synced without sending anything to Joplin.")
@click.pass_context
def joplin(ctx: click.Context, provider: str, project_filter: str | None, dry_run: bool) -> None:
"""Sync exported conversations to Joplin as notes.
Reads the local export cache and pushes exported Markdown files to Joplin
via its local REST API. Requires Joplin desktop to be running with the
Web Clipper service enabled.
Notebooks are created automatically based on provider and project:
exports/chatgpt/my-project/ → "ChatGPT - My Project" notebook
exports/claude/no-project/ → "Claude - No Project" notebook
Re-running is safe: notes are updated (not duplicated) on subsequent runs.
Setup:
1. Open Joplin desktop.
2. Go to Tools → Options → Web Clipper.
3. Enable the Web Clipper service.
4. Copy the Authorization token.
5. Set JOPLIN_API_TOKEN=<token> in your .env file.
"""
debug = ctx.obj.get("debug", False)
cache_obj: Cache = ctx.obj["cache"]
cfg = _load_config_or_exit(debug)
if not cfg.joplin_api_token:
err_console.print(
"[red]JOPLIN_API_TOKEN is not set.[/red]\n"
" 1. Open Joplin → Tools → Options → Web Clipper.\n"
" 2. Enable the Web Clipper service.\n"
" 3. Copy the Authorization token.\n"
" 4. Add [bold]JOPLIN_API_TOKEN=<token>[/bold] to your .env file."
)
sys.exit(1)
from src.joplin import JoplinClient, JoplinError, notebook_title
client = JoplinClient(cfg.joplin_api_url, cfg.joplin_api_token)
if not dry_run:
console.print(f"[dim]Connecting to Joplin at {cfg.joplin_api_url}…[/dim]")
try:
if not client.ping():
err_console.print(
"[red]Joplin is not responding.[/red] "
"Make sure Joplin desktop is open and Web Clipper is enabled."
)
sys.exit(1)
# Ping succeeded but doesn't validate the token — check auth separately
client.validate_token()
except JoplinError as e:
err_console.print(f"[red]Joplin connection error:[/red] {e}")
sys.exit(1)
console.print("[green]Joplin connected and token validated.[/green]")
# Determine which providers to process
providers_to_sync: list[str] = []
if provider in ("chatgpt", "all"):
providers_to_sync.append("chatgpt")
if provider in ("claude", "all"):
providers_to_sync.append("claude")
summary: dict[str, dict[str, int]] = {}
for prov_name in providers_to_sync:
summary[prov_name] = {"created": 0, "updated": 0, "skipped": 0, "failed": 0}
pending = cache_obj.get_joplin_pending(prov_name)
logger.debug("[joplin] %s: %d pending before filter", prov_name, len(pending))
# Apply --project filter against the cached entry's project field
if project_filter is not None:
want_none = project_filter.lower() == "none"
needle = project_filter.lower()
filtered = []
for conv_id, entry in pending:
proj = entry.get("project") or None
if want_none:
if proj is None or proj == "no-project":
filtered.append((conv_id, entry))
else:
if proj and needle in proj.lower():
filtered.append((conv_id, entry))
logger.debug(
"[joplin] %s: --project %r filtered %d%d",
prov_name, project_filter, len(pending), len(filtered),
)
pending = filtered
if not pending:
console.print(f"\n[bold cyan][{prov_name.upper()}][/bold cyan] All up to date — nothing to sync.")
continue
console.print(
f"\n[bold cyan][{prov_name.upper()}][/bold cyan] "
f"{len(pending)} conversation(s) to sync to Joplin."
)
if dry_run:
_print_joplin_dry_run_table(prov_name, pending)
continue
from rich.progress import Progress, SpinnerColumn, TextColumn, BarColumn, TaskProgressColumn
with Progress(
SpinnerColumn(),
TextColumn("[progress.description]{task.description}"),
BarColumn(),
TaskProgressColumn(),
console=console,
) as progress:
task = progress.add_task(f"Syncing {prov_name}", total=len(pending))
for conv_id, entry in pending:
file_path = entry.get("file_path", "")
title = entry.get("title") or "Untitled"
project = entry.get("project") or None
existing_note_id = entry.get("joplin_note_id")
action = "update" if existing_note_id else "create"
logger.debug(
"[joplin] %s %s/%s: %s (file=%s)",
action, prov_name, conv_id[:8], title[:60], file_path,
)
try:
# Read the exported Markdown file
body = Path(file_path).read_text(encoding="utf-8")
logger.debug("[joplin] Read %d chars from %s", len(body), file_path)
# Get or create the notebook
nb_title = notebook_title(prov_name, project)
notebook_id = client.get_or_create_notebook(nb_title)
if existing_note_id:
client.update_note(existing_note_id, title, body)
cache_obj.mark_joplin_synced(prov_name, conv_id, existing_note_id)
summary[prov_name]["updated"] += 1
else:
note_id = client.create_note(title, body, notebook_id)
cache_obj.mark_joplin_synced(prov_name, conv_id, note_id)
summary[prov_name]["created"] += 1
except FileNotFoundError:
logger.warning(
"[joplin] Skipping %s/%s — exported file not found: %s",
prov_name, conv_id[:8], file_path,
)
summary[prov_name]["skipped"] += 1
except JoplinError as e:
logger.error(
"[joplin] Failed to %s note for %s/%s: %s",
action, prov_name, conv_id[:8], e,
)
summary[prov_name]["failed"] += 1
except OSError as e:
logger.error(
"[joplin] File read error for %s/%s (%s): %s",
prov_name, conv_id[:8], file_path, e,
)
summary[prov_name]["failed"] += 1
finally:
progress.advance(task)
if not dry_run:
_print_joplin_summary(summary)
def _print_joplin_dry_run_table(prov_name: str, pending: list[tuple[str, dict]]) -> None:
from src.joplin import notebook_title
table = Table(title=f"[DRY RUN] {prov_name.upper()} — Would sync {len(pending)} conversation(s)")
table.add_column("Title")
table.add_column("Project")
table.add_column("Notebook")
table.add_column("Action")
for conv_id, entry in pending[:50]:
title = entry.get("title") or "Untitled"
project = entry.get("project") or "no-project"
nb = notebook_title(prov_name, entry.get("project"))
action = "update" if entry.get("joplin_note_id") else "create"
table.add_row(title[:50], project[:30], nb, action)
if len(pending) > 50:
table.add_row(f"… and {len(pending) - 50} more", "", "", "")
console.print(table)
def _print_joplin_summary(summary: dict[str, dict[str, int]]) -> None:
table = Table(title="Joplin Sync Summary")
table.add_column("Provider", style="bold")
table.add_column("Created", justify="right")
table.add_column("Updated", justify="right")
table.add_column("Skipped", justify="right")
table.add_column("Failed", justify="right")
for prov, counts in summary.items():
table.add_row(
prov.capitalize(),
str(counts["created"]),
str(counts["updated"]),
str(counts["skipped"]),
f"[red]{counts['failed']}[/red]" if counts["failed"] else "0",
)
console.print(table)
# ────────────────────────────────────────────────────────────────────────────── # ──────────────────────────────────────────────────────────────────────────────
# Helpers # Helpers
# ────────────────────────────────────────────────────────────────────────────── # ──────────────────────────────────────────────────────────────────────────────
+18 -3
View File
@@ -11,6 +11,21 @@ import requests
from src.utils import redact_secrets from src.utils import redact_secrets
# curl_cffi has its own exception hierarchy (rooted at CurlError → OSError),
# completely separate from requests.exceptions. Import them so _make_request
# can catch both when a curl_cffi session is in use.
try:
from curl_cffi.requests.exceptions import (
HTTPError as _CurlHTTPError,
ConnectionError as _CurlConnectionError,
Timeout as _CurlTimeout,
)
except ImportError:
# Fall back to requests types — catching them twice is harmless.
_CurlHTTPError = requests.HTTPError # type: ignore[misc,assignment]
_CurlConnectionError = requests.ConnectionError # type: ignore[misc,assignment]
_CurlTimeout = requests.Timeout # type: ignore[misc,assignment]
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# Request timeouts (connect, read) in seconds # Request timeouts (connect, read) in seconds
@@ -271,7 +286,7 @@ class BaseProvider(ABC):
except ProviderError: except ProviderError:
raise raise
except (requests.ConnectionError, requests.Timeout) as e: except (requests.ConnectionError, requests.Timeout, _CurlConnectionError, _CurlTimeout) as e:
last_exc = e last_exc = e
if attempt > MAX_RETRIES: if attempt > MAX_RETRIES:
raise ProviderError( raise ProviderError(
@@ -293,7 +308,7 @@ class BaseProvider(ABC):
) )
time.sleep(wait) time.sleep(wait)
except requests.HTTPError as e: except (requests.HTTPError, _CurlHTTPError) as e:
raise ProviderError( raise ProviderError(
self.provider_name, f"{method} {url}", e self.provider_name, f"{method} {url}", e
) from e ) from e
@@ -311,7 +326,7 @@ class BaseProvider(ABC):
msg = ( msg = (
f"[{self.provider_name}] Authentication failed (401 Unauthorized). " f"[{self.provider_name}] Authentication failed (401 Unauthorized). "
"Your session token has likely expired. " "Your session token has likely expired. "
"Run 'python -m src.main auth' to refresh your token." "Run 'ai-chat-exporter auth' to refresh your token."
) )
logger.error(msg) logger.error(msg)
raise ProviderError( raise ProviderError(
+536 -37
View File
@@ -1,28 +1,76 @@
"""ChatGPT provider — accesses chat.openai.com internal web API.""" """ChatGPT provider — accesses chat.openai.com internal web API.
ChatGPT Projects discovery
--------------------------
ChatGPT Projects are internally implemented as "snorlax"-type gizmos with IDs
starting with "g-p-". They are *not* returned by any gizmo listing endpoint
(/gizmos/mine, /gizmos/pinned, /gizmos/discovery, /gizmos/search). The
frontend appears to load project IDs from page-level state, not a dedicated
listing API.
Therefore, project IDs must be supplied by the user via CHATGPT_PROJECT_IDS.
Each project gizmo ID looks like "g-p-68c2b2b3037c8191890036fb4ae3ed9f" and
can be read from the browser URL when viewing a project:
https://chatgpt.com/g/{project-gizmo-id}-{slug}/project
Project conversations are fetched via cursor-based pagination at:
GET /backend-api/gizmos/{project_gizmo_id}/conversations?cursor=0
Response: {"items": [...], "cursor": "<opaque_base64_or_null>"}
Pagination ends when cursor is null or an empty string.
"""
import logging import logging
import os import os
from typing import Any from typing import Any
from src.providers.base import BaseProvider, ProviderError from curl_cffi import requests as curl_requests
from src.providers.base import BaseProvider, ProviderError, REQUEST_TIMEOUT
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
BASE_URL = "https://chatgpt.com/backend-api" BASE_URL = "https://chatgpt.com/backend-api"
AUTH_SESSION_URL = "https://chatgpt.com/api/auth/session"
# Chrome version to impersonate — must match a version curl_cffi supports.
# Run: python -c "from curl_cffi.requests import BrowserType; print(list(BrowserType))"
IMPERSONATE = "chrome120"
class ChatGPTProvider(BaseProvider): class ChatGPTProvider(BaseProvider):
"""Provider for ChatGPT conversations via the internal web API. """Provider for ChatGPT conversations via the internal web API.
Authentication: Authorization: Bearer <CHATGPT_SESSION_TOKEN> Uses curl_cffi to impersonate Chrome's TLS fingerprint, bypassing
Token: __Secure-next-auth.session-token cookie value (a JWT). Cloudflare's bot detection which blocks standard Python requests.
Typical validity: ~7 days.
Authentication is a two-step process:
1. Send __Secure-next-auth.session-token as a Cookie to /api/auth/session
to obtain a short-lived accessToken.
2. Use that accessToken as the Bearer token for all backend-api calls.
Token: __Secure-next-auth.session-token cookie (~7 day lifetime).
""" """
provider_name = "chatgpt" provider_name = "chatgpt"
def __init__(self, session_token: str | None = None) -> None: def __init__(
super().__init__() self,
session_token: str | None = None,
session_token_1: str | None = None,
project_ids: list[str] | None = None,
) -> None:
# Pass a curl_cffi session to the base class instead of a requests.Session.
# curl_cffi.requests.Session is API-compatible with requests.Session.
cf_session = curl_requests.Session(impersonate=IMPERSONATE)
super().__init__(session=cf_session) # type: ignore[arg-type]
# Remove headers that curl_cffi manages as part of its Chrome fingerprint.
# Overriding User-Agent, Accept, or Accept-Language with non-Chrome values
# creates header/TLS inconsistencies that Cloudflare's bot detection flags.
self._session.headers.pop("User-Agent", None)
self._session.headers.pop("Accept", None)
self._session.headers.pop("Accept-Language", None)
token = session_token or os.getenv("CHATGPT_SESSION_TOKEN", "").strip() token = session_token or os.getenv("CHATGPT_SESSION_TOKEN", "").strip()
if not token: if not token:
raise ProviderError( raise ProviderError(
@@ -30,26 +78,114 @@ class ChatGPTProvider(BaseProvider):
"init", "init",
RuntimeError( RuntimeError(
"CHATGPT_SESSION_TOKEN is not set. " "CHATGPT_SESSION_TOKEN is not set. "
"Run 'python -m src.main auth' to configure it." "Run 'ai-chat-exporter auth' to configure it."
), ),
) )
# Never log the token value self._session_token = token
# Second chunk of the session token (ChatGPT splits large cookies into
# __Secure-next-auth.session-token.0 and .1 to stay under the 4KB limit).
token_1 = session_token_1 or os.getenv("CHATGPT_SESSION_TOKEN_1", "").strip() or None
# Project gizmo IDs (g-p-xxx) whose conversations we'll fetch.
# ChatGPT project conversations do not appear in the default
# /conversations listing — they require explicit project IDs.
self._project_ids: list[str] = project_ids or []
# Maps conv_id → project_name; populated by fetch_all_conversations()
self._project_map: dict[str, str] = {}
# Cache of project_id → display name (avoids re-fetching gizmo details)
self._project_name_cache: dict[str, str] = {}
# ChatGPT now splits large session cookies into .0 / .1 chunks.
# Always send both named chunks; the server reassembles them.
self._session.cookies.set(
"__Secure-next-auth.session-token.0",
token,
domain="chatgpt.com",
path="/",
)
if token_1:
self._session.cookies.set(
"__Secure-next-auth.session-token.1",
token_1,
domain="chatgpt.com",
path="/",
)
logger.debug("[chatgpt] Set both session cookie chunks (.0 and .1)")
else:
logger.debug("[chatgpt] Set session cookie chunk .0 only (no .1 configured)")
# Set only Referer and sec-fetch-* headers for the auth exchange.
# Origin is intentionally omitted: Chrome does not send Origin on
# same-origin GET requests, and its presence alongside
# sec-fetch-site: same-origin contradicts the browser fingerprint.
self._session.headers.update( self._session.headers.update(
{ {
"Authorization": f"Bearer {token}",
"Referer": "https://chatgpt.com/", "Referer": "https://chatgpt.com/",
"Origin": "https://chatgpt.com", "sec-fetch-dest": "empty",
"sec-fetch-mode": "cors",
"sec-fetch-site": "same-origin",
} }
) )
logger.debug("[chatgpt] Session initialised (token: [REDACTED])")
# Exchange the session cookie for an access token
self._access_token: str = self._fetch_access_token()
# Now set backend-api headers (after auth, so they don't interfere with
# the auth exchange which expects a browser-style request).
self._session.headers["Authorization"] = f"Bearer {self._access_token}"
self._session.headers["Accept"] = "application/json"
self._session.headers["Origin"] = "https://chatgpt.com"
logger.debug(
"[chatgpt] Session initialised (Chrome TLS impersonation, %d project ID(s) configured)",
len(self._project_ids),
)
def _fetch_access_token(self) -> str:
"""Exchange the session cookie for a Bearer access token.
Calls GET /api/auth/session — the cookie jar already contains the
session token, so no manual Cookie header is needed.
Returns {"accessToken": "...", "user": {...}}.
"""
logger.debug("[chatgpt] Fetching access token from %s", AUTH_SESSION_URL)
try:
resp = self._session.get(AUTH_SESSION_URL, timeout=REQUEST_TIMEOUT)
resp.raise_for_status()
data = resp.json()
except Exception as e:
raise ProviderError(
self.provider_name,
"fetch_access_token",
RuntimeError(
f"Could not exchange session token for access token: {e}. "
"Check that your CHATGPT_SESSION_TOKEN is current and not expired."
),
) from e
access_token = data.get("accessToken")
if not access_token:
raise ProviderError(
self.provider_name,
"fetch_access_token",
RuntimeError(
"No accessToken in /api/auth/session response. "
"Your session token may be expired — run 'ai-chat-exporter auth' to refresh."
),
)
return access_token
def _handle_401(self) -> None: def _handle_401(self) -> None:
msg = ( msg = (
"[chatgpt] Authentication failed (401 Unauthorized). " "[chatgpt] Authentication failed (401 Unauthorized). "
"Your __Secure-next-auth.session-token has likely expired (~7 day lifetime). " "Your __Secure-next-auth.session-token has likely expired (~7 day lifetime). "
"The session token is used to obtain a short-lived access token via /api/auth/session. "
"To refresh: open chatgpt.com in Chrome → F12 → Application → Cookies " "To refresh: open chatgpt.com in Chrome → F12 → Application → Cookies "
"→ find '__Secure-next-auth.session-token' → copy the value. " "→ find '__Secure-next-auth.session-token' → copy the value. "
"Then run 'python -m src.main auth' or update CHATGPT_SESSION_TOKEN in .env." "Then run 'ai-chat-exporter auth' or update CHATGPT_SESSION_TOKEN in .env."
) )
logger.error(msg) logger.error(msg)
raise ProviderError( raise ProviderError(
@@ -58,14 +194,22 @@ class ChatGPTProvider(BaseProvider):
RuntimeError("401 Unauthorized — ChatGPT token expired"), RuntimeError("401 Unauthorized — ChatGPT token expired"),
) )
# ------------------------------------------------------------------
# Default workspace conversations (offset-based pagination)
# ------------------------------------------------------------------
def list_conversations(self, offset: int = 0, limit: int = 100) -> list[dict]: def list_conversations(self, offset: int = 0, limit: int = 100) -> list[dict]:
"""Fetch one page of conversations. """Fetch one page of conversations from the default workspace.
Note: Project conversations are NOT included here. They require
separate fetching via list_project_conversations().
Returns: Returns:
List of conversation summary dicts. List of conversation summary dicts.
""" """
url = f"{BASE_URL}/conversations" url = f"{BASE_URL}/conversations"
params = {"offset": offset, "limit": limit, "order": "updated"} params = {"offset": offset, "limit": limit, "order": "updated"}
logger.debug("[chatgpt] list_conversations: GET %s params=%s", url, params)
try: try:
data = self._make_request("GET", url, params=params) data = self._make_request("GET", url, params=params)
except ProviderError: except ProviderError:
@@ -75,18 +219,315 @@ class ChatGPTProvider(BaseProvider):
if not isinstance(data, dict): if not isinstance(data, dict):
self._warn_unexpected_schema("list_conversations", "root") self._warn_unexpected_schema("list_conversations", "root")
logger.debug("[chatgpt] list_conversations: unexpected root type %s", type(data))
return [] return []
items = data.get("items") items = data.get("items")
if items is None: if items is None:
self._warn_unexpected_schema("list_conversations", "items") self._warn_unexpected_schema("list_conversations", "items")
logger.debug("[chatgpt] list_conversations: response keys = %s", list(data.keys()))
return [] return []
logger.debug("[chatgpt] list_conversations: got %d items (offset=%d)", len(items), offset)
return items return items
# ------------------------------------------------------------------
# Project conversations (cursor-based pagination)
# ------------------------------------------------------------------
def _fetch_project_name(self, project_id: str) -> str:
"""Fetch the display name for a project gizmo.
Calls GET /backend-api/gizmos/{project_id} and returns the display
name from gizmo.display.name. Falls back to the project_id itself
if the fetch fails or the name is missing.
Result is cached in self._project_name_cache.
"""
if project_id in self._project_name_cache:
return self._project_name_cache[project_id]
url = f"{BASE_URL}/gizmos/{project_id}"
logger.debug("[chatgpt] _fetch_project_name: GET %s", url)
try:
data = self._make_request("GET", url)
gizmo = data.get("gizmo", {}) if isinstance(data, dict) else {}
name = (gizmo.get("display") or {}).get("name") or gizmo.get("name") or ""
name = name.strip() or project_id
gizmo_type = gizmo.get("gizmo_type", "?")
logger.debug(
"[chatgpt] _fetch_project_name[%s]: name=%r gizmo_type=%r",
project_id[:12],
name,
gizmo_type,
)
except ProviderError as e:
logger.warning(
"[chatgpt] Could not fetch project name for %s: %s — using ID as name",
project_id,
e,
)
name = project_id
self._project_name_cache[project_id] = name
return name
def list_project_conversations(
self, project_id: str, cursor: str = "0"
) -> tuple[list[dict], str | None]:
"""Fetch one page of conversations for a project gizmo.
Uses cursor-based pagination (not offset). The initial cursor is "0".
Subsequent cursors come from the response's "cursor" field.
Endpoint: GET /backend-api/gizmos/{project_id}/conversations?cursor=<cursor>
Returns:
(items, next_cursor) — next_cursor is None or "" when exhausted.
"""
url = f"{BASE_URL}/gizmos/{project_id}/conversations"
params = {"cursor": cursor}
logger.debug(
"[chatgpt] list_project_conversations[%s]: GET %s cursor=%r",
project_id[:12],
url,
cursor,
)
try:
data = self._make_request("GET", url, params=params)
except ProviderError:
raise
except Exception as e:
raise ProviderError(self.provider_name, "list_project_conversations", e) from e
logger.debug(
"[chatgpt] list_project_conversations[%s]: response type=%s",
project_id[:12],
type(data).__name__,
)
if isinstance(data, list):
# Bare list — no next cursor available
logger.debug(
"[chatgpt] list_project_conversations[%s]: bare list with %d items",
project_id[:12],
len(data),
)
return data, None
if not isinstance(data, dict):
self._warn_unexpected_schema("list_project_conversations", "root")
logger.debug(
"[chatgpt] list_project_conversations[%s]: unexpected type %s value=%r",
project_id[:12],
type(data),
data,
)
return [], None
logger.debug(
"[chatgpt] list_project_conversations[%s]: response keys=%s",
project_id[:12],
list(data.keys()),
)
items = data.get("items") or data.get("conversations") or []
next_cursor = data.get("cursor") or None # empty string → treat as None
if not items and data:
logger.debug(
"[chatgpt] list_project_conversations[%s]: no items found; full response=%r",
project_id[:12],
data,
)
logger.debug(
"[chatgpt] list_project_conversations[%s]: %d items, next_cursor=%r",
project_id[:12],
len(items),
next_cursor[:20] + "" if next_cursor and len(next_cursor) > 20 else next_cursor,
)
return items, next_cursor
# ------------------------------------------------------------------
# Combined fetch (default workspace + all configured projects)
# ------------------------------------------------------------------
def fetch_all_conversations(self, since=None) -> list[dict]:
"""Fetch all conversations: default workspace + every configured project.
ChatGPT project conversations are not included in the default
/conversations listing. They must be fetched separately via the
gizmos conversations endpoint using project IDs from CHATGPT_PROJECT_IDS.
Builds self._project_map (conv_id → project_name) as a side effect so
that normalize_conversation() can attach the project name without an
additional API call.
Args:
since: Optional datetime — only return conversations updated at or
after this time (client-side filter, same as base class).
Returns:
Combined list of raw conversation summary dicts.
"""
# Reset maps so a fresh fetch always rebuilds them cleanly
self._project_map = {}
# --- Default workspace (base class handles offset-based pagination) ---
logger.info("[chatgpt] Fetching default workspace conversations…")
default_convs = super().fetch_all_conversations(since=None)
logger.info("[chatgpt] Default workspace: %d conversations", len(default_convs))
# --- Project conversations ---
if not self._project_ids:
logger.info(
"[chatgpt] No project IDs configured — skipping project conversations. "
"To include projects, set CHATGPT_PROJECT_IDS in .env "
"(see 'ai-chat-exporter auth' for instructions)."
)
return self._apply_since_filter(default_convs, since)
logger.info(
"[chatgpt] Fetching conversations for %d project(s): %s",
len(self._project_ids),
self._project_ids,
)
project_convs: list[dict] = []
for project_id in self._project_ids:
project_name = self._fetch_project_name(project_id)
logger.info(
"[chatgpt] Project '%s' (%s): fetching conversations…",
project_name,
project_id,
)
cursor: str = "0"
page = 0
project_total = 0
while True:
page += 1
logger.debug(
"[chatgpt] Project '%s': page %d cursor=%r",
project_name,
page,
cursor[:20] + "" if len(cursor) > 20 else cursor,
)
try:
batch, next_cursor = self.list_project_conversations(
project_id, cursor=cursor
)
except ProviderError as e:
logger.warning(
"[chatgpt] Project '%s': failed to fetch page %d: %s — stopping pagination",
project_name,
page,
e,
)
break
if not batch:
logger.debug(
"[chatgpt] Project '%s': empty batch on page %d — done",
project_name,
page,
)
break
for conv in batch:
conv_id = conv.get("id")
if conv_id:
self._project_map[conv_id] = project_name
else:
logger.debug(
"[chatgpt] Project '%s': conversation with no id: %r",
project_name,
conv,
)
# Annotate so callers can filter by project without the map
conv["_project_name"] = project_name
project_convs.extend(batch)
project_total += len(batch)
logger.debug(
"[chatgpt] Project '%s': page %d%d items (project total: %d)",
project_name,
page,
len(batch),
project_total,
)
if not next_cursor:
logger.debug(
"[chatgpt] Project '%s': no next cursor — pagination complete",
project_name,
)
break
cursor = next_cursor
logger.info(
"[chatgpt] Project '%s': %d conversations fetched",
project_name,
project_total,
)
all_convs = default_convs + project_convs
logger.info(
"[chatgpt] Total: %d conversations (%d default + %d from %d project(s))",
len(all_convs),
len(default_convs),
len(project_convs),
len(self._project_ids),
)
logger.debug(
"[chatgpt] _project_map: %d entries → %s",
len(self._project_map),
{k[:8]: v for k, v in self._project_map.items()},
)
return self._apply_since_filter(all_convs, since)
def _apply_since_filter(self, convs: list[dict], since) -> list[dict]:
"""Filter conversations to those updated at or after `since`."""
if since is None:
return convs
since_naive = since.replace(tzinfo=None)
filtered = []
for c in convs:
raw_ts = c.get("updated_at") or c.get("update_time") or ""
if raw_ts:
try:
from src.utils import _parse_dt
updated = _parse_dt(str(raw_ts)).replace(tzinfo=None)
if updated >= since_naive:
filtered.append(c)
except Exception:
filtered.append(c) # include if date unparseable
else:
filtered.append(c)
logger.info(
"[chatgpt] After --since filter: %d/%d conversations",
len(filtered),
len(convs),
)
return filtered
# ------------------------------------------------------------------
# Single conversation detail
# ------------------------------------------------------------------
def get_conversation(self, conv_id: str) -> dict: def get_conversation(self, conv_id: str) -> dict:
"""Fetch full conversation detail for a single ID.""" """Fetch full conversation detail for a single ID."""
url = f"{BASE_URL}/conversation/{conv_id}" url = f"{BASE_URL}/conversation/{conv_id}"
logger.debug("[chatgpt] get_conversation: GET %s", url)
try: try:
data = self._make_request("GET", url) data = self._make_request("GET", url)
except ProviderError: except ProviderError:
@@ -98,25 +539,45 @@ class ChatGPTProvider(BaseProvider):
self._warn_unexpected_schema("get_conversation", "root") self._warn_unexpected_schema("get_conversation", "root")
return {} return {}
logger.debug(
"[chatgpt] get_conversation[%s]: keys=%s mapping_size=%d",
conv_id[:8],
list(data.keys()),
len(data.get("mapping", {})),
)
return data return data
# ------------------------------------------------------------------
# Normalization
# ------------------------------------------------------------------
def normalize_conversation(self, raw: dict) -> dict: def normalize_conversation(self, raw: dict) -> dict:
"""Transform ChatGPT raw schema to the common normalized schema. """Transform ChatGPT raw schema to the common normalized schema.
ChatGPT stores messages in a nested ``mapping`` dict where each node ChatGPT stores messages in a nested ``mapping`` dict where each node
has an ``id``, ``message``, and ``children`` list. We walk the tree has an ``id``, ``message``, and ``children`` list. We walk the tree
from the root node to build a flat ordered message list. from the root node to build a flat ordered message list.
Project name is looked up from self._project_map (populated by
fetch_all_conversations). The conversation detail endpoint does not
include project information.
""" """
conv_id = raw.get("id", "") conv_id = raw.get("id", "")
title = raw.get("title") or "Untitled" title = raw.get("title") or "Untitled"
created_at = _ts_to_iso(raw.get("create_time")) created_at = _ts_to_iso(raw.get("create_time"))
updated_at = _ts_to_iso(raw.get("update_time")) updated_at = _ts_to_iso(raw.get("update_time"))
# Project info — ChatGPT calls it "gizmo_id" or stores project info differently. # Prefer _project_name annotation injected from the listing summary
# As of 2024, personal projects appear as a separate projects API; conversations # (propagated by the export loop). Fall back to _project_map lookup.
# linked to a project have a non-null `workspace_id` or similar field. project = raw.get("_project_name") or (
# We use `project_title` if present, else None. self._project_map.get(conv_id) if conv_id else None
project: str | None = raw.get("project_title") or raw.get("workspace_title") or None )
logger.debug(
"[chatgpt] normalize_conversation[%s]: project=%r (source=%s)",
conv_id[:8] if conv_id else "?",
project,
"_project_name" if raw.get("_project_name") else "_project_map",
)
mapping: dict = raw.get("mapping", {}) mapping: dict = raw.get("mapping", {})
messages = _extract_messages(mapping, raw, conv_id) messages = _extract_messages(mapping, raw, conv_id)
@@ -181,18 +642,22 @@ def _extract_messages(
if role in ("user", "assistant"): if role in ("user", "assistant"):
content_obj = msg_data.get("content", {}) content_obj = msg_data.get("content", {})
content_type = content_obj.get("content_type", "text") content_type = content_obj.get("content_type", "text")
text = _extract_text(content_obj, conv_id, node_id)
if content_type != "text":
logger.warning(
"[chatgpt] Skipping %s content in conversation %s message %s "
"— rich content not yet supported (see FUTURE.md)",
content_type,
conv_id[:8],
node_id[:8],
)
elif text:
ts = msg_data.get("create_time") ts = msg_data.get("create_time")
# Content types whose parts[] contain plain text strings.
# model_editable_context / user_editable_context = project instructions
# thoughts / reasoning_recap = o1/o3 reasoning traces
_TEXT_PARTS_TYPES = {
"text",
"model_editable_context",
"user_editable_context",
"thoughts",
"reasoning_recap",
}
if content_type in _TEXT_PARTS_TYPES:
text = _extract_text(content_obj, conv_id, node_id)
if text:
messages.append( messages.append(
{ {
"role": role, "role": role,
@@ -203,7 +668,32 @@ def _extract_messages(
) )
else: else:
logger.debug( logger.debug(
"[chatgpt] Skipping empty message in conversation %s", conv_id[:8] "[chatgpt] Skipping empty %s message in conversation %s",
content_type,
conv_id[:8],
)
elif content_type == "code":
# Inline code response — extract and wrap in a fenced code block
code_text = content_obj.get("text") or "\n".join(
p for p in content_obj.get("parts", []) if isinstance(p, str)
)
language = content_obj.get("language", "")
if code_text:
messages.append(
{
"role": role,
"content": f"```{language}\n{code_text}\n```",
"content_type": "code",
"timestamp": _ts_to_iso(ts) if ts else None,
}
)
else:
logger.warning(
"[chatgpt] Skipping %s content in conversation %s message %s "
"— rich content not yet supported (see FUTURE.md)",
content_type,
conv_id[:8],
node_id[:8],
) )
# Walk children in order (ChatGPT typically has one child per node in a linear chat) # Walk children in order (ChatGPT typically has one child per node in a linear chat)
@@ -229,7 +719,13 @@ def _find_root(mapping: dict[str, Any]) -> str | None:
def _extract_text(content_obj: dict, conv_id: str, node_id: str) -> str: def _extract_text(content_obj: dict, conv_id: str, node_id: str) -> str:
"""Extract plain text from a ChatGPT content object.""" """Extract plain text from a ChatGPT content object.
Handles three part shapes:
- str — plain text (most messages)
- dict with content_type="text" — wrapped text part
- dict with "content" key — o1/o3 thoughts/reasoning parts
"""
parts = content_obj.get("parts", []) parts = content_obj.get("parts", [])
if not parts: if not parts:
return "" return ""
@@ -239,16 +735,19 @@ def _extract_text(content_obj: dict, conv_id: str, node_id: str) -> str:
if isinstance(part, str): if isinstance(part, str):
text_parts.append(part) text_parts.append(part)
elif isinstance(part, dict): elif isinstance(part, dict):
# Could be an image or file reference — skip and warn part_type = part.get("content_type", "")
part_type = part.get("content_type", "unknown") if part_type == "text":
if part_type != "text": text_parts.append(part.get("text", ""))
elif "content" in part:
# o1/o3 thoughts parts: {"summary": "...", "content": "..."}
text_parts.append(part["content"])
elif part_type:
# Image, file, or other binary attachment — skip and warn
logger.warning( logger.warning(
"[chatgpt] Skipping %s attachment in conversation %s " "[chatgpt] Skipping %s attachment in conversation %s "
"— rich content not yet supported (see FUTURE.md)", "— rich content not yet supported (see FUTURE.md)",
part_type, part_type,
conv_id[:8], conv_id[:8],
) )
else:
text_parts.append(part.get("text", ""))
return "\n".join(t for t in text_parts if t) return "\n".join(t for t in text_parts if t)
+19 -9
View File
@@ -3,25 +3,35 @@
import logging import logging
import os import os
from curl_cffi import requests as curl_requests
from src.providers.base import BaseProvider, ProviderError from src.providers.base import BaseProvider, ProviderError
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
BASE_URL = "https://claude.ai/api" BASE_URL = "https://claude.ai/api"
IMPERSONATE = "chrome120"
class ClaudeProvider(BaseProvider): class ClaudeProvider(BaseProvider):
"""Provider for Claude conversations via the internal web API. """Provider for Claude conversations via the internal web API.
Authentication: Cookie: sessionKey=<CLAUDE_SESSION_KEY> Uses curl_cffi to impersonate Chrome's TLS fingerprint, bypassing
Token: sessionKey cookie value from claude.ai. Cloudflare's bot detection (same issue as chatgpt.com).
Typical validity: ~30 days (opaque; expiry cannot be decoded client-side).
Authentication: sessionKey cookie (~30 day lifetime, opaque string).
Expiry cannot be decoded client-side — a 401 is the only signal.
""" """
provider_name = "claude" provider_name = "claude"
def __init__(self, session_key: str | None = None) -> None: def __init__(self, session_key: str | None = None) -> None:
super().__init__() cf_session = curl_requests.Session(impersonate=IMPERSONATE)
super().__init__(session=cf_session) # type: ignore[arg-type]
# Remove base class User-Agent so curl_cffi uses its Chrome-matched UA
self._session.headers.pop("User-Agent", None)
key = session_key or os.getenv("CLAUDE_SESSION_KEY", "").strip() key = session_key or os.getenv("CLAUDE_SESSION_KEY", "").strip()
if not key: if not key:
raise ProviderError( raise ProviderError(
@@ -29,19 +39,19 @@ class ClaudeProvider(BaseProvider):
"init", "init",
RuntimeError( RuntimeError(
"CLAUDE_SESSION_KEY is not set. " "CLAUDE_SESSION_KEY is not set. "
"Run 'python -m src.main auth' to configure it." "Run 'ai-chat-exporter auth' to configure it."
), ),
) )
# Set cookie header; never log the key value # Set sessionKey in the cookie jar
self._session.cookies.set("sessionKey", key, domain="claude.ai", path="/")
self._session.headers.update( self._session.headers.update(
{ {
"Cookie": f"sessionKey={key}",
"Referer": "https://claude.ai/", "Referer": "https://claude.ai/",
"Origin": "https://claude.ai", "Origin": "https://claude.ai",
} }
) )
self._org_id: str | None = None # cached per session self._org_id: str | None = None # cached per session
logger.debug("[claude] Session initialised (key: [REDACTED])") logger.debug("[claude] Session initialised with Chrome TLS impersonation (key: [REDACTED])")
def _handle_401(self) -> None: def _handle_401(self) -> None:
msg = ( msg = (
@@ -50,7 +60,7 @@ class ClaudeProvider(BaseProvider):
"Note: Claude session keys are opaque — a 401 is the only expiry signal. " "Note: Claude session keys are opaque — a 401 is the only expiry signal. "
"To refresh: open claude.ai in Chrome → F12 → Application → Cookies " "To refresh: open claude.ai in Chrome → F12 → Application → Cookies "
"→ find 'sessionKey' → copy the value. " "→ find 'sessionKey' → copy the value. "
"Then run 'python -m src.main auth' or update CLAUDE_SESSION_KEY in .env." "Then run 'ai-chat-exporter auth' or update CLAUDE_SESSION_KEY in .env."
) )
logger.error(msg) logger.error(msg)
raise ProviderError( raise ProviderError(
+129
View File
@@ -0,0 +1,129 @@
"""CLI-level tests using Click's CliRunner — no live API calls required."""
import pytest
from click.testing import CliRunner
from src.cache import Cache
from src.main import _filter_by_project, cli
# ---------------------------------------------------------------------------
# _filter_by_project (T-27)
# ---------------------------------------------------------------------------
class TestFilterByProject:
"""Unit tests for the project filter logic used by export/list/joplin."""
# ChatGPT conversations use the _project_name annotation key
def _chatgpt(self, conv_id, project_name):
return {"id": conv_id, "_project_name": project_name}
# Claude conversations use the project dict key
def _claude(self, conv_id, project_name):
proj = {"name": project_name} if project_name else None
return {"id": conv_id, "project": proj}
def test_none_filter_keeps_no_project_chatgpt(self):
convs = [self._chatgpt("a", None), self._chatgpt("b", "Python Course")]
result = _filter_by_project(convs, "none")
assert len(result) == 1
assert result[0]["id"] == "a"
def test_none_filter_keeps_no_project_claude(self):
convs = [self._claude("a", None), self._claude("b", "Python Course")]
result = _filter_by_project(convs, "none")
assert len(result) == 1
assert result[0]["id"] == "a"
def test_name_filter_case_insensitive(self):
convs = [
self._chatgpt("a", "Python Course"),
self._chatgpt("b", "Java Course"),
self._chatgpt("c", None),
]
result = _filter_by_project(convs, "PYTHON")
assert len(result) == 1
assert result[0]["id"] == "a"
def test_name_filter_substring_match(self):
convs = [
self._chatgpt("a", "Python Advanced Course"),
self._chatgpt("b", "Python Basics"),
self._chatgpt("c", "JavaScript"),
]
result = _filter_by_project(convs, "python")
assert len(result) == 2
assert {c["id"] for c in result} == {"a", "b"}
def test_no_matches_returns_empty(self):
convs = [self._chatgpt("a", "Python Course"), self._chatgpt("b", None)]
result = _filter_by_project(convs, "ruby")
assert result == []
def test_none_filter_excludes_all_with_projects(self):
convs = [self._chatgpt("a", "Project A"), self._chatgpt("b", "Project B")]
result = _filter_by_project(convs, "none")
assert result == []
def test_empty_string_project_treated_as_no_project(self):
convs = [{"id": "a", "_project_name": ""}, {"id": "b", "_project_name": "Real"}]
result = _filter_by_project(convs, "none")
assert len(result) == 1
assert result[0]["id"] == "a"
def test_claude_project_string_matched(self):
# Claude can also have project as a plain string
convs = [{"id": "a", "project": "python-course"}, {"id": "b", "project": None}]
result = _filter_by_project(convs, "python")
assert len(result) == 1
assert result[0]["id"] == "a"
# ---------------------------------------------------------------------------
# export --since validation (T-25)
# ---------------------------------------------------------------------------
class TestExportSinceValidation:
"""Test that --since with an invalid date exits cleanly with an error message."""
def _pre_populated_cache(self, tmp_path) -> Cache:
"""Create a cache that passes the ToS gate and first-run doctor check."""
cache = Cache(tmp_path)
cache.acknowledge_tos()
cache.mark_exported("chatgpt", "dummy-conv", {"updated_at": "2024-01-01T00:00:00Z"})
return cache
def test_invalid_since_date_exits_with_error(self, tmp_path):
self._pre_populated_cache(tmp_path)
runner = CliRunner(mix_stderr=True)
result = runner.invoke(
cli,
["--no-log-file", "export", "--since", "notadate"],
env={
"CHATGPT_SESSION_TOKEN": "eyJtesttoken",
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
},
)
assert result.exit_code == 1
assert "Invalid --since date" in result.output
assert "YYYY-MM-DD" in result.output
def test_valid_since_date_does_not_error(self, tmp_path):
"""A valid date should not produce the invalid-date error (may fail later on API)."""
self._pre_populated_cache(tmp_path)
runner = CliRunner(mix_stderr=True)
result = runner.invoke(
cli,
["--no-log-file", "export", "--since", "2024-01-01"],
env={
"CHATGPT_SESSION_TOKEN": "eyJtesttoken",
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
},
)
assert "Invalid --since date" not in result.output
+56
View File
@@ -0,0 +1,56 @@
"""Tests for src/config.py — token validation logic (T-14)."""
import logging
import time
import jwt
import pytest
from src.config import _validate_chatgpt_token
class TestValidateChatGPTToken:
def test_expired_token_logs_warning(self, caplog):
# T-14: expired JWT must produce a clear warning
payload = {"exp": int(time.time()) - 3600} # expired 1 hour ago
token = jwt.encode(payload, "secret", algorithm="HS256")
with caplog.at_level(logging.WARNING, logger="src.config"):
result = _validate_chatgpt_token(token)
assert any("expired" in r.message.lower() for r in caplog.records)
assert result is not None # still returns the expiry datetime
def test_expiring_within_24h_logs_warning(self, caplog):
payload = {"exp": int(time.time()) + 3600} # expires in 1 hour
token = jwt.encode(payload, "secret", algorithm="HS256")
with caplog.at_level(logging.WARNING, logger="src.config"):
_validate_chatgpt_token(token)
assert any("less than 24 hours" in r.message for r in caplog.records)
def test_valid_token_no_expiry_warning(self, caplog):
payload = {"exp": int(time.time()) + 86400 * 5} # valid for 5 days
token = jwt.encode(payload, "secret", algorithm="HS256")
with caplog.at_level(logging.WARNING, logger="src.config"):
result = _validate_chatgpt_token(token)
assert not any("expired" in r.message.lower() for r in caplog.records)
assert result is not None
def test_token_without_exp_claim_logs_warning(self, caplog):
payload = {"sub": "user123"} # no exp
token = jwt.encode(payload, "secret", algorithm="HS256")
with caplog.at_level(logging.WARNING, logger="src.config"):
result = _validate_chatgpt_token(token)
assert any("'exp'" in r.message or "no 'exp'" in r.message for r in caplog.records)
assert result is None
def test_jwe_encrypted_token_returns_none(self, caplog):
# JWE tokens (alg=dir) cannot be decoded client-side — this is normal for ChatGPT
jwe_like = "eyJhbGciOiJkaXIiLCJlbmMiOiJBMjU2R0NNIn0.fake.token.data.here"
with caplog.at_level(logging.DEBUG, logger="src.config"):
result = _validate_chatgpt_token(jwe_like)
assert result is None # cannot decode, but not an error
def test_non_jwt_string_logs_warning(self, caplog):
with caplog.at_level(logging.WARNING, logger="src.config"):
result = _validate_chatgpt_token("notajwttoken")
assert any("does not look like a JWT" in r.message for r in caplog.records)
assert result is None
+28
View File
@@ -199,6 +199,34 @@ class TestJSONExporter:
assert " " in raw assert " " in raw
class TestBothFormats:
"""T-38: Markdown and JSON exporters produce matching filenames for the same conversation."""
def test_both_formats_produce_files(self, tmp_path):
md_exp = MarkdownExporter(tmp_path)
json_exp = JSONExporter(tmp_path)
md_path = md_exp.export(SAMPLE_CONV)
json_path = json_exp.export(SAMPLE_CONV)
assert md_path.exists()
assert json_path.exists()
def test_both_formats_have_matching_stems(self, tmp_path):
md_exp = MarkdownExporter(tmp_path)
json_exp = JSONExporter(tmp_path)
md_path = md_exp.export(SAMPLE_CONV)
json_path = json_exp.export(SAMPLE_CONV)
assert md_path.suffix == ".md"
assert json_path.suffix == ".json"
assert md_path.stem == json_path.stem
def test_both_formats_same_directory(self, tmp_path):
md_exp = MarkdownExporter(tmp_path)
json_exp = JSONExporter(tmp_path)
md_path = md_exp.export(SAMPLE_CONV)
json_path = json_exp.export(SAMPLE_CONV)
assert md_path.parent == json_path.parent
class TestYamlEscape: class TestYamlEscape:
def test_escapes_double_quotes(self): def test_escapes_double_quotes(self):
assert _yaml_escape('Say "hello"') == 'Say \\"hello\\"' assert _yaml_escape('Say "hello"') == 'Say \\"hello\\"'
+341
View File
@@ -0,0 +1,341 @@
"""Unit tests for src/joplin.py (JoplinClient)."""
from unittest.mock import MagicMock, patch
import pytest
import requests
from src.joplin import JoplinClient, JoplinError, _http_error_message, _timeout_message, notebook_title
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _make_client() -> JoplinClient:
return JoplinClient(base_url="http://localhost:41184", token="test-token")
def _mock_response(json_data=None, text="", status_code=200):
resp = MagicMock()
resp.status_code = status_code
resp.text = text
resp.json.return_value = json_data or {}
resp.raise_for_status = MagicMock()
if status_code >= 400:
resp.raise_for_status.side_effect = requests.exceptions.HTTPError(
response=resp
)
return resp
# ---------------------------------------------------------------------------
# notebook_title helper
# ---------------------------------------------------------------------------
class TestNotebookTitle:
def test_no_project(self):
assert notebook_title("chatgpt", None) == "ChatGPT - No Project"
def test_no_project_string(self):
assert notebook_title("chatgpt", "no-project") == "ChatGPT - No Project"
def test_project_with_hyphens(self):
assert notebook_title("chatgpt", "my-project") == "ChatGPT - My Project"
def test_claude_provider(self):
assert notebook_title("claude", "budget-tracker") == "Claude - Budget Tracker"
def test_multi_word_project(self):
assert notebook_title("claude", "ai-research-notes") == "Claude - Ai Research Notes"
# ---------------------------------------------------------------------------
# ping
# ---------------------------------------------------------------------------
class TestPing:
def test_ping_success(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(text="JoplinClipperServer")
assert client.ping() is True
def test_ping_not_joplin(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(text="SomeOtherServer")
assert client.ping() is False
def test_ping_connection_refused(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.side_effect = requests.exceptions.ConnectionError()
assert client.ping() is False
def test_ping_timeout_returns_false(self):
"""Ping timeout is not an error — Joplin just isn't responding."""
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.side_effect = requests.exceptions.Timeout()
assert client.ping() is False
def test_ping_invalid_url_raises_joplin_error(self):
"""Non-connection, non-timeout errors (e.g. invalid URL) surface as JoplinError."""
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.side_effect = requests.exceptions.InvalidURL("bad url")
with pytest.raises(JoplinError):
client.ping()
class TestValidateToken:
def test_validate_token_success(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(json_data={"items": [], "has_more": False})
client.validate_token() # should not raise
def test_validate_token_401_raises_joplin_error(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(status_code=401)
with pytest.raises(JoplinError, match="401"):
client.validate_token()
class TestTimeoutMessage:
def test_includes_timeout_duration(self):
import src.joplin as joplin_module
msg = _timeout_message("POST", "/notes")
assert "POST" in msg
assert "/notes" in msg
assert str(joplin_module._REQUEST_TIMEOUT) in msg
def test_includes_actionable_hints(self):
msg = _timeout_message("PUT", "/notes/abc")
assert "JOPLIN_REQUEST_TIMEOUT" in msg
# Should mention at least one cause
assert "large" in msg.lower() or "busy" in msg.lower() or "frozen" in msg.lower()
class TestTimeoutHandling:
def test_get_timeout_raises_joplin_error_with_clear_message(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.side_effect = requests.exceptions.Timeout()
with pytest.raises(JoplinError) as exc_info:
client._get("/folders")
assert "timed out" in str(exc_info.value).lower()
assert "JOPLIN_REQUEST_TIMEOUT" in str(exc_info.value)
def test_post_timeout_raises_joplin_error_with_clear_message(self):
client = _make_client()
with patch("requests.post") as mock_post:
mock_post.side_effect = requests.exceptions.Timeout()
with pytest.raises(JoplinError) as exc_info:
client._post("/notes", {"title": "Test"})
assert "timed out" in str(exc_info.value).lower()
def test_put_timeout_raises_joplin_error_with_clear_message(self):
client = _make_client()
with patch("requests.put") as mock_put:
mock_put.side_effect = requests.exceptions.Timeout()
with pytest.raises(JoplinError) as exc_info:
client._put("/notes/abc", {"title": "Test"})
assert "timed out" in str(exc_info.value).lower()
def test_create_note_timeout_propagates(self):
"""Timeout on create_note surfaces as JoplinError, not raw requests exception."""
client = _make_client()
with patch("requests.post") as mock_post:
mock_post.side_effect = requests.exceptions.Timeout()
with pytest.raises(JoplinError, match="timed out"):
client.create_note("Big Note", "x" * 100_000, "nb-123")
def test_update_note_timeout_propagates(self):
client = _make_client()
with patch("requests.put") as mock_put:
mock_put.side_effect = requests.exceptions.Timeout()
with pytest.raises(JoplinError, match="timed out"):
client.update_note("note-id", "Big Note", "x" * 100_000)
class TestHttpErrorMessage:
def test_401_gives_token_hint(self):
resp = MagicMock()
resp.status_code = 401
resp.text = "Unauthorized"
e = requests.exceptions.HTTPError(response=resp)
msg = _http_error_message("GET", "/folders", e)
assert "401" in msg
assert "token" in msg.lower()
def test_404_gives_deleted_note_hint(self):
resp = MagicMock()
resp.status_code = 404
resp.text = "Not Found"
e = requests.exceptions.HTTPError(response=resp)
msg = _http_error_message("PUT", "/notes/abc", e)
assert "404" in msg
assert "deleted" in msg.lower()
def test_other_error_includes_status_and_body(self):
resp = MagicMock()
resp.status_code = 500
resp.text = "Internal Server Error"
e = requests.exceptions.HTTPError(response=resp)
msg = _http_error_message("POST", "/notes", e)
assert "500" in msg
# ---------------------------------------------------------------------------
# list_notebooks
# ---------------------------------------------------------------------------
class TestListNotebooks:
def test_list_notebooks_single_page(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={"items": [{"id": "nb1", "title": "ChatGPT - No Project"}], "has_more": False}
)
result = client.list_notebooks()
assert len(result) == 1
assert result[0]["id"] == "nb1"
def test_list_notebooks_paginated(self):
client = _make_client()
page1 = _mock_response(
json_data={"items": [{"id": "nb1", "title": "A"}], "has_more": True}
)
page2 = _mock_response(
json_data={"items": [{"id": "nb2", "title": "B"}], "has_more": False}
)
with patch("requests.get") as mock_get:
mock_get.side_effect = [page1, page2]
result = client.list_notebooks()
assert len(result) == 2
assert {nb["id"] for nb in result} == {"nb1", "nb2"}
def test_list_notebooks_connection_error(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.side_effect = requests.exceptions.ConnectionError()
with pytest.raises(JoplinError, match="Joplin"):
client.list_notebooks()
# ---------------------------------------------------------------------------
# get_or_create_notebook
# ---------------------------------------------------------------------------
class TestGetOrCreateNotebook:
def test_returns_existing_notebook_id(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={
"items": [{"id": "nb-existing", "title": "ChatGPT - No Project"}],
"has_more": False,
}
)
nb_id = client.get_or_create_notebook("ChatGPT - No Project")
assert nb_id == "nb-existing"
def test_creates_new_notebook_when_not_found(self):
client = _make_client()
with patch("requests.get") as mock_get, patch("requests.post") as mock_post:
mock_get.return_value = _mock_response(
json_data={"items": [], "has_more": False}
)
mock_post.return_value = _mock_response(
json_data={"id": "nb-new", "title": "ChatGPT - New Project"}
)
nb_id = client.get_or_create_notebook("ChatGPT - New Project")
assert nb_id == "nb-new"
mock_post.assert_called_once()
def test_caches_notebook_after_first_load(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={
"items": [{"id": "nb1", "title": "Claude - No Project"}],
"has_more": False,
}
)
# Call twice — GET /folders should only happen once
client.get_or_create_notebook("Claude - No Project")
client.get_or_create_notebook("Claude - No Project")
assert mock_get.call_count == 1
# ---------------------------------------------------------------------------
# create_note
# ---------------------------------------------------------------------------
class TestCreateNote:
def test_create_note_returns_id(self):
client = _make_client()
with patch("requests.post") as mock_post:
mock_post.return_value = _mock_response(
json_data={"id": "note-123", "title": "My Note"}
)
note_id = client.create_note("My Note", "Note body", "nb-456")
assert note_id == "note-123"
_, kwargs = mock_post.call_args
assert kwargs["json"]["title"] == "My Note"
assert kwargs["json"]["body"] == "Note body"
assert kwargs["json"]["parent_id"] == "nb-456"
def test_create_note_connection_error(self):
client = _make_client()
with patch("requests.post") as mock_post:
mock_post.side_effect = requests.exceptions.ConnectionError()
with pytest.raises(JoplinError, match="Joplin"):
client.create_note("Title", "Body", "nb-id")
def test_create_note_http_error(self):
client = _make_client()
with patch("requests.post") as mock_post:
mock_post.return_value = _mock_response(status_code=401)
with pytest.raises(JoplinError):
client.create_note("Title", "Body", "nb-id")
# ---------------------------------------------------------------------------
# update_note
# ---------------------------------------------------------------------------
class TestUpdateNote:
def test_update_note_calls_put(self):
client = _make_client()
with patch("requests.put") as mock_put:
mock_put.return_value = _mock_response(json_data={"id": "note-123"})
client.update_note("note-123", "New Title", "New Body")
mock_put.assert_called_once()
_, kwargs = mock_put.call_args
assert kwargs["json"]["title"] == "New Title"
assert kwargs["json"]["body"] == "New Body"
def test_update_note_connection_error(self):
client = _make_client()
with patch("requests.put") as mock_put:
mock_put.side_effect = requests.exceptions.ConnectionError()
with pytest.raises(JoplinError, match="Joplin"):
client.update_note("note-id", "Title", "Body")
def test_update_note_http_error(self):
client = _make_client()
with patch("requests.put") as mock_put:
mock_put.return_value = _mock_response(status_code=404)
with pytest.raises(JoplinError):
client.update_note("note-id", "Title", "Body")
+48 -3
View File
@@ -13,15 +13,17 @@ class TestChatGPTNormalization:
def _get_provider(self): def _get_provider(self):
from src.providers.chatgpt import ChatGPTProvider from src.providers.chatgpt import ChatGPTProvider
import unittest.mock as mock
# Bypass __init__ token check # Bypass __init__ token check
p = ChatGPTProvider.__new__(ChatGPTProvider) p = ChatGPTProvider.__new__(ChatGPTProvider)
import requests import requests
p._session = requests.Session() p._session = requests.Session()
p._org_id = None p._org_id = None
p._project_ids = []
p._project_map = {}
p._project_name_cache = {}
return p return p
def test_normalizes_with_project(self): def test_normalizes_conversation(self):
raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text()) raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text())
p = self._get_provider() p = self._get_provider()
result = p.normalize_conversation(raw) result = p.normalize_conversation(raw)
@@ -29,7 +31,8 @@ class TestChatGPTNormalization:
assert result["id"] == "chatgpt-conv-001" assert result["id"] == "chatgpt-conv-001"
assert result["title"] == "Python Async Tutorial" assert result["title"] == "Python Async Tutorial"
assert result["provider"] == "chatgpt" assert result["provider"] == "chatgpt"
assert result["project"] == "Learning Python" # No entry in _project_map → project is None
assert result["project"] is None
assert result["created_at"] != "" assert result["created_at"] != ""
assert result["updated_at"] != "" assert result["updated_at"] != ""
assert isinstance(result["messages"], list) assert isinstance(result["messages"], list)
@@ -42,6 +45,15 @@ class TestChatGPTNormalization:
assert result["project"] is None assert result["project"] is None
assert result["id"] == "chatgpt-conv-002" assert result["id"] == "chatgpt-conv-002"
def test_normalizes_with_project_from_map(self):
"""Project name from _project_map (populated by fetch_all_conversations) flows through."""
raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text())
p = self._get_provider()
p._project_map["chatgpt-conv-001"] = "My Research Project"
result = p.normalize_conversation(raw)
assert result["project"] == "My Research Project"
def test_extracts_text_messages(self): def test_extracts_text_messages(self):
raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text()) raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text())
p = self._get_provider() p = self._get_provider()
@@ -63,6 +75,39 @@ class TestChatGPTNormalization:
for r in caplog.records for r in caplog.records
) )
def test_model_editable_context_included_without_warning(self, caplog):
"""model_editable_context messages (project instructions) should be included, not warned about."""
import logging
conv = {
"id": "test-conv-mec",
"title": "Test",
"create_time": 1700000000.0,
"update_time": 1700000001.0,
"mapping": {
"root": {"id": "root", "message": None, "parent": None, "children": ["msg1"]},
"msg1": {
"id": "msg1",
"message": {
"id": "msg1",
"author": {"role": "user"},
"content": {
"content_type": "model_editable_context",
"parts": ["These are the project instructions."],
},
"create_time": 1700000001.0,
"status": "finished_successfully",
},
"parent": "root",
"children": [],
},
},
}
p = self._get_provider()
with caplog.at_level(logging.WARNING):
result = p.normalize_conversation(conv)
assert any(m["content"] == "These are the project instructions." for m in result["messages"])
assert not any("model_editable_context" in r.message for r in caplog.records)
def test_message_roles_are_valid(self): def test_message_roles_are_valid(self):
raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text()) raw = json.loads((FIXTURES / "chatgpt_conversation.json").read_text())
p = self._get_provider() p = self._get_provider()
+147
View File
@@ -0,0 +1,147 @@
"""Tests for src/utils.py — filename generation, path building, redaction."""
from pathlib import Path
import pytest
from src.utils import (
build_export_path,
format_token_status,
generate_filename,
redact_secrets,
)
class TestGenerateFilename:
def test_basic_format(self):
name = generate_filename("Hello World", "abc12345def", "2024-06-10T14:00:00Z")
assert name == "2024-06-10_hello-world_abc12345.md"
def test_special_chars_slugified(self):
# T-36: titles with punctuation must produce safe, OS-compatible filenames
name = generate_filename("What's this?! A test.", "abc12345", "2024-06-01T00:00:00Z")
assert "?" not in name
assert "!" not in name
assert "'" not in name
assert " " not in name
assert name.startswith("2024-06-01_")
assert name.endswith("_abc12345.md")
def test_unicode_chars_handled(self):
name = generate_filename("Héllo Wörld", "abc12345", "2024-06-01T00:00:00Z")
assert " " not in name
assert name.endswith("_abc12345.md")
def test_empty_title_becomes_untitled(self):
name = generate_filename("", "abc12345", "2024-06-01T00:00:00Z")
assert "untitled" in name
def test_id_truncated_to_8_chars(self):
name = generate_filename("Test", "abcdefghijklmnop", "2024-06-01T00:00:00Z")
assert name.endswith("_abcdefgh.md")
def test_long_title_truncated(self):
long_title = "a" * 200
name = generate_filename(long_title, "abc12345", "2024-06-01T00:00:00Z")
# Slug is capped at 60 chars by max_length
slug_part = name.split("_")[1]
assert len(slug_part) <= 60
def test_date_comes_from_created_at(self):
name = generate_filename("Test", "abc12345", "2023-11-25T00:00:00Z")
assert name.startswith("2023-11-25_")
class TestBuildExportPath:
def test_default_structure_provider_project_year(self):
path = build_export_path(
Path("/exports"), "claude", "my-project", "2024-06-01T00:00:00Z", "file.md"
)
assert str(path) == "/exports/claude/my-project/2024/file.md"
def test_no_project_uses_no_project_slug(self):
path = build_export_path(
Path("/exports"), "chatgpt", None, "2024-06-01T00:00:00Z", "file.md"
)
assert "no-project" in str(path)
def test_provider_project_structure_omits_year(self):
path = build_export_path(
Path("/exports"), "claude", "proj", "2024-06-01T00:00:00Z", "file.md",
structure="provider/project",
)
assert "2024" not in str(path)
assert "proj" in str(path)
def test_provider_year_structure_omits_project(self):
path = build_export_path(
Path("/exports"), "claude", "proj", "2024-06-01T00:00:00Z", "file.md",
structure="provider/year",
)
assert "proj" not in str(path)
assert "2024" in str(path)
def test_project_name_with_spaces_is_slugified(self):
path = build_export_path(
Path("/exports"), "claude", "My Project Name!", "2024-06-01T00:00:00Z", "file.md"
)
assert " " not in str(path)
assert "!" not in str(path)
class TestRedactSecrets:
def test_token_value_redacted(self):
data = {"token": "supersecret"}
result = redact_secrets(data)
assert result["token"] == "[REDACTED]"
def test_session_key_redacted(self):
result = redact_secrets({"sessionKey": "abc123"})
assert result["sessionKey"] == "[REDACTED]"
def test_non_sensitive_key_unchanged(self):
result = redact_secrets({"title": "My Chat", "id": "abc123"})
assert result["title"] == "My Chat"
assert result["id"] == "abc123"
def test_nested_dict_redacted(self):
data = {"user": {"token": "secret", "name": "Alice"}}
result = redact_secrets(data)
assert result["user"]["token"] == "[REDACTED]"
assert result["user"]["name"] == "Alice"
def test_list_of_dicts(self):
data = [{"password": "p@ss"}, {"title": "chat"}]
result = redact_secrets(data)
assert result[0]["password"] == "[REDACTED]"
assert result[1]["title"] == "chat"
class TestFormatTokenStatus:
def test_none_token_returns_not_set(self):
assert format_token_status(None) == "[NOT SET]"
def test_empty_token_returns_not_set(self):
assert format_token_status("") == "[NOT SET]"
def test_set_token_no_expiry(self):
assert format_token_status("sometoken") == "[SET]"
def test_expired_token(self):
from datetime import datetime, timezone, timedelta
expiry = datetime.now(tz=timezone.utc) - timedelta(days=1)
result = format_token_status("tok", expiry)
assert "EXPIRED" in result
def test_expiring_today_shows_hours(self):
from datetime import datetime, timezone, timedelta
expiry = datetime.now(tz=timezone.utc) + timedelta(hours=3)
result = format_token_status("tok", expiry)
assert "expires in" in result
assert "h" in result
def test_expiring_in_days(self):
from datetime import datetime, timezone, timedelta
expiry = datetime.now(tz=timezone.utc) + timedelta(days=10, hours=12)
result = format_token_status("tok", expiry)
assert "10 days" in result