# Planned Future Work Items completed in each release are moved to the changelog. Items here are designed for but not yet implemented. The codebase is structured to make each of these additions straightforward. **Completed:** - v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output - v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation) - v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks - v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders - v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6) --- # Roadmap (decided 2026-06-12) Priorities reflect the tool's primary purpose — a trustworthy backup so that conversation data is not lost if a provider account is ever closed — plus the day-to-day friction of weekly ChatGPT token refresh, and the long-term goal of running headless as a StartOS service. **Now (in order):** 1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0** (full-archive re-export still pending; see entry below) 2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0** (note: runs on whichever machine hosts the browser — Brave is not on this Linux box) 3. ~~Claude Code session provider~~ — **shipped in v0.6.0** 4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0** **Soon, but later:** 5. ~~Binary content downloads~~ — **shipped in v0.6.0** 6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0** 7. Scheduled / watch mode → StartOS service packaging (long-term destination) **Deprioritized** (entries kept at the bottom of this file; revisit on demand): `--force` flags, per-conversation cache reset, official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian output, token expiry notifications (largely superseded by cookie auto-extraction), search command. Additional web providers (Gemini/Grok/Perplexity) are explicitly out of scope — no significant usage to archive. --- ## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0 **Implemented 2026-06-12** as designed below, with one scoping correction from live recon: the retrieval dumps are NOT flagged `is_visually_hidden_from_conversation` — `author.name == "file_search"` is the discriminator (the hidden flag only marks Custom Instructions and small system stubs). Verified live: worst files shrink 93% (524KB → 36KB). Remaining step: full-archive re-export (`cache --clear` + `export`) + Joplin re-sync, at the user's chosen time. **Problem (measured 2026-06-12 against a full fresh export):** 45% of the entire 11.2MB archive (260 files) is tool-role messages; 29 files are majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64 messages) is ChatGPT's file-retrieval tool re-injecting the full text of the user's own attached documents, the same files dumped dozens of times per conversation. These messages were invisible in the ChatGPT web UI; they appear in exports because v0.4.0 lifted the role filter to fix silent data loss. Custom Instructions hidden-context blocks are a minor secondary case (~2KB, once per conversation) — the originally planned `EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst file contains zero hidden-context blocks. **Fix: collapse, don't drop** (consistent with the no-silent-drop rule): - `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default `placeholder`, plus a `--hidden-content` CLI override on `export`. - `placeholder` renders affected messages as one line with type and size: `> 🔧 Tool output (file_search, 24KB) — omitted (EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive roughly halves; worst files shrink ~90%. - **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw `author.name` (`file_search`, `myfiles_browser`, …) in the API response (the rendered Markdown only shows a generic "🔧 Tool" label, so the decision must happen in the provider, not the renderer); (b) messages flagged `is_visually_hidden_from_conversation`, including `user_editable_context` / `model_editable_context` (Custom Instructions) — subsumes the old suppress-hidden-context idea. - **Scope — keep at full size:** code-execution `tool_result` blocks and web-search results; those are usually content the user wants. - Count collapsed messages in the post-export summary so the omission stays visible (mirror the LossReport presentation, but as intentional policy, not loss). Re-export workflow after shipping: `cache --clear` + `export` (same as the v0.4.0 migration). ## 2. Brave Cookie Auto-Extraction — SHIPPED v0.6.0 **Implemented 2026-06-12**: `auth --from-browser [browser]` (`src/browser_tokens.py`, browser-cookie3 dependency, defaults to brave; chrome/chromium/edge/firefox also supported). Extracted tokens are validated against the live API before `.env` is touched. **Discovered during implementation: Brave is not installed on this machine** — the logged-in browser lives elsewhere, so the flag only helps when the exporter runs on that machine (verified the graceful-failure path here; the happy path is covered by mocked tests). This strengthens the case for the StartOS token-push mechanism (entry 8). Pull session tokens directly from the Brave browser profile instead of the weekly manual DevTools dance. `auth` gains an "extract from browser" path (with the manual flow kept as fallback); a later iteration could let `doctor` or a 401 handler suggest/perform a re-extract automatically. - Brave on Linux follows the Chromium pattern: cookies SQLite at `~/.config/BraveSoftware/Brave-Browser/Default/Cookies`, values AES-128-CBC encrypted with a key derived from the OS keyring secret ("Brave Safe Storage" via SecretService/kwallet, or the `peanuts` v10 fallback). Well-trodden territory — evaluate a small dependency (`browser_cookie3`-style) vs. a focused in-repo implementation. - Extract `__Secure-next-auth.session-token.0` / `.1` from `chatgpt.com` and `sessionKey` from `claude.ai`; write to `.env` through the existing auth-wizard write path (tokens never echoed). - Brave must be the user's logged-in browser (it is); detect a locked / unreadable cookie DB (Brave running with the DB exclusively locked) and fall back to the manual flow with a clear message. - Does **not** solve headless/StartOS — no browser on the server. That needs a token-push mechanism, tracked under the StartOS entry. ## 3. Claude Code Session Provider — SHIPPED v0.6.0 **Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`, `--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings during implementation: `isSidechain` records are subagent transcripts (skipped), `isMeta` marks harness-generated user records (skipped), and listing/normalized `updated_at` must both use file mtime or the cache would re-export every session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown. Archive local Claude Code session transcripts. No tokens, no rate limits, no ToS risk — the data is already on disk but lives in a single JSONL per session that Claude Code may clean up, and it contains deliverables (reviews, plans, analyses) that exist nowhere else. Decisions (2026-06-12): - **Rendering: prose-only.** Keep user prompts and assistant text (including full deliverable write-ups); collapse tool activity to one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2), 86KB — omitted`); **exclude thinking blocks**. - **Joplin: sync enabled.** Each coding project becomes a notebook nested under an **"AI-Claude"** parent notebook (nested-notebook support shipped in v0.5.0). Data facts (measured 2026-06-12): - Source: `~/.claude/projects//.jsonl`. Currently 29 sessions, 18.8MB total, largest 3.8MB. - Representative 2.1MB session: tool_result 436KB, tool_use 122KB, thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool activity is what makes these exports readable. - Record types: `user` / `assistant` (Anthropic-style `message.content` block arrays) plus harness records: `ai-title` (use for note title and filename slug), `last-prompt`, `file-history-snapshot`, `attachment`, `permission-mode`, `system` (skip). Strip harness noise from user messages (``, `` blocks). Implementation shape: new `src/providers/claude_code.py` implementing the `BaseProvider` interface — `list_conversations` scans project dirs, `get_conversation` parses the JSONL, `normalize_conversation` maps onto the existing block schema (content is already block-shaped: text / tool_use / tool_result / thinking). Incremental sync via file mtime/size recorded in the existing manifest. Project name derives from the munged cwd dirname. ## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0 **Implemented 2026-06-12** as designed below, plus an empty-manifest guard (refuses to prune right after `cache --clear`). First live run removed 420 stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports manifest↔disk integrity (293/293 after the run). A backup is only trustworthy if the on-disk tree matches the manifest. Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`) and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in v0.4.1) sit alongside current exports and would double-sync into Joplin. - `prune` command: delete export files not referenced by the manifest. `--dry-run` (default off, but always print the list before deleting) shows what would be removed and why (old layout / orphan / unknown). - `doctor` extension: verify every manifest entry's `file_path` exists on disk; report missing files (re-export candidates) and unreferenced files (prune candidates). ## 5. Binary Content Downloads — SHIPPED v0.6.0 **Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`, `download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin `create_resource` + `upload_media_and_rewrite`). Live recon settled the download mechanism: `GET /backend-api/files/{id}/download` returns a signed `download_url`; a second GET yields the bytes (works for user uploads; older AI-generated images 404 — expired server-side, handled gracefully). Asset refs come in three shapes — `sediment://file_…`, `sediment://#file_…#p_N.png` (generated), `file-service://…`. Archive scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so images-only is the default and audio is opt-in via `all`. Original notes below. **Priority note (2026-06-12): "later, but soon" — under the backup-if-account-closes goal, embedded images are part of the data that would be lost; placeholders alone don't preserve them.** v0.4.0 ships placeholders for images and audio assets but does not download the binary content. The `_safe_fence`-wrapped placeholders include the asset reference (`sediment://...` or `file-service://...`), MIME type, size, and duration where available; the actual bytes are not preserved. Next steps: - Download attached images alongside the Markdown export, save under a `media/` sibling directory with a stable filename derived from the asset reference. - Replace `image_placeholder` rendering with an inline `![](relative/path)` reference once the file is on disk. - Joplin integration: upload binaries as Joplin resources via `POST /resources`, rewrite the rendered Markdown to use `:/resourceId` references, and track the resource ID in the cache manifest so re-syncs stay idempotent. - DALL-E images on the assistant side: not observed in this user's data; the code path exists (`source = "model_generated"`) but is untested. The block-level schema is already in place — only the file-fetch + rewrite layer needs to be added. See the `image_placeholder` and `file_placeholder` block definitions in `src/blocks.py`. ## 6. Per-Session Download Limiter — SHIPPED v0.6.0 **Implemented 2026-06-12** as designed below: `--max-conversations N` / `MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and `REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`. Verified live: a 2-pending run with cap 1 exported one, deferred one, and the re-run picked it up. Cap how many conversations are downloaded in a single `export` run so the tool never hammers the ChatGPT/Claude internal APIs with a large burst — most importantly on the very first export, which otherwise fetches the entire conversation history in one session. Because every run is resumable (the manifest records each conversation immediately), a capped run simply exports the first N pending conversations and the next run picks up where it left off. This keeps traffic looking like a human-paced session rather than a scraper, reducing the risk of rate limiting or account flags. Two complementary pieces: 1. **Session cap** — `--max-conversations N` flag (and `MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the `export` command, slice the pending list after the cache filter: `to_export = to_export[:n]`. On exit, print exported-vs-remaining counts (reuse the message format from the 429 early-exit path) and remind the user to re-run to continue. 2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small random jitter) slept between per-conversation detail fetches in `BaseProvider`, so even a capped run doesn't fire requests back-to-back. The existing 429 backoff in `_request` stays as the reactive safety net. Note: the conversation *listing* (paginated, 100/page) still runs in full each time so the cache comparison works — the cap applies to the heavy per-conversation detail fetches, which dominate request volume. Together with watch mode, this is a stepping stone to the StartOS service: a scheduled, capped, politely-paced export is the traffic profile a headless deployment needs. ## 7. Scheduled / Watch Mode Add a `watch` command (or cron integration helper) to run exports automatically on a schedule: ```bash python -m src.main watch --interval 6h # poll every 6 hours ``` This would run `export` + `joplin` in sequence, then sleep. Alternatively, provide a `cron` command that prints the correct crontab line for the user's setup. Implementation: simple loop with `time.sleep()`, or emit a crontab entry string that calls the export and joplin commands in sequence. A `--once` flag would do a single run then exit (useful for cron itself). Stepping stone to the StartOS service (below). ## 8. StartOS Service Packaging (long-term destination) Run the exporter headless as a StartOS service: scheduled export + sync to a Joplin Server instance, so the backup happens without a desktop in the loop. - Building blocks needed first: watch mode (#7), session limiter (#6), and a Joplin Server (vs. desktop Web Clipper) sync target. - **Hard problem, flagged early:** session-token freshness without a browser. ChatGPT tokens last ~7 days; there is no browser on the server to re-extract from. Candidate mechanisms: a service config action the user pastes fresh tokens into periodically, or a future browser extension that pushes tokens to the server. Cookie auto-extraction (#2) only solves the desktop case. --- # Deprioritized Kept for reference; not on the active roadmap. ## Export `--force` Flag — SHIPPED v0.6.0 Implemented 2026-06-12: `export --force` passes `force=True` to `cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that preserves Joplin links across re-exports, so a forced re-render + `joplin` updates existing notes instead of duplicating them. ## Joplin `--force` Flag Similarly, add `--force` to the `joplin` command to re-sync all cached conversations to Joplin regardless of whether they've been synced before. Useful after making formatting changes to the Markdown exporter. Implementation: in `get_joplin_pending()`, return all entries that have a `file_path` when `force=True`, ignoring `joplin_synced_at`. ## Per-Conversation Cache Reset Add `cache --reset --conversation ` to force re-export or re-sync of a single conversation without clearing the entire provider cache. Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and delete the entry, then re-run export. ## Official API Fallback If the unofficial internal web API approach breaks, migrate to official export file parsing as a fallback: - ChatGPT: parse `conversations.json` from Settings → Export Data - Claude: parse `conversations.json` from Settings → Privacy → Export Data The `BaseProvider` abstract class is intentionally designed so that a `FileProvider` subclass can implement the same interface (`list_conversations`, `get_conversation`, `normalize_conversation`) without any changes to cache, exporters, or CLI code. To add this: implement `src/providers/file_chatgpt.py` and `src/providers/file_claude.py`, then add `--input-file` flag to the export command to accept a pre-downloaded export ZIP or JSON. Deprioritized 2026-06-12: the official ChatGPT export does not cover what this user needs (project data), so it isn't a real fallback here. ## Reclassify o1/o3 Reasoning Subparts v0.4.0 leaves dict parts inside `text` content_type messages with shape `{"summary": ..., "content": ...}` rendered as plain text (defensive — the shape was inferred from a code comment, not captured live). Once a real reasoning conversation is captured, reclassify these as `thinking` blocks. ## Obsidian Vault Output Add an `obsidian` command (or `--target obsidian` flag) to sync exported conversations into an Obsidian vault directory. The current Markdown format is already largely compatible; the main differences are: - Obsidian uses YAML frontmatter `properties` (same format, already supported) - Tags should use `#tag` inline or `tags:` list in frontmatter (already done) - Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian supports both Implementation: the existing `MarkdownExporter` output is already valid in Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply copy files to the vault directory and maintain a flat or nested folder structure matching the user's Obsidian setup. No API needed — just file I/O. ## Token Expiry Notifications Proactively warn when a token is close to expiry (within 48h for ChatGPT), rather than only surfacing the warning at startup. Options: - Add an `expiry` subcommand that prints token status and exits non-zero if any token is expired or expiring soon (useful in scripts/cron) - Send a desktop notification via `notify-send` (Linux) or `osascript` (macOS) when a token is within 24h of expiry Largely superseded by Brave cookie auto-extraction (#2): once refresh is one command (or automatic), expiry warnings matter much less. ## Search Command Add a `search` command to full-text search across all exported Markdown files: ```bash python -m src.main search "kubernetes ingress" python -m src.main search "kubernetes ingress" --provider claude --project devops ``` Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with conversation title, date, and a snippet. No index needed — Markdown files are small enough to grep directly.