399 lines
20 KiB
Markdown
399 lines
20 KiB
Markdown
# Planned Future Work
|
||
|
||
Items completed in each release are moved to the changelog. Items here are
|
||
designed for but not yet implemented. The codebase is structured to make each
|
||
of these additions straightforward.
|
||
|
||
**Completed:**
|
||
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
||
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
||
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
|
||
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
||
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
|
||
|
||
---
|
||
|
||
# Roadmap (decided 2026-06-12)
|
||
|
||
Priorities reflect the tool's primary purpose — a trustworthy backup so that
|
||
conversation data is not lost if a provider account is ever closed — plus the
|
||
day-to-day friction of weekly ChatGPT token refresh, and the long-term goal
|
||
of running headless as a StartOS service.
|
||
|
||
**Now (in order):**
|
||
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
|
||
(full-archive re-export still pending; see entry below)
|
||
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0** (note: runs on
|
||
whichever machine hosts the browser — Brave is not on this Linux box)
|
||
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
|
||
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
|
||
|
||
**Soon, but later:**
|
||
|
||
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
|
||
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
|
||
7. Scheduled / watch mode → StartOS service packaging (long-term destination)
|
||
|
||
**Deprioritized** (entries kept at the bottom of this file; revisit on
|
||
demand): `--force` flags, per-conversation cache reset, official export-ZIP
|
||
fallback, o1/o3 reasoning reclassification, Obsidian output, token expiry
|
||
notifications (largely superseded by cookie auto-extraction), search command.
|
||
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
|
||
scope — no significant usage to archive.
|
||
|
||
---
|
||
|
||
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
|
||
|
||
**Implemented 2026-06-12** as designed below, with one scoping correction
|
||
from live recon: the retrieval dumps are NOT flagged
|
||
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
|
||
the discriminator (the hidden flag only marks Custom Instructions and small
|
||
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
|
||
Remaining step: full-archive re-export (`cache --clear` + `export`) +
|
||
Joplin re-sync, at the user's chosen time.
|
||
|
||
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
|
||
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
|
||
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
|
||
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
|
||
the user's own attached documents, the same files dumped dozens of times
|
||
per conversation. These messages were invisible in the ChatGPT web UI;
|
||
they appear in exports because v0.4.0 lifted the role filter to fix silent
|
||
data loss. Custom Instructions hidden-context blocks are a minor secondary
|
||
case (~2KB, once per conversation) — the originally planned
|
||
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
|
||
file contains zero hidden-context blocks.
|
||
|
||
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
|
||
|
||
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
|
||
`placeholder`, plus a `--hidden-content` CLI override on `export`.
|
||
- `placeholder` renders affected messages as one line with type and size:
|
||
`> 🔧 Tool output (file_search, 24KB) — omitted
|
||
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
|
||
roughly halves; worst files shrink ~90%.
|
||
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
|
||
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
|
||
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
|
||
decision must happen in the provider, not the renderer); (b) messages
|
||
flagged `is_visually_hidden_from_conversation`, including
|
||
`user_editable_context` / `model_editable_context` (Custom
|
||
Instructions) — subsumes the old suppress-hidden-context idea.
|
||
- **Scope — keep at full size:** code-execution `tool_result` blocks and
|
||
web-search results; those are usually content the user wants.
|
||
- Count collapsed messages in the post-export summary so the omission
|
||
stays visible (mirror the LossReport presentation, but as intentional
|
||
policy, not loss).
|
||
|
||
Re-export workflow after shipping: `cache --clear` + `export` (same as
|
||
the v0.4.0 migration).
|
||
|
||
## 2. Brave Cookie Auto-Extraction — SHIPPED v0.6.0
|
||
|
||
**Implemented 2026-06-12**: `auth --from-browser [browser]`
|
||
(`src/browser_tokens.py`, browser-cookie3 dependency, defaults to brave;
|
||
chrome/chromium/edge/firefox also supported). Extracted tokens are validated
|
||
against the live API before `.env` is touched. **Discovered during
|
||
implementation: Brave is not installed on this machine** — the logged-in
|
||
browser lives elsewhere, so the flag only helps when the exporter runs on
|
||
that machine (verified the graceful-failure path here; the happy path is
|
||
covered by mocked tests). This strengthens the case for the StartOS
|
||
token-push mechanism (entry 8).
|
||
|
||
Pull session tokens directly from the Brave browser profile instead of the
|
||
weekly manual DevTools dance. `auth` gains an "extract from browser" path
|
||
(with the manual flow kept as fallback); a later iteration could let
|
||
`doctor` or a 401 handler suggest/perform a re-extract automatically.
|
||
|
||
- Brave on Linux follows the Chromium pattern: cookies SQLite at
|
||
`~/.config/BraveSoftware/Brave-Browser/Default/Cookies`, values
|
||
AES-128-CBC encrypted with a key derived from the OS keyring secret
|
||
("Brave Safe Storage" via SecretService/kwallet, or the `peanuts` v10
|
||
fallback). Well-trodden territory — evaluate a small dependency
|
||
(`browser_cookie3`-style) vs. a focused in-repo implementation.
|
||
- Extract `__Secure-next-auth.session-token.0` / `.1` from `chatgpt.com`
|
||
and `sessionKey` from `claude.ai`; write to `.env` through the existing
|
||
auth-wizard write path (tokens never echoed).
|
||
- Brave must be the user's logged-in browser (it is); detect a locked /
|
||
unreadable cookie DB (Brave running with the DB exclusively locked) and
|
||
fall back to the manual flow with a clear message.
|
||
- Does **not** solve headless/StartOS — no browser on the server. That
|
||
needs a token-push mechanism, tracked under the StartOS entry.
|
||
|
||
## 3. Claude Code Session Provider — SHIPPED v0.6.0
|
||
|
||
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
|
||
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
|
||
during implementation: `isSidechain` records are subagent transcripts (skipped),
|
||
`isMeta` marks harness-generated user records (skipped), and listing/normalized
|
||
`updated_at` must both use file mtime or the cache would re-export every
|
||
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
|
||
|
||
Archive local Claude Code session transcripts. No tokens, no rate limits,
|
||
no ToS risk — the data is already on disk but lives in a single JSONL per
|
||
session that Claude Code may clean up, and it contains deliverables
|
||
(reviews, plans, analyses) that exist nowhere else.
|
||
|
||
Decisions (2026-06-12):
|
||
- **Rendering: prose-only.** Keep user prompts and assistant text
|
||
(including full deliverable write-ups); collapse tool activity to
|
||
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
|
||
86KB — omitted`); **exclude thinking blocks**.
|
||
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
|
||
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
|
||
in v0.5.0).
|
||
|
||
Data facts (measured 2026-06-12):
|
||
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
|
||
Currently 29 sessions, 18.8MB total, largest 3.8MB.
|
||
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
|
||
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
|
||
activity is what makes these exports readable.
|
||
- Record types: `user` / `assistant` (Anthropic-style `message.content`
|
||
block arrays) plus harness records: `ai-title` (use for note title and
|
||
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
|
||
`permission-mode`, `system` (skip). Strip harness noise from user
|
||
messages (`<local-command-caveat>`, `<command-name>` blocks).
|
||
|
||
Implementation shape: new `src/providers/claude_code.py` implementing the
|
||
`BaseProvider` interface — `list_conversations` scans project dirs,
|
||
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
|
||
existing block schema (content is already block-shaped: text / tool_use /
|
||
tool_result / thinking). Incremental sync via file mtime/size recorded in
|
||
the existing manifest. Project name derives from the munged cwd dirname.
|
||
|
||
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
|
||
|
||
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
|
||
(refuses to prune right after `cache --clear`). First live run removed 420
|
||
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
|
||
manifest↔disk integrity (293/293 after the run).
|
||
|
||
A backup is only trustworthy if the on-disk tree matches the manifest.
|
||
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
|
||
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
|
||
v0.4.1) sit alongside current exports and would double-sync into Joplin.
|
||
|
||
- `prune` command: delete export files not referenced by the manifest.
|
||
`--dry-run` (default off, but always print the list before deleting)
|
||
shows what would be removed and why (old layout / orphan / unknown).
|
||
- `doctor` extension: verify every manifest entry's `file_path` exists on
|
||
disk; report missing files (re-export candidates) and unreferenced files
|
||
(prune candidates).
|
||
|
||
## 5. Binary Content Downloads — SHIPPED v0.6.0
|
||
|
||
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
|
||
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
|
||
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
|
||
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
|
||
`download_url`; a second GET yields the bytes (works for user uploads;
|
||
older AI-generated images 404 — expired server-side, handled gracefully).
|
||
Asset refs come in three shapes — `sediment://file_…`,
|
||
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
|
||
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
|
||
images-only is the default and audio is opt-in via `all`.
|
||
|
||
Original notes below.
|
||
|
||
**Priority note (2026-06-12): "later, but soon" — under the
|
||
backup-if-account-closes goal, embedded images are part of the data that
|
||
would be lost; placeholders alone don't preserve them.**
|
||
|
||
v0.4.0 ships placeholders for images and audio assets but does not download
|
||
the binary content. The `_safe_fence`-wrapped placeholders include the asset
|
||
reference (`sediment://...` or `file-service://...`), MIME type, size, and
|
||
duration where available; the actual bytes are not preserved.
|
||
|
||
Next steps:
|
||
- Download attached images alongside the Markdown export, save under a
|
||
`media/` sibling directory with a stable filename derived from the asset
|
||
reference.
|
||
- Replace `image_placeholder` rendering with an inline ``
|
||
reference once the file is on disk.
|
||
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
|
||
rewrite the rendered Markdown to use `:/resourceId` references, and track
|
||
the resource ID in the cache manifest so re-syncs stay idempotent.
|
||
- DALL-E images on the assistant side: not observed in this user's data; the
|
||
code path exists (`source = "model_generated"`) but is untested.
|
||
|
||
The block-level schema is already in place — only the file-fetch + rewrite
|
||
layer needs to be added. See the `image_placeholder` and `file_placeholder`
|
||
block definitions in `src/blocks.py`.
|
||
|
||
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
|
||
|
||
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
|
||
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
|
||
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
|
||
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
|
||
the re-run picked it up.
|
||
|
||
Cap how many conversations are downloaded in a single `export` run so the
|
||
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
|
||
most importantly on the very first export, which otherwise fetches the
|
||
entire conversation history in one session. Because every run is resumable
|
||
(the manifest records each conversation immediately), a capped run simply
|
||
exports the first N pending conversations and the next run picks up where
|
||
it left off. This keeps traffic looking like a human-paced session rather
|
||
than a scraper, reducing the risk of rate limiting or account flags.
|
||
|
||
Two complementary pieces:
|
||
|
||
1. **Session cap** — `--max-conversations N` flag (and
|
||
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
|
||
`export` command, slice the pending list after the cache filter:
|
||
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
|
||
counts (reuse the message format from the 429 early-exit path) and
|
||
remind the user to re-run to continue.
|
||
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
|
||
random jitter) slept between per-conversation detail fetches in
|
||
`BaseProvider`, so even a capped run doesn't fire requests
|
||
back-to-back. The existing 429 backoff in `_request` stays as the
|
||
reactive safety net.
|
||
|
||
Note: the conversation *listing* (paginated, 100/page) still runs in full
|
||
each time so the cache comparison works — the cap applies to the heavy
|
||
per-conversation detail fetches, which dominate request volume.
|
||
|
||
Together with watch mode, this is a stepping stone to the StartOS service:
|
||
a scheduled, capped, politely-paced export is the traffic profile a
|
||
headless deployment needs.
|
||
|
||
## 7. Scheduled / Watch Mode
|
||
|
||
Add a `watch` command (or cron integration helper) to run exports automatically
|
||
on a schedule:
|
||
|
||
```bash
|
||
python -m src.main watch --interval 6h # poll every 6 hours
|
||
```
|
||
|
||
This would run `export` + `joplin` in sequence, then sleep. Alternatively,
|
||
provide a `cron` command that prints the correct crontab line for the user's
|
||
setup.
|
||
|
||
Implementation: simple loop with `time.sleep()`, or emit a crontab entry
|
||
string that calls the export and joplin commands in sequence. A `--once`
|
||
flag would do a single run then exit (useful for cron itself).
|
||
|
||
Stepping stone to the StartOS service (below).
|
||
|
||
## 8. StartOS Service Packaging (long-term destination)
|
||
|
||
Run the exporter headless as a StartOS service: scheduled export + sync to
|
||
a Joplin Server instance, so the backup happens without a desktop in the
|
||
loop.
|
||
|
||
- Building blocks needed first: watch mode (#7), session limiter (#6),
|
||
and a Joplin Server (vs. desktop Web Clipper) sync target.
|
||
- **Hard problem, flagged early:** session-token freshness without a
|
||
browser. ChatGPT tokens last ~7 days; there is no browser on the server
|
||
to re-extract from. Candidate mechanisms: a service config action the
|
||
user pastes fresh tokens into periodically, or a future browser
|
||
extension that pushes tokens to the server. Cookie auto-extraction (#2)
|
||
only solves the desktop case.
|
||
|
||
---
|
||
|
||
# Deprioritized
|
||
|
||
Kept for reference; not on the active roadmap.
|
||
|
||
## Export `--force` Flag
|
||
|
||
Add `--force` to the `export` command to re-export already-cached conversations
|
||
without permanently clearing the entire manifest. Useful for re-generating files
|
||
after changing the Markdown template or output structure.
|
||
|
||
Implementation: pass a `force=True` flag to `cache.get_new_or_updated()`, which
|
||
returns all conversations regardless of cache state when force is True.
|
||
|
||
Current workaround: `python -m src.main cache --clear` then re-run export.
|
||
|
||
## Joplin `--force` Flag
|
||
|
||
Similarly, add `--force` to the `joplin` command to re-sync all cached
|
||
conversations to Joplin regardless of whether they've been synced before.
|
||
Useful after making formatting changes to the Markdown exporter.
|
||
|
||
Implementation: in `get_joplin_pending()`, return all entries that have a
|
||
`file_path` when `force=True`, ignoring `joplin_synced_at`.
|
||
|
||
## Per-Conversation Cache Reset
|
||
|
||
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
|
||
single conversation without clearing the entire provider cache.
|
||
|
||
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
|
||
delete the entry, then re-run export.
|
||
|
||
## Official API Fallback
|
||
|
||
If the unofficial internal web API approach breaks, migrate to official export
|
||
file parsing as a fallback:
|
||
- ChatGPT: parse `conversations.json` from Settings → Export Data
|
||
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
|
||
|
||
The `BaseProvider` abstract class is intentionally designed so that a
|
||
`FileProvider` subclass can implement the same interface
|
||
(`list_conversations`, `get_conversation`, `normalize_conversation`)
|
||
without any changes to cache, exporters, or CLI code.
|
||
|
||
To add this: implement `src/providers/file_chatgpt.py` and
|
||
`src/providers/file_claude.py`, then add `--input-file` flag to the
|
||
export command to accept a pre-downloaded export ZIP or JSON.
|
||
|
||
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
|
||
this user needs (project data), so it isn't a real fallback here.
|
||
|
||
## Reclassify o1/o3 Reasoning Subparts
|
||
|
||
v0.4.0 leaves dict parts inside `text` content_type messages with shape
|
||
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
|
||
shape was inferred from a code comment, not captured live). Once a real
|
||
reasoning conversation is captured, reclassify these as `thinking` blocks.
|
||
|
||
## Obsidian Vault Output
|
||
|
||
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
|
||
conversations into an Obsidian vault directory. The current Markdown format
|
||
is already largely compatible; the main differences are:
|
||
|
||
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
|
||
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
|
||
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
|
||
supports both
|
||
|
||
Implementation: the existing `MarkdownExporter` output is already valid in
|
||
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
|
||
copy files to the vault directory and maintain a flat or nested folder
|
||
structure matching the user's Obsidian setup. No API needed — just file I/O.
|
||
|
||
## Token Expiry Notifications
|
||
|
||
Proactively warn when a token is close to expiry (within 48h for ChatGPT),
|
||
rather than only surfacing the warning at startup. Options:
|
||
|
||
- Add an `expiry` subcommand that prints token status and exits non-zero if
|
||
any token is expired or expiring soon (useful in scripts/cron)
|
||
- Send a desktop notification via `notify-send` (Linux) or `osascript` (macOS)
|
||
when a token is within 24h of expiry
|
||
|
||
Largely superseded by Brave cookie auto-extraction (#2): once refresh is
|
||
one command (or automatic), expiry warnings matter much less.
|
||
|
||
## Search Command
|
||
|
||
Add a `search` command to full-text search across all exported Markdown files:
|
||
|
||
```bash
|
||
python -m src.main search "kubernetes ingress"
|
||
python -m src.main search "kubernetes ingress" --provider claude --project devops
|
||
```
|
||
|
||
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
|
||
conversation title, date, and a snippet. No index needed — Markdown files are
|
||
small enough to grep directly.
|