release: v0.9.0, and rewrite FUTURE.md around what is actually planned

FUTURE.md had become a 648-line archaeological record: six shipped roadmap
items, two dropped, two implemented, an investigation trail for each, and a
backlog closed as not needed — with the two genuinely planned items buried
at the bottom. It is now 139 lines, planned work first.

- Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon
  behind decisions now recorded in one line each — why Brave cookie
  extraction is not viable, why the ChatGPT token's expiry cannot be read
  client-side, how the drift canary was designed — which would be expensive
  to rediscover.

- Roadmap is now two items: the StartOS service (including the 2026-08-18
  decision that clients upload to StartOS storage and the service owns the
  Joplin connection), and the README split.

- Everything shipped or dropped is off the roadmap. Decisions not to build
  survive as a one-line table so they are not re-proposed.

- Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0.

Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased]
section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the
launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy
notifications, gizmo_id project attribution and the `projects` command, and
the splitlines data-loss fix in both local providers.

Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose —
exported conversations are Markdown and may contain private content — with
each doc re-included by name, so a new doc is silently untracked rather than
rejected. The archive would not have been committed at all. Recorded in the
README-split roadmap item, since `docs/*.md` will hit exactly this.

Also refreshes the package description, which still named only ChatGPT and
Claude after two local providers were added.

Carries the documentation and table-of-contents work from earlier today.
This commit is contained in:
JesseMarkowitz
2026-08-18 13:51:34 -04:00
parent 55b9ce12f6
commit 2d5fcb26f5
7 changed files with 909 additions and 579 deletions
+7
View File
@@ -48,6 +48,13 @@ CLAUDE_SESSION_KEY=
# touched. To never tag specific repos, list their names here (comma-separated).
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
# --- Launcher ---
# Read by the ai-chat-exporter wrapper scripts, not by the Python code. The
# wrapper warns when run from outside the repo, because cache/ and exports/
# resolve against the current directory and the wrong one silently starts a
# separate archive. Set to 1 to silence that warning.
#AI_CHAT_EXPORTER_QUIET_CWD=1
# --- Notifications (ntfy) ---
# Push the result of a run to ntfy so an unattended archive reports back — the
# log file, the systemd journal and Task Scheduler's exit code are all pull-only.
+1
View File
@@ -22,6 +22,7 @@ exports/
!tests/fixtures/*.json
!README.md
!FUTURE.md
!FUTURE-ARCHIVE.md
!CHANGELOG.md
# Cache and logs
+1 -1
View File
@@ -3,7 +3,7 @@
All notable changes to this project will be documented here.
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
## [Unreleased]
## [0.9.0] - 2026-08-18
### Fixed
- **An em dash in a notification title silently dropped the notification.** HTTP header values are latin-1 at best and `requests` raises on anything outside it, so the first real send failed with `'latin-1' codec can't encode character '\u2014'`. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
+660
View File
@@ -0,0 +1,660 @@
# FUTURE.md — archived 2026-08-18 (pre-v0.9.0 cleanup)
This is the full `FUTURE.md` as it stood before the v0.9.0 cleanup, kept
because it carries the investigation trails behind decisions that are now
recorded in one line each: why Brave cookie extraction is not viable, why the
ChatGPT token's expiry cannot be read client-side, how the drift canary was
designed, and the reasoning behind each closed backlog item.
Nothing here is planned work. The live roadmap is in `FUTURE.md`.
---
# Planned Future Work
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
>
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
> active roadmap is empty and the remaining backlog below has been **closed as
> not needed** — the tool does what it's needed to do as a local, manually-run
> backup CLI. Items are kept for reference only; revisit on demand if a real
> need shows up. Nothing here is planned work.
Items completed in each release are moved to the changelog. Items below the
roadmap were designed for but intentionally not implemented. The codebase is
structured to make each of these additions straightforward if ever revived.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
---
# Roadmap (decided 2026-06-12)
Priorities reflect the tool's primary purpose — a trustworthy backup so that
conversation data is not lost if a provider account is ever closed — plus the
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
local, manually-run CLI; the headless/StartOS direction was dropped
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
(manual refresh is sufficient).
**Now (in order):**
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
(full-archive `export --force` re-export completed 2026-06-13)
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
needs admin + SYSTEM impersonation that AV flags as credential theft, and
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
**Soon, but later:**
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
manually-run CLI, so no in-app polling loop is needed
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
sufficient (source convos live in the cloud and can be re-downloaded;
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
also retires the headless token-freshness problem — manual weekly refresh
is fine.
**Active:**
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
via the `/api/auth/session` `error` field (not `expires`). See §9.
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
`canary` command. See §10.
**Deprioritized** (entries kept at the bottom of this file; revisit on
demand): `--force` flags, per-conversation cache reset, official export-ZIP
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
scope — no significant usage to archive.
---
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, with one scoping correction
from live recon: the retrieval dumps are NOT flagged
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
the discriminator (the hidden flag only marks Custom Instructions and small
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
completed 2026-06-13.
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
the user's own attached documents, the same files dumped dozens of times
per conversation. These messages were invisible in the ChatGPT web UI;
they appear in exports because v0.4.0 lifted the role filter to fix silent
data loss. Custom Instructions hidden-context blocks are a minor secondary
case (~2KB, once per conversation) — the originally planned
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
file contains zero hidden-context blocks.
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
`placeholder`, plus a `--hidden-content` CLI override on `export`.
- `placeholder` renders affected messages as one line with type and size:
`> 🔧 Tool output (file_search, 24KB) — omitted
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
roughly halves; worst files shrink ~90%.
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
decision must happen in the provider, not the renderer); (b) messages
flagged `is_visually_hidden_from_conversation`, including
`user_editable_context` / `model_editable_context` (Custom
Instructions) — subsumes the old suppress-hidden-context idea.
- **Scope — keep at full size:** code-execution `tool_result` blocks and
web-search results; those are usually content the user wants.
- Count collapsed messages in the post-export summary so the omission
stays visible (mirror the LossReport presentation, but as intentional
policy, not loss).
Re-export workflow after shipping: `cache --clear` + `export` (same as
the v0.4.0 migration).
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
different mechanism (see below).
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
symptom hit in the field:
```
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
```
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
rights** and behavior that AV/EDR flags as infostealer activity. The one
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
this, so no off-disk reader is a reliable, non-invasive fit.
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
(launch the browser with `--remote-debugging-port`, read cookies via
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
hand. With the headless/StartOS direction dropped (#8), manual DevTools
refresh is the accepted approach — no automated extraction is needed.
## 3. Claude Code Session Provider — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
during implementation: `isSidechain` records are subagent transcripts (skipped),
`isMeta` marks harness-generated user records (skipped), and listing/normalized
`updated_at` must both use file mtime or the cache would re-export every
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
Archive local Claude Code session transcripts. No tokens, no rate limits,
no ToS risk — the data is already on disk but lives in a single JSONL per
session that Claude Code may clean up, and it contains deliverables
(reviews, plans, analyses) that exist nowhere else.
Decisions (2026-06-12):
- **Rendering: prose-only.** Keep user prompts and assistant text
(including full deliverable write-ups); collapse tool activity to
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
86KB — omitted`); **exclude thinking blocks**.
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
in v0.5.0).
Data facts (measured 2026-06-12):
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
Currently 29 sessions, 18.8MB total, largest 3.8MB.
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
activity is what makes these exports readable.
- Record types: `user` / `assistant` (Anthropic-style `message.content`
block arrays) plus harness records: `ai-title` (use for note title and
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
`permission-mode`, `system` (skip). Strip harness noise from user
messages (`<local-command-caveat>`, `<command-name>` blocks).
Implementation shape: new `src/providers/claude_code.py` implementing the
`BaseProvider` interface — `list_conversations` scans project dirs,
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
existing block schema (content is already block-shaped: text / tool_use /
tool_result / thinking). Incremental sync via file mtime/size recorded in
the existing manifest. Project name derives from the munged cwd dirname.
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
(refuses to prune right after `cache --clear`). First live run removed 420
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
manifest↔disk integrity (293/293 after the run).
A backup is only trustworthy if the on-disk tree matches the manifest.
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
v0.4.1) sit alongside current exports and would double-sync into Joplin.
- `prune` command: delete export files not referenced by the manifest.
`--dry-run` (default off, but always print the list before deleting)
shows what would be removed and why (old layout / orphan / unknown).
- `doctor` extension: verify every manifest entry's `file_path` exists on
disk; report missing files (re-export candidates) and unreferenced files
(prune candidates).
## 5. Binary Content Downloads — SHIPPED v0.6.0
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
`download_url`; a second GET yields the bytes (works for user uploads;
older AI-generated images 404 — expired server-side, handled gracefully).
Asset refs come in three shapes — `sediment://file_…`,
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
images-only is the default and audio is opt-in via `all`.
Original notes below.
**Priority note (2026-06-12): "later, but soon" — under the
backup-if-account-closes goal, embedded images are part of the data that
would be lost; placeholders alone don't preserve them.**
v0.4.0 ships placeholders for images and audio assets but does not download
the binary content. The `_safe_fence`-wrapped placeholders include the asset
reference (`sediment://...` or `file-service://...`), MIME type, size, and
duration where available; the actual bytes are not preserved.
Next steps:
- Download attached images alongside the Markdown export, save under a
`media/` sibling directory with a stable filename derived from the asset
reference.
- Replace `image_placeholder` rendering with an inline `![](relative/path)`
reference once the file is on disk.
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
rewrite the rendered Markdown to use `:/resourceId` references, and track
the resource ID in the cache manifest so re-syncs stay idempotent.
- DALL-E images on the assistant side: not observed in this user's data; the
code path exists (`source = "model_generated"`) but is untested.
The block-level schema is already in place — only the file-fetch + rewrite
layer needs to be added. See the `image_placeholder` and `file_placeholder`
block definitions in `src/blocks.py`.
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
the re-run picked it up.
Cap how many conversations are downloaded in a single `export` run so the
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
most importantly on the very first export, which otherwise fetches the
entire conversation history in one session. Because every run is resumable
(the manifest records each conversation immediately), a capped run simply
exports the first N pending conversations and the next run picks up where
it left off. This keeps traffic looking like a human-paced session rather
than a scraper, reducing the risk of rate limiting or account flags.
Two complementary pieces:
1. **Session cap** — `--max-conversations N` flag (and
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
`export` command, slice the pending list after the cache filter:
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
counts (reuse the message format from the 429 early-exit path) and
remind the user to re-run to continue.
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
random jitter) slept between per-conversation detail fetches in
`BaseProvider`, so even a capped run doesn't fire requests
back-to-back. The existing 429 backoff in `_request` stays as the
reactive safety net.
Note: the conversation *listing* (paginated, 100/page) still runs in full
each time so the cache comparison works — the cap applies to the heavy
per-conversation detail fetches, which dominate request volume.
This is a stepping stone to the StartOS service: a capped, politely-paced
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
traffic profile a headless deployment needs.
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
whatever hosts the tool: a user cron line locally, and on the long-term
StartOS target the platform's own scheduling. Either way the tool only needs
to do one capped, politely-paced `export` + `joplin` run and exit — which it
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
that chains `export` then `joplin` for a single cron line is a trivial
add-on, but the polling loop itself is off the roadmap.
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
Not pursuing a headless StartOS service. The local, manually-run CLI is
sufficient: the source conversations live in the providers' clouds and can be
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
so durability is covered without a server in the loop.
Dropping this also retires the one genuinely hard sub-problem it carried —
session-token freshness without a browser. There is no headless context to
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
extraction remains a dead end regardless — see #2.)
### REOPENED as a TODO (2026-08-18) — centralization, not durability
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
decision rested on "the source conversations live in the providers' clouds and
can be re-downloaded". **That is no longer true of half the providers.**
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
that exist *only* on the machine that produced them — Codex prunes its rollout
files, and neither is recoverable from any cloud. The re-download premise now
covers the web providers only.
The new motivation is consolidation rather than durability: work is split across
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
conversations from everywhere, instead of per-machine islands that only meet
inside Joplin.
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what it is uniquely able to do —
read that machine's local transcripts, and hold the browser session for the web
providers. What changes is where the output goes: instead of syncing to Joplin
itself, a local run **uploads its conversations to the StartOS storage area**,
and the StartOS service owns the Joplin connection for the whole corpus.
That inverts today's arrangement, where every machine talks to its own Joplin
desktop, and it removes two problems we already have:
- **The Joplin-availability race disappears from the clients.** A scheduled run
currently has to find Joplin desktop open on that same machine — the
2026-08-18 09:02 timer run exported fine and then skipped the sync because
Joplin did not start until 09:07. Uploading to a server that is always up has
no such window, and `--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note-update logic run once, server-side, against one manifest — rather than
each machine independently deciding what a notebook is called and racing to
update the same note.
What this would need, and what it would *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers can keep running interactively on the
machine that has the browser, pushing their output to the server. Only the
local providers need to run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter currently does not
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing would make that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should probably keep them rather than move them.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
Shipped: `doctor` now adds a "ChatGPT token active" check via
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
decode path was removed, and `_fetch_access_token` now fails fast on a set
`error` instead of returning a stale token. Tests in
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
below for the record.
Goal: if it's cheap to tell how much longer a token will work, show it on
`doctor`. Findings from live recon:
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
never yields an expiry for the real tokens (falls to the "not decodable"
branch).
- **`/api/auth/session` exposes an `expires`** (the provider already calls
this endpoint in `_fetch_access_token`; the response includes `expires`
alongside `accessToken`). Live value observed 2026-06-28:
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
assumption and the lived weekly-refresh cadence, so it is almost certainly
the NextAuth **rolling session window** (re-extended on every call), not
the point at which the pasted token actually 401s. Displaying it verbatim
would give false confidence.
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
was actually dead (conversations API → 401), `/api/auth/session` still
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
So `expires` is a **rolling session window that rolls forward on every call
even for a dead token**; displaying it would actively lie. The same response
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
token `error` is absent/null; when the session token is dead NextAuth can't
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
health check.
- **Design (ready to build):**
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
falsy and `accessToken` present; on `RefreshAccessTokenError` report
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
work) and do NOT surface `expires`.
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
without checking `error`, so it proceeds with a stale token and yields a
confusing downstream 401 instead of a clear "refresh your token" message.
Check `error` there and fail fast.
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
ChatGPT and Claude). It fetches one listing page + one conversation per
provider and asserts only the normalizer's load-bearing fields, emitting
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
trail below for the record.
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
change shape without notice. The worst failure for a backup tool is *silent*:
a response-schema change that makes the exporter skip or mis-parse content
without erroring. `doctor` currently checks token validity, reachability, and
manifest↔disk integrity — but not "does the provider's response still look
like what the parser expects."
To investigate: a lightweight schema/shape assertion on a known-good sample
of each provider's listing + conversation-detail responses (presence and type
of the fields the normalizers rely on), surfaced as a `doctor` check or a
dedicated canary.
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
- **Dependency surface (assert ONLY these — see below for why):**
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
`project.name`, `chat_messages[]`.
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
(list of typed blocks), `created_at`.
- **Key finding — full-shape diffing is the wrong design.** Claude's
`settings` object is full of volatile internal codenames that churn
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
canary would fire on every UI experiment. The canary MUST target the
normalizer's load-bearing fields only, not the whole response. (Aligns with
the drop-noise-don't-retain-it principle.)
- **Real Claude messages are flat `text`/`sender`** — in this archive every
message had a string `text` and NO `content` block list (0 rich blocks
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
thinking, image, …) is an **unexercised theoretical path**; drift there
can't be "caught" by a canary because it never runs on real data — it's a
safety net for if Claude ever switches to block content. The canary should
assert the flat shape and *warn if `content` ever appears as a list* (that
itself is the drift event that would activate the dormant code).
- **Possible silent-loss spot (separate from drift):** Claude messages carry
`attachments` and `files` arrays (empty in this sample) that the normalizer
ignores entirely. If a user ever attaches files in Claude, they'd be
dropped without a LossReport entry. Worth a follow-up check.
**Live recon — ChatGPT captured 2026-06-28:**
- **Dependency surface (assert ONLY these):**
- listing item: `id`, `title`, `update_time`/`create_time`.
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
`update_time`, `mapping` (non-empty).
- mapping node: `message`, `children` (the tree walk depends on both);
message: `author.role`, `author.name`, `content.content_type`,
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
- **content_type vocabulary observed (all currently handled):** `text`,
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
`execution_output`, `reasoning_recap`, `user_editable_context`,
`tether_browsing_display`. A *new* content_type already degrades gracefully
(visible `unknown` block + WARNING + LossReport tally) — so content_type
drift is **already non-silent**. The canary just needs to confirm the known
set still parses to non-empty blocks.
- **The genuinely silent drift risk — `author.name` collapse keys.**
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
confirms `file_search` is live (and `web.run`/`python` are correctly left
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
stops** and the archive re-bloats with no error or LossReport entry. This is
the top canary target: assert that retrieval-dump tool authors are still
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
- **Second silent risk — empty `content.parts`.** A `text` message whose
`parts` field is renamed/emptied yields zero blocks and is skipped with only
a debug log = silent loss. Canary should assert a sampled `text` message
produces a non-empty block.
**Recon complete for both providers. Canary design (ready to build):** a
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
listing page + one conversation and asserts the dependency-surface fields
above by presence+type — NOT full shape (Claude `settings` codenames prove
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
surface as a warning, never a hard error (a backup tool must still run).
---
# Deprioritized — CLOSED as not needed (2026-06-28)
These were considered and intentionally **not** built. Closed, not planned —
the tool is feature-complete for its purpose. Kept for reference in case a
real need ever revives one: Joplin `--force`, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
output, token-expiry notifications (also moot — see §9), and a search
command. (Also closed, outside this list: handling Claude `attachments`/
`files`, which the canary will flag if they ever appear in real data.)
## Export `--force` Flag — SHIPPED v0.6.0
Implemented 2026-06-12: `export --force` passes `force=True` to
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
preserves Joplin links across re-exports, so a forced re-render + `joplin`
updates existing notes instead of duplicating them.
## Joplin `--force` Flag
Similarly, add `--force` to the `joplin` command to re-sync all cached
conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
Implementation: in `get_joplin_pending()`, return all entries that have a
`file_path` when `force=True`, ignoring `joplin_synced_at`.
## Per-Conversation Cache Reset
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
single conversation without clearing the entire provider cache.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
## Official API Fallback
If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface
(`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON.
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
this user needs (project data), so it isn't a real fallback here.
## Reclassify o1/o3 Reasoning Subparts
v0.4.0 leaves dict parts inside `text` content_type messages with shape
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
shape was inferred from a code comment, not captured live). Once a real
reasoning conversation is captured, reclassify these as `thinking` blocks.
## Obsidian Vault Output
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
## Token Expiry Notifications
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
"proactively notify before expiry" idea is blocked by the same finding: a
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
window rather than the real refresh cadence). Any heads-up — a `doctor`
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
resolving the §9 open question of what signal is actually trustworthy.
## Search Command
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
## Split the README into Separate Documents
**TODO (2026-08-18).** The README is 830 lines / 5,562 words / 37 KB — about a
25-minute read, with 61 headings. An H2-only table of contents was added the
same day and helps navigation, but it treats the symptom: the file is doing at
least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
Not done yet, and not urgent, because it has a real cost the TOC does not: any
existing link into a README section (a bookmark, a note, another repo, a commit
message) breaks when that section moves to another file. Worth doing when the
README next needs substantial editing anyway, rather than as a change of its own.
Two things to decide when it happens:
- Whether `docs/` renders acceptably on the Gitea instance that hosts this repo
(relative links between Markdown files do work there, but worth confirming
before splitting rather than after).
- Whether the anchors in the split files stay stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors are derived from heading text, so they break silently on a reword —
the same fragility the TOC already carries.
+101 -573
View File
@@ -1,616 +1,144 @@
# Planned Future Work
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
>
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
> active roadmap is empty and the remaining backlog below has been **closed as
> not needed** — the tool does what it's needed to do as a local, manually-run
> backup CLI. Items are kept for reference only; revisit on demand if a real
> need shows up. Nothing here is planned work.
> **Status 2026-08-18 (v0.9.0).** The tool archives four providers — two web
> (`chatgpt`, `claude`) and two local agent-transcript (`claude-code`, `codex`)
> — on a schedule, on Linux and Windows, reporting results by push notification.
> Two items below are genuinely planned. Everything else has shipped or been
> decided against.
Items completed in each release are moved to the changelog. Items below the
roadmap were designed for but intentionally not implemented. The codebase is
structured to make each of these additions straightforward if ever revived.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
Completed work moves to the changelog; this file holds only what is *not* built
yet. Decisions not to build something are recorded at the bottom in one line
each, so they are not re-proposed — the full investigation trails behind them
are in `FUTURE-ARCHIVE.md`.
---
# Roadmap (decided 2026-06-12)
# Roadmap
Priorities reflect the tool's primary purpose — a trustworthy backup so that
conversation data is not lost if a provider account is ever closed — plus the
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
local, manually-run CLI; the headless/StartOS direction was dropped
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
(manual refresh is sufficient).
## 1. StartOS Service — one corpus, not per-machine islands
**Now (in order):**
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
(full-archive `export --force` re-export completed 2026-06-13)
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
needs admin + SYSTEM impersonation that AV flags as credential theft, and
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
Each machine currently archives to its own `exports/` and its own Joplin. Work
is split across boxes — coding sessions (`claude-code`, `codex`) on the Linux
machine, web chats (`chatgpt`, `claude`) on the Windows one — so there is no
single place where all conversations exist together.
**Soon, but later:**
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
manually-run CLI, so no in-app polling loop is needed
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
sufficient (source convos live in the cloud and can be re-downloaded;
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
also retires the headless token-freshness problem — manual weekly refresh
is fine.
**Active:**
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
via the `/api/auth/session` `error` field (not `expires`). See §9.
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
`canary` command. See §10.
**Deprioritized** (entries kept at the bottom of this file; revisit on
demand): `--force` flags, per-conversation cache reset, official export-ZIP
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
scope — no significant usage to archive.
---
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, with one scoping correction
from live recon: the retrieval dumps are NOT flagged
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
the discriminator (the hidden flag only marks Custom Instructions and small
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
completed 2026-06-13.
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
the user's own attached documents, the same files dumped dozens of times
per conversation. These messages were invisible in the ChatGPT web UI;
they appear in exports because v0.4.0 lifted the role filter to fix silent
data loss. Custom Instructions hidden-context blocks are a minor secondary
case (~2KB, once per conversation) — the originally planned
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
file contains zero hidden-context blocks.
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
`placeholder`, plus a `--hidden-content` CLI override on `export`.
- `placeholder` renders affected messages as one line with type and size:
`> 🔧 Tool output (file_search, 24KB) — omitted
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
roughly halves; worst files shrink ~90%.
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
decision must happen in the provider, not the renderer); (b) messages
flagged `is_visually_hidden_from_conversation`, including
`user_editable_context` / `model_editable_context` (Custom
Instructions) — subsumes the old suppress-hidden-context idea.
- **Scope — keep at full size:** code-execution `tool_result` blocks and
web-search results; those are usually content the user wants.
- Count collapsed messages in the post-export summary so the omission
stays visible (mirror the LossReport presentation, but as intentional
policy, not loss).
Re-export workflow after shipping: `cache --clear` + `export` (same as
the v0.4.0 migration).
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
different mechanism (see below).
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
symptom hit in the field:
```
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
```
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
rights** and behavior that AV/EDR flags as infostealer activity. The one
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
this, so no off-disk reader is a reliable, non-invasive fit.
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
(launch the browser with `--remote-debugging-port`, read cookies via
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
hand. With the headless/StartOS direction dropped (#8), manual DevTools
refresh is the accepted approach — no automated extraction is needed.
## 3. Claude Code Session Provider — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
during implementation: `isSidechain` records are subagent transcripts (skipped),
`isMeta` marks harness-generated user records (skipped), and listing/normalized
`updated_at` must both use file mtime or the cache would re-export every
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
Archive local Claude Code session transcripts. No tokens, no rate limits,
no ToS risk — the data is already on disk but lives in a single JSONL per
session that Claude Code may clean up, and it contains deliverables
(reviews, plans, analyses) that exist nowhere else.
Decisions (2026-06-12):
- **Rendering: prose-only.** Keep user prompts and assistant text
(including full deliverable write-ups); collapse tool activity to
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
86KB — omitted`); **exclude thinking blocks**.
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
in v0.5.0).
Data facts (measured 2026-06-12):
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
Currently 29 sessions, 18.8MB total, largest 3.8MB.
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
activity is what makes these exports readable.
- Record types: `user` / `assistant` (Anthropic-style `message.content`
block arrays) plus harness records: `ai-title` (use for note title and
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
`permission-mode`, `system` (skip). Strip harness noise from user
messages (`<local-command-caveat>`, `<command-name>` blocks).
Implementation shape: new `src/providers/claude_code.py` implementing the
`BaseProvider` interface — `list_conversations` scans project dirs,
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
existing block schema (content is already block-shaped: text / tool_use /
tool_result / thinking). Incremental sync via file mtime/size recorded in
the existing manifest. Project name derives from the munged cwd dirname.
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
(refuses to prune right after `cache --clear`). First live run removed 420
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
manifest↔disk integrity (293/293 after the run).
A backup is only trustworthy if the on-disk tree matches the manifest.
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
v0.4.1) sit alongside current exports and would double-sync into Joplin.
- `prune` command: delete export files not referenced by the manifest.
`--dry-run` (default off, but always print the list before deleting)
shows what would be removed and why (old layout / orphan / unknown).
- `doctor` extension: verify every manifest entry's `file_path` exists on
disk; report missing files (re-export candidates) and unreferenced files
(prune candidates).
## 5. Binary Content Downloads — SHIPPED v0.6.0
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
`download_url`; a second GET yields the bytes (works for user uploads;
older AI-generated images 404 — expired server-side, handled gracefully).
Asset refs come in three shapes — `sediment://file_…`,
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
images-only is the default and audio is opt-in via `all`.
Original notes below.
**Priority note (2026-06-12): "later, but soon" — under the
backup-if-account-closes goal, embedded images are part of the data that
would be lost; placeholders alone don't preserve them.**
v0.4.0 ships placeholders for images and audio assets but does not download
the binary content. The `_safe_fence`-wrapped placeholders include the asset
reference (`sediment://...` or `file-service://...`), MIME type, size, and
duration where available; the actual bytes are not preserved.
Next steps:
- Download attached images alongside the Markdown export, save under a
`media/` sibling directory with a stable filename derived from the asset
reference.
- Replace `image_placeholder` rendering with an inline `![](relative/path)`
reference once the file is on disk.
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
rewrite the rendered Markdown to use `:/resourceId` references, and track
the resource ID in the cache manifest so re-syncs stay idempotent.
- DALL-E images on the assistant side: not observed in this user's data; the
code path exists (`source = "model_generated"`) but is untested.
The block-level schema is already in place — only the file-fetch + rewrite
layer needs to be added. See the `image_placeholder` and `file_placeholder`
block definitions in `src/blocks.py`.
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
the re-run picked it up.
Cap how many conversations are downloaded in a single `export` run so the
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
most importantly on the very first export, which otherwise fetches the
entire conversation history in one session. Because every run is resumable
(the manifest records each conversation immediately), a capped run simply
exports the first N pending conversations and the next run picks up where
it left off. This keeps traffic looking like a human-paced session rather
than a scraper, reducing the risk of rate limiting or account flags.
Two complementary pieces:
1. **Session cap** — `--max-conversations N` flag (and
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
`export` command, slice the pending list after the cache filter:
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
counts (reuse the message format from the 429 early-exit path) and
remind the user to re-run to continue.
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
random jitter) slept between per-conversation detail fetches in
`BaseProvider`, so even a capped run doesn't fire requests
back-to-back. The existing 429 backoff in `_request` stays as the
reactive safety net.
Note: the conversation *listing* (paginated, 100/page) still runs in full
each time so the cache comparison works — the cap applies to the heavy
per-conversation detail fetches, which dominate request volume.
This is a stepping stone to the StartOS service: a capped, politely-paced
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
traffic profile a headless deployment needs.
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
whatever hosts the tool: a user cron line locally, and on the long-term
StartOS target the platform's own scheduling. Either way the tool only needs
to do one capped, politely-paced `export` + `joplin` run and exit — which it
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
that chains `export` then `joplin` for a single cron line is a trivial
add-on, but the polling loop itself is off the roadmap.
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
Not pursuing a headless StartOS service. The local, manually-run CLI is
sufficient: the source conversations live in the providers' clouds and can be
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
so durability is covered without a server in the loop.
Dropping this also retires the one genuinely hard sub-problem it carried —
session-token freshness without a browser. There is no headless context to
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
extraction remains a dead end regardless — see #2.)
### REOPENED as a TODO (2026-08-18) — centralization, not durability
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
decision rested on "the source conversations live in the providers' clouds and
can be re-downloaded". **That is no longer true of half the providers.**
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
that exist *only* on the machine that produced them — Codex prunes its rollout
files, and neither is recoverable from any cloud. The re-download premise now
covers the web providers only.
The new motivation is consolidation rather than durability: work is split across
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
conversations from everywhere, instead of per-machine islands that only meet
inside Joplin.
This was dropped on 2026-06-28 and **reopened 2026-08-18**, because the
reasoning behind the drop has gone stale. It rested on "the source conversations
live in the providers' clouds and can be re-downloaded". That is no longer true
of half the providers: `claude-code` and `codex` transcripts exist *only* on the
machine that produced them, Codex prunes its rollout files, and neither is
recoverable from any cloud. The motivation is also different from the one
weighed then — consolidation, not durability.
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what it is uniquely able to do —
read that machine's local transcripts, and hold the browser session for the web
providers. What changes is where the output goes: instead of syncing to Joplin
itself, a local run **uploads its conversations to the StartOS storage area**,
and the StartOS service owns the Joplin connection for the whole corpus.
running the exporter locally and keeps doing what only it can do: read that
machine's local transcripts, and hold the browser session for the web providers.
What changes is where the output goes. Instead of syncing to Joplin itself, a
local run **uploads its conversations to the StartOS storage area**, and the
StartOS service owns the Joplin connection for the whole corpus.
That inverts today's arrangement, where every machine talks to its own Joplin
desktop, and it removes two problems we already have:
That inverts today's arrangement and removes two problems we already have:
- **The Joplin-availability race disappears from the clients.** A scheduled run
currently has to find Joplin desktop open on that same machine — the
2026-08-18 09:02 timer run exported fine and then skipped the sync because
Joplin did not start until 09:07. Uploading to a server that is always up has
no such window, and `--joplin-optional` stops being load-bearing.
- **The Joplin-availability race leaves the clients.** A scheduled run currently
has to find Joplin desktop open on that same machine — the 09:02 timer run on
2026-08-18 exported fine and then skipped the sync because Joplin did not
start until 09:07. A server that is always up has no such window, and
`--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note-update logic run once, server-side, against one manifest — rather than
each machine independently deciding what a notebook is called and racing to
update the same note.
note updates run once, server-side, against one manifest — rather than each
machine independently deciding what a notebook is called and racing to update
the same note.
What this would need, and what it would *not*:
What this needs, and what it does *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers can keep running interactively on the
machine that has the browser, pushing their output to the server. Only the
local providers need to run server-side, and they need no tokens at all.
retired stays retired: the web providers keep running interactively on the
machine that has the browser, and push their output to the server. Only the
local providers would run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter currently does not
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing would make that label load-bearing.
`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter does not track today:
`claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing makes that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should probably keep them rather than move them.
rollouts, so the client should keep them rather than hand them off.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
## 2. Split the README into separate documents
Shipped: `doctor` now adds a "ChatGPT token active" check via
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
decode path was removed, and `_fetch_access_token` now fails fast on a set
`error` instead of returning a stale token. Tests in
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
below for the record.
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61
headings. An H2-only table of contents was added 2026-08-18 and helps, but it
treats the symptom: the file is doing at least four unrelated jobs at once.
Goal: if it's cheap to tell how much longer a token will work, show it on
`doctor`. Findings from live recon:
Rough shape of a split:
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
never yields an expiry for the real tokens (falls to the "not decodable"
branch).
- **`/api/auth/session` exposes an `expires`** (the provider already calls
this endpoint in `_fetch_access_token`; the response includes `expires`
alongside `accessToken`). Live value observed 2026-06-28:
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
assumption and the lived weekly-refresh cadence, so it is almost certainly
the NextAuth **rolling session window** (re-extended on every call), not
the point at which the pasted token actually 401s. Displaying it verbatim
would give false confidence.
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
was actually dead (conversations API → 401), `/api/auth/session` still
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
So `expires` is a **rolling session window that rolls forward on every call
even for a dead token**; displaying it would actively lie. The same response
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
token `error` is absent/null; when the session token is dead NextAuth can't
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
health check.
- **Design (ready to build):**
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
falsy and `accessToken` present; on `RefreshAccessTokenError` report
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
work) and do NOT surface `expires`.
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
without checking `error`, so it proceeds with a stale token and yields a
confusing downstream 401 instead of a clear "refresh your token" message.
Check `error` there and fail fast.
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
Not urgent, because it has a real cost the TOC did not: any existing link into a
README section — a bookmark, a note, another repo, a commit message — breaks
when that section moves to another file. Worth folding into the next substantial
README edit rather than doing as a change of its own.
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
ChatGPT and Claude). It fetches one listing page + one conversation per
provider and asserts only the normalizer's load-bearing fields, emitting
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
trail below for the record.
Two things to settle when it happens:
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
change shape without notice. The worst failure for a backup tool is *silent*:
a response-schema change that makes the exporter skip or mis-parse content
without erroring. `doctor` currently checks token validity, reachability, and
manifest↔disk integrity — but not "does the provider's response still look
like what the parser expects."
To investigate: a lightweight schema/shape assertion on a known-good sample
of each provider's listing + conversation-detail responses (presence and type
of the fields the normalizers rely on), surfaced as a `doctor` check or a
dedicated canary.
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
- **Dependency surface (assert ONLY these — see below for why):**
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
`project.name`, `chat_messages[]`.
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
(list of typed blocks), `created_at`.
- **Key finding — full-shape diffing is the wrong design.** Claude's
`settings` object is full of volatile internal codenames that churn
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
canary would fire on every UI experiment. The canary MUST target the
normalizer's load-bearing fields only, not the whole response. (Aligns with
the drop-noise-don't-retain-it principle.)
- **Real Claude messages are flat `text`/`sender`** — in this archive every
message had a string `text` and NO `content` block list (0 rich blocks
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
thinking, image, …) is an **unexercised theoretical path**; drift there
can't be "caught" by a canary because it never runs on real data — it's a
safety net for if Claude ever switches to block content. The canary should
assert the flat shape and *warn if `content` ever appears as a list* (that
itself is the drift event that would activate the dormant code).
- **Possible silent-loss spot (separate from drift):** Claude messages carry
`attachments` and `files` arrays (empty in this sample) that the normalizer
ignores entirely. If a user ever attaches files in Claude, they'd be
dropped without a LossReport entry. Worth a follow-up check.
**Live recon — ChatGPT captured 2026-06-28:**
- **Dependency surface (assert ONLY these):**
- listing item: `id`, `title`, `update_time`/`create_time`.
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
`update_time`, `mapping` (non-empty).
- mapping node: `message`, `children` (the tree walk depends on both);
message: `author.role`, `author.name`, `content.content_type`,
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
- **content_type vocabulary observed (all currently handled):** `text`,
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
`execution_output`, `reasoning_recap`, `user_editable_context`,
`tether_browsing_display`. A *new* content_type already degrades gracefully
(visible `unknown` block + WARNING + LossReport tally) — so content_type
drift is **already non-silent**. The canary just needs to confirm the known
set still parses to non-empty blocks.
- **The genuinely silent drift risk — `author.name` collapse keys.**
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
confirms `file_search` is live (and `web.run`/`python` are correctly left
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
stops** and the archive re-bloats with no error or LossReport entry. This is
the top canary target: assert that retrieval-dump tool authors are still
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
- **Second silent risk — empty `content.parts`.** A `text` message whose
`parts` field is renamed/emptied yields zero blocks and is skipped with only
a debug log = silent loss. Canary should assert a sampled `text` message
produces a non-empty block.
**Recon complete for both providers. Canary design (ready to build):** a
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
listing page + one conversation and asserts the dependency-surface fields
above by presence+type — NOT full shape (Claude `settings` codenames prove
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
surface as a warning, never a hard error (a backup tool must still run).
- **`.gitignore` ignores `*.md` on purpose** — exported conversations are
Markdown and may contain private content — and re-includes each doc by name
(`!README.md`, `!FUTURE.md`, …). New files under `docs/` will be silently
ignored, with no error, until `!docs/*.md` is added. This already bit the
creation of `FUTURE-ARCHIVE.md` on 2026-08-18.
- Whether `docs/` renders acceptably on the Gitea instance hosting this repo.
Relative links between Markdown files do work there, but confirm before
splitting rather than after.
- Whether anchors in the split files are stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors derive from heading text and break silently on a reword — the same
fragility the TOC already carries.
---
# Deprioritized — CLOSED as not needed (2026-06-28)
# Decided against
These were considered and intentionally **not** built. Closed, not planned —
the tool is feature-complete for its purpose. Kept for reference in case a
real need ever revives one: Joplin `--force`, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
output, token-expiry notifications (also moot — see §9), and a search
command. (Also closed, outside this list: handling Claude `attachments`/
`files`, which the canary will flag if they ever appear in real data.)
Recorded so they are not re-proposed. Full reasoning and the recon behind each
is in `FUTURE-ARCHIVE.md`.
## Export `--force` Flag — SHIPPED v0.6.0
| Item | Verdict |
|------|---------|
| Brave/Chromium cookie auto-extraction | **Not viable** (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
| In-app watch/polling loop | **Dropped** (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as `sync` plus a systemd timer and a Windows scheduled task. |
| Proactive token-expiry notification | **Blocked, and moot** (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and `/api/auth/session`'s `expires` is a rolling window that advances even for a dead token. `doctor`'s `error`-based health check covers the real need. |
| Official export-ZIP fallback | **Closed** (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. `BaseProvider` still admits a `FileProvider` if that changes. |
| Joplin `--force` flag | Closed as not needed (2026-06-28). |
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
| Search command | Closed as not needed (2026-06-28). `grep`/`ripgrep` over `EXPORT_DIR` covers it. |
| Claude `attachments` / `files` handling | Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
Implemented 2026-06-12: `export --force` passes `force=True` to
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
preserves Joplin links across re-exports, so a forced re-render + `joplin`
updates existing notes instead of duplicating them.
---
## Joplin `--force` Flag
# Completed
Similarly, add `--force` to the `joplin` command to re-sync all cached
conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
Detail for each release is in `CHANGELOG.md`.
Implementation: in `get_joplin_pending()`, return all entries that have a
`file_path` when `force=True`, ignoring `joplin_synced_at`.
## Per-Conversation Cache Reset
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
single conversation without clearing the entire provider cache.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
## Official API Fallback
If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface
(`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON.
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
this user needs (project data), so it isn't a real fallback here.
## Reclassify o1/o3 Reasoning Subparts
v0.4.0 leaves dict parts inside `text` content_type messages with shape
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
shape was inferred from a code comment, not captured live). Once a real
reasoning conversation is captured, reclassify these as `thinking` blocks.
## Obsidian Vault Output
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
## Token Expiry Notifications
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
"proactively notify before expiry" idea is blocked by the same finding: a
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
window rather than the real refresh cadence). Any heads-up — a `doctor`
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
resolving the §9 open question of what signal is actually trustworthy.
## Search Command
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
- **v0.1.0** — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- **v0.2.0** — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- **v0.4.0** — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via `LossReport` and visible `unknown` blocks
- **v0.5.0** — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- **v0.6.0** — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT`); Claude Code session provider; `prune` + manifest integrity; binary content downloads; session limiter and request pacing; `export --force`
- **v0.7.0** — `canary` drift detection; real ChatGPT token-health check on `doctor`; removal of the non-viable browser-cookie path
- **v0.8.0** — Claude Code coverage reopened: subagent capture as folded `<details>`, repo `[tags]` in titles, its own `AI-ClaudeCode` notebook with self-healing note moves, multi-root scanning (`CLAUDE_CODE_DIR`, `CLAUDE_CONFIG_DIR`)
- **v0.9.0** — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms; `sync` with a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution via `gizmo_id` and the `projects` command; a silent data-loss fix in both local providers (`splitlines` breaking on U+0085/U+2028/U+2029)
+137 -3
View File
@@ -6,7 +6,28 @@ Supports incremental sync — only new or updated conversations are exported on
---
## ⚠️ Terms of Service Warning
## Contents
- [Terms of Service Warning](#terms-of-service-warning)
- [Installation](#installation)
- [First Run: Run Doctor](#first-run-run-doctor)
- [Getting Your Session Tokens](#getting-your-session-tokens)
- [The `auth` Command](#the-auth-command)
- [`.env` Setup](#env-setup)
- [ChatGPT Projects](#chatgpt-projects)
- [Claude Code Sessions](#claude-code-sessions)
- [Codex Sessions](#codex-sessions)
- [Scheduling a Daily Run](#scheduling-a-daily-run)
- [Output Structure](#output-structure)
- [CLI Reference](#cli-reference)
- [How the Cache Works](#how-the-cache-works)
- [Troubleshooting](#troubleshooting)
- [Future Work](#future-work)
- [Security Notes](#security-notes)
---
## Terms of Service Warning
**Read this before using this tool.**
@@ -223,12 +244,26 @@ cp .env.example .env
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
### Notifications
| Variable | Default | Description |
|----------|---------|-------------|
| `NTFY_TOPIC` | — | [ntfy](https://ntfy.sh) topic to push run results to. Unset disables notifications entirely. |
| `NTFY_SERVER` | `https://ntfy.sh` | Point at your own host if self-hosting. |
| `NTFY_TOKEN` | — | Bearer token, for access-controlled topics. |
| `NTFY_NOTIFY` | `always` | `always` notifies on every run, `failure` only when something failed, `off` never. |
A topic on public ntfy.sh is readable by anyone who knows its name, so
notifications carry per-provider counts and a machine name only — never
conversation titles. See [Getting notified](#getting-notified).
### Cache & logging
| Variable | Default | Description |
|----------|---------|-------------|
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
| `AI_CHAT_EXPORTER_QUIET_CWD` | — | Set to `1` to silence the launcher's warning when run from outside the repo. Read by the `ai-chat-exporter` wrapper scripts, not by Python; the scheduler installers set it, since they always set the correct working directory. |
---
@@ -589,7 +624,59 @@ Reads the local export cache and pushes each exported Markdown file to Joplin as
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
4. Joplin desktop must be open when you run this command
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--project NAME`, `--dry-run`
### `sync` — Export and sync in one run
```bash
# The whole archive run: export, then push to Joplin
ai-chat-exporter sync
# One provider
ai-chat-exporter sync --provider codex
# Export only; don't touch Joplin
ai-chat-exporter sync --skip-joplin
# Joplin being closed is a warning, not a failure (used by the schedulers)
ai-chat-exporter sync --joplin-optional
```
Equivalent to `export` followed by `joplin` with the same `--provider`. Intended
for scheduled runs — see [Scheduling a Daily Run](#scheduling-a-daily-run).
Unlike the individual commands, `sync` sets a **meaningful exit code**: non-zero
if any conversation failed to export or any note failed to sync. A provider whose
listing call fails outright (an expired web session token being the usual cause)
counts its whole batch as failed. A provider that is simply unconfigured, or that
had nothing new, is ordinary success. `export` on its own always exits 0, which
is fine when you're reading the summary table and useless to a scheduler.
`--joplin-optional` downgrades an unreachable Joplin to a warning: the export has
already captured the local transcripts, and the notes are rebuilt from the cache
by the next run that finds Joplin open.
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--since YYYY-MM-DD`, `--hidden-content [full|placeholder|omit]`, `--max-conversations N`, `--skip-joplin`, `--joplin-optional`, `--notify/--no-notify`, `--dry-run`
Note this is a deliberate subset of `export`'s options — `--format`, `--output`,
`--project`, `--download-media` and `--force` are not passed through. Use
`export` directly for those. (`--download-media` still applies from `.env`; the
flag is only a per-run override.)
### `notify` — Push-notification settings and test
```bash
# Show the current settings
ai-chat-exporter notify
# Send a test push to confirm the topic works
ai-chat-exporter notify --test
```
Shows the resolved ntfy configuration and which machine name will appear in the
title. See [Getting notified](#getting-notified) for what a scheduled run sends.
Options: `--test`
### `prune` — Delete stale export files
@@ -607,6 +694,50 @@ Joplin. Refuses to run when the manifest is empty (e.g. right after
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
command separately verifies that every manifest entry's file exists on disk.
### `projects` — Discover ChatGPT project IDs
```bash
# List the projects your conversations belong to
ai-chat-exporter projects
# Also inspect conversations whose listing entry doesn't name a project
ai-chat-exporter projects --deep
# Write the discovered IDs straight into .env
ai-chat-exporter projects --write
```
`CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is
invisible to the listing pass — conversations that live *only* inside that
project are never fetched at all. This reports every project your conversations
belong to, marks the ones absent from `.env`, and prints a paste-ready line.
`--deep` fetches each conversation's detail when the listing doesn't name its
project: complete, but one request per conversation, so it's slow.
Options: `--deep`, `--write`
### `canary` — Check for provider API drift
```bash
ai-chat-exporter canary
ai-chat-exporter canary --provider chatgpt
```
The web providers are undocumented internal APIs that can change shape without
notice, and the failure mode is silent — a renamed field means content is
quietly dropped rather than an error being raised. The canary fetches one
listing page and one conversation per provider and asserts only the fields the
normalizer actually depends on.
Findings are `ERROR` (a load-bearing field is missing or mistyped — the parser
will break or silently lose data) or `WARN` (something unfamiliar appeared;
worth investigating, not necessarily broken). **Exits non-zero on any ERROR**,
so it can be scheduled or run in CI. Local providers have no remote schema and
are not probed.
Options: `--provider [chatgpt|claude|all]`
### `cache` — Manage the sync manifest
```bash
@@ -706,7 +837,10 @@ No new or updated conversations since your last run. To verify: `ai-chat-exporte
See `FUTURE.md` for the full roadmap. Current priorities:
- **Watch/scheduled mode** on the way to a headless StartOS service
- **A StartOS service** that centralises every machine's conversations into one
corpus and owns the Joplin connection, so each machine only has to upload
(`FUTURE.md` §8)
- **Splitting this README** into a short overview plus separate documents
---
+2 -2
View File
@@ -4,8 +4,8 @@ build-backend = "setuptools.build_meta"
[project]
name = "ai-chat-exporter"
version = "0.8.0"
description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin"
version = "0.9.0"
description = "Archive ChatGPT, Claude, Claude Code and Codex conversation history to Markdown for personal backup in Joplin"
requires-python = ">=3.11"
dependencies = [
"requests==2.31.0",