release: v0.9.0, and rewrite FUTURE.md around what is actually planned
FUTURE.md had become a 648-line archaeological record: six shipped roadmap items, two dropped, two implemented, an investigation trail for each, and a backlog closed as not needed — with the two genuinely planned items buried at the bottom. It is now 139 lines, planned work first. - Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon behind decisions now recorded in one line each — why Brave cookie extraction is not viable, why the ChatGPT token's expiry cannot be read client-side, how the drift canary was designed — which would be expensive to rediscover. - Roadmap is now two items: the StartOS service (including the 2026-08-18 decision that clients upload to StartOS storage and the service owns the Joplin connection), and the README split. - Everything shipped or dropped is off the roadmap. Decisions not to build survive as a one-line table so they are not re-proposed. - Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0. Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased] section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy notifications, gizmo_id project attribution and the `projects` command, and the splitlines data-loss fix in both local providers. Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose — exported conversations are Markdown and may contain private content — with each doc re-included by name, so a new doc is silently untracked rather than rejected. The archive would not have been committed at all. Recorded in the README-split roadmap item, since `docs/*.md` will hit exactly this. Also refreshes the package description, which still named only ChatGPT and Claude after two local providers were added. Carries the documentation and table-of-contents work from earlier today.
This commit is contained in:
@@ -48,6 +48,13 @@ CLAUDE_SESSION_KEY=
|
||||
# touched. To never tag specific repos, list their names here (comma-separated).
|
||||
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
|
||||
|
||||
# --- Launcher ---
|
||||
# Read by the ai-chat-exporter wrapper scripts, not by the Python code. The
|
||||
# wrapper warns when run from outside the repo, because cache/ and exports/
|
||||
# resolve against the current directory and the wrong one silently starts a
|
||||
# separate archive. Set to 1 to silence that warning.
|
||||
#AI_CHAT_EXPORTER_QUIET_CWD=1
|
||||
|
||||
# --- Notifications (ntfy) ---
|
||||
# Push the result of a run to ntfy so an unattended archive reports back — the
|
||||
# log file, the systemd journal and Task Scheduler's exit code are all pull-only.
|
||||
|
||||
@@ -22,6 +22,7 @@ exports/
|
||||
!tests/fixtures/*.json
|
||||
!README.md
|
||||
!FUTURE.md
|
||||
!FUTURE-ARCHIVE.md
|
||||
!CHANGELOG.md
|
||||
|
||||
# Cache and logs
|
||||
|
||||
+1
-1
@@ -3,7 +3,7 @@
|
||||
All notable changes to this project will be documented here.
|
||||
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
||||
|
||||
## [Unreleased]
|
||||
## [0.9.0] - 2026-08-18
|
||||
|
||||
### Fixed
|
||||
- **An em dash in a notification title silently dropped the notification.** HTTP header values are latin-1 at best and `requests` raises on anything outside it, so the first real send failed with `'latin-1' codec can't encode character '\u2014'`. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
|
||||
|
||||
@@ -0,0 +1,660 @@
|
||||
# FUTURE.md — archived 2026-08-18 (pre-v0.9.0 cleanup)
|
||||
|
||||
This is the full `FUTURE.md` as it stood before the v0.9.0 cleanup, kept
|
||||
because it carries the investigation trails behind decisions that are now
|
||||
recorded in one line each: why Brave cookie extraction is not viable, why the
|
||||
ChatGPT token's expiry cannot be read client-side, how the drift canary was
|
||||
designed, and the reasoning behind each closed backlog item.
|
||||
|
||||
Nothing here is planned work. The live roadmap is in `FUTURE.md`.
|
||||
|
||||
---
|
||||
|
||||
# Planned Future Work
|
||||
|
||||
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
|
||||
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
|
||||
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
|
||||
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
|
||||
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
|
||||
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
|
||||
>
|
||||
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
|
||||
> active roadmap is empty and the remaining backlog below has been **closed as
|
||||
> not needed** — the tool does what it's needed to do as a local, manually-run
|
||||
> backup CLI. Items are kept for reference only; revisit on demand if a real
|
||||
> need shows up. Nothing here is planned work.
|
||||
|
||||
Items completed in each release are moved to the changelog. Items below the
|
||||
roadmap were designed for but intentionally not implemented. The codebase is
|
||||
structured to make each of these additions straightforward if ever revived.
|
||||
|
||||
**Completed:**
|
||||
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
||||
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
||||
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
|
||||
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
||||
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
|
||||
|
||||
---
|
||||
|
||||
# Roadmap (decided 2026-06-12)
|
||||
|
||||
Priorities reflect the tool's primary purpose — a trustworthy backup so that
|
||||
conversation data is not lost if a provider account is ever closed — plus the
|
||||
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
|
||||
local, manually-run CLI; the headless/StartOS direction was dropped
|
||||
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
|
||||
(manual refresh is sufficient).
|
||||
|
||||
**Now (in order):**
|
||||
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
|
||||
(full-archive `export --force` re-export completed 2026-06-13)
|
||||
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
|
||||
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
|
||||
needs admin + SYSTEM impersonation that AV flags as credential theft, and
|
||||
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
|
||||
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
|
||||
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
|
||||
|
||||
**Soon, but later:**
|
||||
|
||||
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
|
||||
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
|
||||
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
|
||||
manually-run CLI, so no in-app polling loop is needed
|
||||
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
|
||||
sufficient (source convos live in the cloud and can be re-downloaded;
|
||||
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
|
||||
also retires the headless token-freshness problem — manual weekly refresh
|
||||
is fine.
|
||||
|
||||
**Active:**
|
||||
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
|
||||
via the `/api/auth/session` `error` field (not `expires`). See §9.
|
||||
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
|
||||
`canary` command. See §10.
|
||||
|
||||
**Deprioritized** (entries kept at the bottom of this file; revisit on
|
||||
demand): `--force` flags, per-conversation cache reset, official export-ZIP
|
||||
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
|
||||
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
|
||||
scope — no significant usage to archive.
|
||||
|
||||
---
|
||||
|
||||
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below, with one scoping correction
|
||||
from live recon: the retrieval dumps are NOT flagged
|
||||
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
|
||||
the discriminator (the hidden flag only marks Custom Instructions and small
|
||||
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
|
||||
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
|
||||
completed 2026-06-13.
|
||||
|
||||
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
|
||||
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
|
||||
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
|
||||
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
|
||||
the user's own attached documents, the same files dumped dozens of times
|
||||
per conversation. These messages were invisible in the ChatGPT web UI;
|
||||
they appear in exports because v0.4.0 lifted the role filter to fix silent
|
||||
data loss. Custom Instructions hidden-context blocks are a minor secondary
|
||||
case (~2KB, once per conversation) — the originally planned
|
||||
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
|
||||
file contains zero hidden-context blocks.
|
||||
|
||||
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
|
||||
|
||||
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
|
||||
`placeholder`, plus a `--hidden-content` CLI override on `export`.
|
||||
- `placeholder` renders affected messages as one line with type and size:
|
||||
`> 🔧 Tool output (file_search, 24KB) — omitted
|
||||
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
|
||||
roughly halves; worst files shrink ~90%.
|
||||
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
|
||||
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
|
||||
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
|
||||
decision must happen in the provider, not the renderer); (b) messages
|
||||
flagged `is_visually_hidden_from_conversation`, including
|
||||
`user_editable_context` / `model_editable_context` (Custom
|
||||
Instructions) — subsumes the old suppress-hidden-context idea.
|
||||
- **Scope — keep at full size:** code-execution `tool_result` blocks and
|
||||
web-search results; those are usually content the user wants.
|
||||
- Count collapsed messages in the post-export summary so the omission
|
||||
stays visible (mirror the LossReport presentation, but as intentional
|
||||
policy, not loss).
|
||||
|
||||
Re-export workflow after shipping: `cache --clear` + `export` (same as
|
||||
the v0.4.0 migration).
|
||||
|
||||
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
|
||||
|
||||
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
|
||||
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
|
||||
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
|
||||
different mechanism (see below).
|
||||
|
||||
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
|
||||
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
|
||||
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
|
||||
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
|
||||
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
|
||||
symptom hit in the field:
|
||||
|
||||
```
|
||||
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
|
||||
```
|
||||
|
||||
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
|
||||
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
|
||||
rights** and behavior that AV/EDR flags as infostealer activity. The one
|
||||
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
|
||||
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
|
||||
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
|
||||
this, so no off-disk reader is a reliable, non-invasive fit.
|
||||
|
||||
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
|
||||
(launch the browser with `--remote-debugging-port`, read cookies via
|
||||
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
|
||||
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
|
||||
hand. With the headless/StartOS direction dropped (#8), manual DevTools
|
||||
refresh is the accepted approach — no automated extraction is needed.
|
||||
|
||||
## 3. Claude Code Session Provider — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
|
||||
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
|
||||
during implementation: `isSidechain` records are subagent transcripts (skipped),
|
||||
`isMeta` marks harness-generated user records (skipped), and listing/normalized
|
||||
`updated_at` must both use file mtime or the cache would re-export every
|
||||
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
|
||||
|
||||
Archive local Claude Code session transcripts. No tokens, no rate limits,
|
||||
no ToS risk — the data is already on disk but lives in a single JSONL per
|
||||
session that Claude Code may clean up, and it contains deliverables
|
||||
(reviews, plans, analyses) that exist nowhere else.
|
||||
|
||||
Decisions (2026-06-12):
|
||||
- **Rendering: prose-only.** Keep user prompts and assistant text
|
||||
(including full deliverable write-ups); collapse tool activity to
|
||||
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
|
||||
86KB — omitted`); **exclude thinking blocks**.
|
||||
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
|
||||
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
|
||||
in v0.5.0).
|
||||
|
||||
Data facts (measured 2026-06-12):
|
||||
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
|
||||
Currently 29 sessions, 18.8MB total, largest 3.8MB.
|
||||
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
|
||||
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
|
||||
activity is what makes these exports readable.
|
||||
- Record types: `user` / `assistant` (Anthropic-style `message.content`
|
||||
block arrays) plus harness records: `ai-title` (use for note title and
|
||||
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
|
||||
`permission-mode`, `system` (skip). Strip harness noise from user
|
||||
messages (`<local-command-caveat>`, `<command-name>` blocks).
|
||||
|
||||
Implementation shape: new `src/providers/claude_code.py` implementing the
|
||||
`BaseProvider` interface — `list_conversations` scans project dirs,
|
||||
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
|
||||
existing block schema (content is already block-shaped: text / tool_use /
|
||||
tool_result / thinking). Incremental sync via file mtime/size recorded in
|
||||
the existing manifest. Project name derives from the munged cwd dirname.
|
||||
|
||||
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
|
||||
(refuses to prune right after `cache --clear`). First live run removed 420
|
||||
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
|
||||
manifest↔disk integrity (293/293 after the run).
|
||||
|
||||
A backup is only trustworthy if the on-disk tree matches the manifest.
|
||||
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
|
||||
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
|
||||
v0.4.1) sit alongside current exports and would double-sync into Joplin.
|
||||
|
||||
- `prune` command: delete export files not referenced by the manifest.
|
||||
`--dry-run` (default off, but always print the list before deleting)
|
||||
shows what would be removed and why (old layout / orphan / unknown).
|
||||
- `doctor` extension: verify every manifest entry's `file_path` exists on
|
||||
disk; report missing files (re-export candidates) and unreferenced files
|
||||
(prune candidates).
|
||||
|
||||
## 5. Binary Content Downloads — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
|
||||
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
|
||||
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
|
||||
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
|
||||
`download_url`; a second GET yields the bytes (works for user uploads;
|
||||
older AI-generated images 404 — expired server-side, handled gracefully).
|
||||
Asset refs come in three shapes — `sediment://file_…`,
|
||||
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
|
||||
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
|
||||
images-only is the default and audio is opt-in via `all`.
|
||||
|
||||
Original notes below.
|
||||
|
||||
**Priority note (2026-06-12): "later, but soon" — under the
|
||||
backup-if-account-closes goal, embedded images are part of the data that
|
||||
would be lost; placeholders alone don't preserve them.**
|
||||
|
||||
v0.4.0 ships placeholders for images and audio assets but does not download
|
||||
the binary content. The `_safe_fence`-wrapped placeholders include the asset
|
||||
reference (`sediment://...` or `file-service://...`), MIME type, size, and
|
||||
duration where available; the actual bytes are not preserved.
|
||||
|
||||
Next steps:
|
||||
- Download attached images alongside the Markdown export, save under a
|
||||
`media/` sibling directory with a stable filename derived from the asset
|
||||
reference.
|
||||
- Replace `image_placeholder` rendering with an inline ``
|
||||
reference once the file is on disk.
|
||||
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
|
||||
rewrite the rendered Markdown to use `:/resourceId` references, and track
|
||||
the resource ID in the cache manifest so re-syncs stay idempotent.
|
||||
- DALL-E images on the assistant side: not observed in this user's data; the
|
||||
code path exists (`source = "model_generated"`) but is untested.
|
||||
|
||||
The block-level schema is already in place — only the file-fetch + rewrite
|
||||
layer needs to be added. See the `image_placeholder` and `file_placeholder`
|
||||
block definitions in `src/blocks.py`.
|
||||
|
||||
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
|
||||
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
|
||||
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
|
||||
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
|
||||
the re-run picked it up.
|
||||
|
||||
Cap how many conversations are downloaded in a single `export` run so the
|
||||
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
|
||||
most importantly on the very first export, which otherwise fetches the
|
||||
entire conversation history in one session. Because every run is resumable
|
||||
(the manifest records each conversation immediately), a capped run simply
|
||||
exports the first N pending conversations and the next run picks up where
|
||||
it left off. This keeps traffic looking like a human-paced session rather
|
||||
than a scraper, reducing the risk of rate limiting or account flags.
|
||||
|
||||
Two complementary pieces:
|
||||
|
||||
1. **Session cap** — `--max-conversations N` flag (and
|
||||
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
|
||||
`export` command, slice the pending list after the cache filter:
|
||||
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
|
||||
counts (reuse the message format from the 429 early-exit path) and
|
||||
remind the user to re-run to continue.
|
||||
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
|
||||
random jitter) slept between per-conversation detail fetches in
|
||||
`BaseProvider`, so even a capped run doesn't fire requests
|
||||
back-to-back. The existing 429 backoff in `_request` stays as the
|
||||
reactive safety net.
|
||||
|
||||
Note: the conversation *listing* (paginated, 100/page) still runs in full
|
||||
each time so the cache comparison works — the cap applies to the heavy
|
||||
per-conversation detail fetches, which dominate request volume.
|
||||
|
||||
This is a stepping stone to the StartOS service: a capped, politely-paced
|
||||
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
|
||||
traffic profile a headless deployment needs.
|
||||
|
||||
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
|
||||
|
||||
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
|
||||
whatever hosts the tool: a user cron line locally, and on the long-term
|
||||
StartOS target the platform's own scheduling. Either way the tool only needs
|
||||
to do one capped, politely-paced `export` + `joplin` run and exit — which it
|
||||
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
|
||||
that chains `export` then `joplin` for a single cron line is a trivial
|
||||
add-on, but the polling loop itself is off the roadmap.
|
||||
|
||||
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
|
||||
|
||||
Not pursuing a headless StartOS service. The local, manually-run CLI is
|
||||
sufficient: the source conversations live in the providers' clouds and can be
|
||||
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
|
||||
so durability is covered without a server in the loop.
|
||||
|
||||
Dropping this also retires the one genuinely hard sub-problem it carried —
|
||||
session-token freshness without a browser. There is no headless context to
|
||||
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
|
||||
extraction remains a dead end regardless — see #2.)
|
||||
|
||||
### REOPENED as a TODO (2026-08-18) — centralization, not durability
|
||||
|
||||
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
|
||||
decision rested on "the source conversations live in the providers' clouds and
|
||||
can be re-downloaded". **That is no longer true of half the providers.**
|
||||
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
|
||||
that exist *only* on the machine that produced them — Codex prunes its rollout
|
||||
files, and neither is recoverable from any cloud. The re-download premise now
|
||||
covers the web providers only.
|
||||
|
||||
The new motivation is consolidation rather than durability: work is split across
|
||||
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
|
||||
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
|
||||
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
|
||||
conversations from everywhere, instead of per-machine islands that only meet
|
||||
inside Joplin.
|
||||
|
||||
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
|
||||
running the exporter locally and keeps doing what it is uniquely able to do —
|
||||
read that machine's local transcripts, and hold the browser session for the web
|
||||
providers. What changes is where the output goes: instead of syncing to Joplin
|
||||
itself, a local run **uploads its conversations to the StartOS storage area**,
|
||||
and the StartOS service owns the Joplin connection for the whole corpus.
|
||||
|
||||
That inverts today's arrangement, where every machine talks to its own Joplin
|
||||
desktop, and it removes two problems we already have:
|
||||
|
||||
- **The Joplin-availability race disappears from the clients.** A scheduled run
|
||||
currently has to find Joplin desktop open on that same machine — the
|
||||
2026-08-18 09:02 timer run exported fine and then skipped the sync because
|
||||
Joplin did not start until 09:07. Uploading to a server that is always up has
|
||||
no such window, and `--joplin-optional` stops being load-bearing.
|
||||
- **One Joplin integration instead of N.** Notebook naming, resource upload and
|
||||
note-update logic run once, server-side, against one manifest — rather than
|
||||
each machine independently deciding what a notebook is called and racing to
|
||||
update the same note.
|
||||
|
||||
What this would need, and what it would *not*:
|
||||
|
||||
- **Not** a headless web-provider login. The hard sub-problem the original drop
|
||||
retired stays retired: the web providers can keep running interactively on the
|
||||
machine that has the browser, pushing their output to the server. Only the
|
||||
local providers need to run server-side, and they need no tokens at all.
|
||||
- An upload step in the client — the counterpart of today's `joplin` command,
|
||||
pointed at the StartOS service instead of a local Joplin API. Probably a
|
||||
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
|
||||
- Per-machine identity in the corpus, which the exporter currently does not
|
||||
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
|
||||
per-machine label". Centralizing would make that label load-bearing.
|
||||
- Conflict handling for one conversation seen by two machines, and a decision
|
||||
about whether the server or the client owns the cache manifest. It is
|
||||
per-machine today, and that is what makes "already up to date" mean anything.
|
||||
- A story for what the client keeps locally after a successful upload. Exports
|
||||
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
|
||||
rollouts, so the client should probably keep them rather than move them.
|
||||
|
||||
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
|
||||
where the archive is actually read. This is a "nice eventually", not a gap.
|
||||
|
||||
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
|
||||
|
||||
Shipped: `doctor` now adds a "ChatGPT token active" check via
|
||||
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
|
||||
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
|
||||
decode path was removed, and `_fetch_access_token` now fails fast on a set
|
||||
`error` instead of returning a stale token. Tests in
|
||||
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
|
||||
below for the record.
|
||||
|
||||
Goal: if it's cheap to tell how much longer a token will work, show it on
|
||||
`doctor`. Findings from live recon:
|
||||
|
||||
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
|
||||
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
|
||||
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
|
||||
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
|
||||
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
|
||||
never yields an expiry for the real tokens (falls to the "not decodable"
|
||||
branch).
|
||||
- **`/api/auth/session` exposes an `expires`** (the provider already calls
|
||||
this endpoint in `_fetch_access_token`; the response includes `expires`
|
||||
alongside `accessToken`). Live value observed 2026-06-28:
|
||||
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
|
||||
assumption and the lived weekly-refresh cadence, so it is almost certainly
|
||||
the NextAuth **rolling session window** (re-extended on every call), not
|
||||
the point at which the pasted token actually 401s. Displaying it verbatim
|
||||
would give false confidence.
|
||||
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
|
||||
was actually dead (conversations API → 401), `/api/auth/session` still
|
||||
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
|
||||
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
|
||||
So `expires` is a **rolling session window that rolls forward on every call
|
||||
even for a dead token**; displaying it would actively lie. The same response
|
||||
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
|
||||
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
|
||||
token `error` is absent/null; when the session token is dead NextAuth can't
|
||||
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
|
||||
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
|
||||
health check.
|
||||
- **Design (ready to build):**
|
||||
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
|
||||
falsy and `accessToken` present; on `RefreshAccessTokenError` report
|
||||
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
|
||||
work) and do NOT surface `expires`.
|
||||
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
|
||||
without checking `error`, so it proceeds with a stale token and yields a
|
||||
confusing downstream 401 instead of a clear "refresh your token" message.
|
||||
Check `error` there and fail fast.
|
||||
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
|
||||
|
||||
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
|
||||
|
||||
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
|
||||
ChatGPT and Claude). It fetches one listing page + one conversation per
|
||||
provider and asserts only the normalizer's load-bearing fields, emitting
|
||||
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
|
||||
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
|
||||
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
|
||||
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
|
||||
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
|
||||
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
|
||||
trail below for the record.
|
||||
|
||||
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
|
||||
|
||||
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
|
||||
change shape without notice. The worst failure for a backup tool is *silent*:
|
||||
a response-schema change that makes the exporter skip or mis-parse content
|
||||
without erroring. `doctor` currently checks token validity, reachability, and
|
||||
manifest↔disk integrity — but not "does the provider's response still look
|
||||
like what the parser expects."
|
||||
|
||||
To investigate: a lightweight schema/shape assertion on a known-good sample
|
||||
of each provider's listing + conversation-detail responses (presence and type
|
||||
of the fields the normalizers rely on), surfaced as a `doctor` check or a
|
||||
dedicated canary.
|
||||
|
||||
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
|
||||
|
||||
- **Dependency surface (assert ONLY these — see below for why):**
|
||||
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
|
||||
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
|
||||
`project.name`, `chat_messages[]`.
|
||||
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
|
||||
(list of typed blocks), `created_at`.
|
||||
- **Key finding — full-shape diffing is the wrong design.** Claude's
|
||||
`settings` object is full of volatile internal codenames that churn
|
||||
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
|
||||
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
|
||||
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
|
||||
canary would fire on every UI experiment. The canary MUST target the
|
||||
normalizer's load-bearing fields only, not the whole response. (Aligns with
|
||||
the drop-noise-don't-retain-it principle.)
|
||||
- **Real Claude messages are flat `text`/`sender`** — in this archive every
|
||||
message had a string `text` and NO `content` block list (0 rich blocks
|
||||
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
|
||||
thinking, image, …) is an **unexercised theoretical path**; drift there
|
||||
can't be "caught" by a canary because it never runs on real data — it's a
|
||||
safety net for if Claude ever switches to block content. The canary should
|
||||
assert the flat shape and *warn if `content` ever appears as a list* (that
|
||||
itself is the drift event that would activate the dormant code).
|
||||
- **Possible silent-loss spot (separate from drift):** Claude messages carry
|
||||
`attachments` and `files` arrays (empty in this sample) that the normalizer
|
||||
ignores entirely. If a user ever attaches files in Claude, they'd be
|
||||
dropped without a LossReport entry. Worth a follow-up check.
|
||||
**Live recon — ChatGPT captured 2026-06-28:**
|
||||
|
||||
- **Dependency surface (assert ONLY these):**
|
||||
- listing item: `id`, `title`, `update_time`/`create_time`.
|
||||
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
|
||||
`update_time`, `mapping` (non-empty).
|
||||
- mapping node: `message`, `children` (the tree walk depends on both);
|
||||
message: `author.role`, `author.name`, `content.content_type`,
|
||||
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
|
||||
- **content_type vocabulary observed (all currently handled):** `text`,
|
||||
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
|
||||
`execution_output`, `reasoning_recap`, `user_editable_context`,
|
||||
`tether_browsing_display`. A *new* content_type already degrades gracefully
|
||||
(visible `unknown` block + WARNING + LossReport tally) — so content_type
|
||||
drift is **already non-silent**. The canary just needs to confirm the known
|
||||
set still parses to non-empty blocks.
|
||||
- **The genuinely silent drift risk — `author.name` collapse keys.**
|
||||
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
|
||||
confirms `file_search` is live (and `web.run`/`python` are correctly left
|
||||
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
|
||||
stops** and the archive re-bloats with no error or LossReport entry. This is
|
||||
the top canary target: assert that retrieval-dump tool authors are still
|
||||
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
|
||||
- **Second silent risk — empty `content.parts`.** A `text` message whose
|
||||
`parts` field is renamed/emptied yields zero blocks and is skipped with only
|
||||
a debug log = silent loss. Canary should assert a sampled `text` message
|
||||
produces a non-empty block.
|
||||
|
||||
**Recon complete for both providers. Canary design (ready to build):** a
|
||||
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
|
||||
listing page + one conversation and asserts the dependency-surface fields
|
||||
above by presence+type — NOT full shape (Claude `settings` codenames prove
|
||||
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
|
||||
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
|
||||
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
|
||||
surface as a warning, never a hard error (a backup tool must still run).
|
||||
|
||||
---
|
||||
|
||||
# Deprioritized — CLOSED as not needed (2026-06-28)
|
||||
|
||||
These were considered and intentionally **not** built. Closed, not planned —
|
||||
the tool is feature-complete for its purpose. Kept for reference in case a
|
||||
real need ever revives one: Joplin `--force`, per-conversation cache reset,
|
||||
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
|
||||
output, token-expiry notifications (also moot — see §9), and a search
|
||||
command. (Also closed, outside this list: handling Claude `attachments`/
|
||||
`files`, which the canary will flag if they ever appear in real data.)
|
||||
|
||||
## Export `--force` Flag — SHIPPED v0.6.0
|
||||
|
||||
Implemented 2026-06-12: `export --force` passes `force=True` to
|
||||
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
|
||||
preserves Joplin links across re-exports, so a forced re-render + `joplin`
|
||||
updates existing notes instead of duplicating them.
|
||||
|
||||
## Joplin `--force` Flag
|
||||
|
||||
Similarly, add `--force` to the `joplin` command to re-sync all cached
|
||||
conversations to Joplin regardless of whether they've been synced before.
|
||||
Useful after making formatting changes to the Markdown exporter.
|
||||
|
||||
Implementation: in `get_joplin_pending()`, return all entries that have a
|
||||
`file_path` when `force=True`, ignoring `joplin_synced_at`.
|
||||
|
||||
## Per-Conversation Cache Reset
|
||||
|
||||
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
|
||||
single conversation without clearing the entire provider cache.
|
||||
|
||||
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
|
||||
delete the entry, then re-run export.
|
||||
|
||||
## Official API Fallback
|
||||
|
||||
If the unofficial internal web API approach breaks, migrate to official export
|
||||
file parsing as a fallback:
|
||||
- ChatGPT: parse `conversations.json` from Settings → Export Data
|
||||
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
|
||||
|
||||
The `BaseProvider` abstract class is intentionally designed so that a
|
||||
`FileProvider` subclass can implement the same interface
|
||||
(`list_conversations`, `get_conversation`, `normalize_conversation`)
|
||||
without any changes to cache, exporters, or CLI code.
|
||||
|
||||
To add this: implement `src/providers/file_chatgpt.py` and
|
||||
`src/providers/file_claude.py`, then add `--input-file` flag to the
|
||||
export command to accept a pre-downloaded export ZIP or JSON.
|
||||
|
||||
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
|
||||
this user needs (project data), so it isn't a real fallback here.
|
||||
|
||||
## Reclassify o1/o3 Reasoning Subparts
|
||||
|
||||
v0.4.0 leaves dict parts inside `text` content_type messages with shape
|
||||
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
|
||||
shape was inferred from a code comment, not captured live). Once a real
|
||||
reasoning conversation is captured, reclassify these as `thinking` blocks.
|
||||
|
||||
## Obsidian Vault Output
|
||||
|
||||
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
|
||||
conversations into an Obsidian vault directory. The current Markdown format
|
||||
is already largely compatible; the main differences are:
|
||||
|
||||
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
|
||||
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
|
||||
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
|
||||
supports both
|
||||
|
||||
Implementation: the existing `MarkdownExporter` output is already valid in
|
||||
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
|
||||
copy files to the vault directory and maintain a flat or nested folder
|
||||
structure matching the user's Obsidian setup. No API needed — just file I/O.
|
||||
|
||||
## Token Expiry Notifications
|
||||
|
||||
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
|
||||
"proactively notify before expiry" idea is blocked by the same finding: a
|
||||
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
|
||||
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
|
||||
window rather than the real refresh cadence). Any heads-up — a `doctor`
|
||||
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
|
||||
resolving the §9 open question of what signal is actually trustworthy.
|
||||
|
||||
## Search Command
|
||||
|
||||
Add a `search` command to full-text search across all exported Markdown files:
|
||||
|
||||
```bash
|
||||
python -m src.main search "kubernetes ingress"
|
||||
python -m src.main search "kubernetes ingress" --provider claude --project devops
|
||||
```
|
||||
|
||||
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
|
||||
conversation title, date, and a snippet. No index needed — Markdown files are
|
||||
small enough to grep directly.
|
||||
|
||||
## Split the README into Separate Documents
|
||||
|
||||
**TODO (2026-08-18).** The README is 830 lines / 5,562 words / 37 KB — about a
|
||||
25-minute read, with 61 headings. An H2-only table of contents was added the
|
||||
same day and helps navigation, but it treats the symptom: the file is doing at
|
||||
least four unrelated jobs at once.
|
||||
|
||||
Rough shape of a split:
|
||||
|
||||
| Document | Content today |
|
||||
|----------|---------------|
|
||||
| `README.md` | What it is, install, first run, a pointer to the rest |
|
||||
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
|
||||
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
|
||||
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
|
||||
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
|
||||
|
||||
Not done yet, and not urgent, because it has a real cost the TOC does not: any
|
||||
existing link into a README section (a bookmark, a note, another repo, a commit
|
||||
message) breaks when that section moves to another file. Worth doing when the
|
||||
README next needs substantial editing anyway, rather than as a change of its own.
|
||||
|
||||
Two things to decide when it happens:
|
||||
|
||||
- Whether `docs/` renders acceptably on the Gitea instance that hosts this repo
|
||||
(relative links between Markdown files do work there, but worth confirming
|
||||
before splitting rather than after).
|
||||
- Whether the anchors in the split files stay stable enough to link *between*
|
||||
documents, or whether cross-references should point at file tops only. GFM
|
||||
anchors are derived from heading text, so they break silently on a reword —
|
||||
the same fragility the TOC already carries.
|
||||
@@ -1,616 +1,144 @@
|
||||
# Planned Future Work
|
||||
|
||||
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
|
||||
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
|
||||
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
|
||||
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
|
||||
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
|
||||
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
|
||||
>
|
||||
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
|
||||
> active roadmap is empty and the remaining backlog below has been **closed as
|
||||
> not needed** — the tool does what it's needed to do as a local, manually-run
|
||||
> backup CLI. Items are kept for reference only; revisit on demand if a real
|
||||
> need shows up. Nothing here is planned work.
|
||||
> **Status 2026-08-18 (v0.9.0).** The tool archives four providers — two web
|
||||
> (`chatgpt`, `claude`) and two local agent-transcript (`claude-code`, `codex`)
|
||||
> — on a schedule, on Linux and Windows, reporting results by push notification.
|
||||
> Two items below are genuinely planned. Everything else has shipped or been
|
||||
> decided against.
|
||||
|
||||
Items completed in each release are moved to the changelog. Items below the
|
||||
roadmap were designed for but intentionally not implemented. The codebase is
|
||||
structured to make each of these additions straightforward if ever revived.
|
||||
|
||||
**Completed:**
|
||||
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
||||
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
||||
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
|
||||
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
||||
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
|
||||
Completed work moves to the changelog; this file holds only what is *not* built
|
||||
yet. Decisions not to build something are recorded at the bottom in one line
|
||||
each, so they are not re-proposed — the full investigation trails behind them
|
||||
are in `FUTURE-ARCHIVE.md`.
|
||||
|
||||
---
|
||||
|
||||
# Roadmap (decided 2026-06-12)
|
||||
# Roadmap
|
||||
|
||||
Priorities reflect the tool's primary purpose — a trustworthy backup so that
|
||||
conversation data is not lost if a provider account is ever closed — plus the
|
||||
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
|
||||
local, manually-run CLI; the headless/StartOS direction was dropped
|
||||
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
|
||||
(manual refresh is sufficient).
|
||||
## 1. StartOS Service — one corpus, not per-machine islands
|
||||
|
||||
**Now (in order):**
|
||||
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
|
||||
(full-archive `export --force` re-export completed 2026-06-13)
|
||||
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
|
||||
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
|
||||
needs admin + SYSTEM impersonation that AV flags as credential theft, and
|
||||
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
|
||||
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
|
||||
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
|
||||
Each machine currently archives to its own `exports/` and its own Joplin. Work
|
||||
is split across boxes — coding sessions (`claude-code`, `codex`) on the Linux
|
||||
machine, web chats (`chatgpt`, `claude`) on the Windows one — so there is no
|
||||
single place where all conversations exist together.
|
||||
|
||||
**Soon, but later:**
|
||||
|
||||
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
|
||||
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
|
||||
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
|
||||
manually-run CLI, so no in-app polling loop is needed
|
||||
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
|
||||
sufficient (source convos live in the cloud and can be re-downloaded;
|
||||
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
|
||||
also retires the headless token-freshness problem — manual weekly refresh
|
||||
is fine.
|
||||
|
||||
**Active:**
|
||||
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
|
||||
via the `/api/auth/session` `error` field (not `expires`). See §9.
|
||||
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
|
||||
`canary` command. See §10.
|
||||
|
||||
**Deprioritized** (entries kept at the bottom of this file; revisit on
|
||||
demand): `--force` flags, per-conversation cache reset, official export-ZIP
|
||||
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
|
||||
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
|
||||
scope — no significant usage to archive.
|
||||
|
||||
---
|
||||
|
||||
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below, with one scoping correction
|
||||
from live recon: the retrieval dumps are NOT flagged
|
||||
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
|
||||
the discriminator (the hidden flag only marks Custom Instructions and small
|
||||
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
|
||||
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
|
||||
completed 2026-06-13.
|
||||
|
||||
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
|
||||
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
|
||||
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
|
||||
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
|
||||
the user's own attached documents, the same files dumped dozens of times
|
||||
per conversation. These messages were invisible in the ChatGPT web UI;
|
||||
they appear in exports because v0.4.0 lifted the role filter to fix silent
|
||||
data loss. Custom Instructions hidden-context blocks are a minor secondary
|
||||
case (~2KB, once per conversation) — the originally planned
|
||||
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
|
||||
file contains zero hidden-context blocks.
|
||||
|
||||
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
|
||||
|
||||
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
|
||||
`placeholder`, plus a `--hidden-content` CLI override on `export`.
|
||||
- `placeholder` renders affected messages as one line with type and size:
|
||||
`> 🔧 Tool output (file_search, 24KB) — omitted
|
||||
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
|
||||
roughly halves; worst files shrink ~90%.
|
||||
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
|
||||
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
|
||||
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
|
||||
decision must happen in the provider, not the renderer); (b) messages
|
||||
flagged `is_visually_hidden_from_conversation`, including
|
||||
`user_editable_context` / `model_editable_context` (Custom
|
||||
Instructions) — subsumes the old suppress-hidden-context idea.
|
||||
- **Scope — keep at full size:** code-execution `tool_result` blocks and
|
||||
web-search results; those are usually content the user wants.
|
||||
- Count collapsed messages in the post-export summary so the omission
|
||||
stays visible (mirror the LossReport presentation, but as intentional
|
||||
policy, not loss).
|
||||
|
||||
Re-export workflow after shipping: `cache --clear` + `export` (same as
|
||||
the v0.4.0 migration).
|
||||
|
||||
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
|
||||
|
||||
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
|
||||
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
|
||||
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
|
||||
different mechanism (see below).
|
||||
|
||||
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
|
||||
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
|
||||
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
|
||||
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
|
||||
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
|
||||
symptom hit in the field:
|
||||
|
||||
```
|
||||
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
|
||||
```
|
||||
|
||||
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
|
||||
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
|
||||
rights** and behavior that AV/EDR flags as infostealer activity. The one
|
||||
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
|
||||
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
|
||||
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
|
||||
this, so no off-disk reader is a reliable, non-invasive fit.
|
||||
|
||||
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
|
||||
(launch the browser with `--remote-debugging-port`, read cookies via
|
||||
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
|
||||
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
|
||||
hand. With the headless/StartOS direction dropped (#8), manual DevTools
|
||||
refresh is the accepted approach — no automated extraction is needed.
|
||||
|
||||
## 3. Claude Code Session Provider — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
|
||||
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
|
||||
during implementation: `isSidechain` records are subagent transcripts (skipped),
|
||||
`isMeta` marks harness-generated user records (skipped), and listing/normalized
|
||||
`updated_at` must both use file mtime or the cache would re-export every
|
||||
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
|
||||
|
||||
Archive local Claude Code session transcripts. No tokens, no rate limits,
|
||||
no ToS risk — the data is already on disk but lives in a single JSONL per
|
||||
session that Claude Code may clean up, and it contains deliverables
|
||||
(reviews, plans, analyses) that exist nowhere else.
|
||||
|
||||
Decisions (2026-06-12):
|
||||
- **Rendering: prose-only.** Keep user prompts and assistant text
|
||||
(including full deliverable write-ups); collapse tool activity to
|
||||
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
|
||||
86KB — omitted`); **exclude thinking blocks**.
|
||||
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
|
||||
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
|
||||
in v0.5.0).
|
||||
|
||||
Data facts (measured 2026-06-12):
|
||||
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
|
||||
Currently 29 sessions, 18.8MB total, largest 3.8MB.
|
||||
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
|
||||
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
|
||||
activity is what makes these exports readable.
|
||||
- Record types: `user` / `assistant` (Anthropic-style `message.content`
|
||||
block arrays) plus harness records: `ai-title` (use for note title and
|
||||
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
|
||||
`permission-mode`, `system` (skip). Strip harness noise from user
|
||||
messages (`<local-command-caveat>`, `<command-name>` blocks).
|
||||
|
||||
Implementation shape: new `src/providers/claude_code.py` implementing the
|
||||
`BaseProvider` interface — `list_conversations` scans project dirs,
|
||||
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
|
||||
existing block schema (content is already block-shaped: text / tool_use /
|
||||
tool_result / thinking). Incremental sync via file mtime/size recorded in
|
||||
the existing manifest. Project name derives from the munged cwd dirname.
|
||||
|
||||
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
|
||||
(refuses to prune right after `cache --clear`). First live run removed 420
|
||||
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
|
||||
manifest↔disk integrity (293/293 after the run).
|
||||
|
||||
A backup is only trustworthy if the on-disk tree matches the manifest.
|
||||
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
|
||||
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
|
||||
v0.4.1) sit alongside current exports and would double-sync into Joplin.
|
||||
|
||||
- `prune` command: delete export files not referenced by the manifest.
|
||||
`--dry-run` (default off, but always print the list before deleting)
|
||||
shows what would be removed and why (old layout / orphan / unknown).
|
||||
- `doctor` extension: verify every manifest entry's `file_path` exists on
|
||||
disk; report missing files (re-export candidates) and unreferenced files
|
||||
(prune candidates).
|
||||
|
||||
## 5. Binary Content Downloads — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
|
||||
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
|
||||
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
|
||||
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
|
||||
`download_url`; a second GET yields the bytes (works for user uploads;
|
||||
older AI-generated images 404 — expired server-side, handled gracefully).
|
||||
Asset refs come in three shapes — `sediment://file_…`,
|
||||
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
|
||||
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
|
||||
images-only is the default and audio is opt-in via `all`.
|
||||
|
||||
Original notes below.
|
||||
|
||||
**Priority note (2026-06-12): "later, but soon" — under the
|
||||
backup-if-account-closes goal, embedded images are part of the data that
|
||||
would be lost; placeholders alone don't preserve them.**
|
||||
|
||||
v0.4.0 ships placeholders for images and audio assets but does not download
|
||||
the binary content. The `_safe_fence`-wrapped placeholders include the asset
|
||||
reference (`sediment://...` or `file-service://...`), MIME type, size, and
|
||||
duration where available; the actual bytes are not preserved.
|
||||
|
||||
Next steps:
|
||||
- Download attached images alongside the Markdown export, save under a
|
||||
`media/` sibling directory with a stable filename derived from the asset
|
||||
reference.
|
||||
- Replace `image_placeholder` rendering with an inline ``
|
||||
reference once the file is on disk.
|
||||
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
|
||||
rewrite the rendered Markdown to use `:/resourceId` references, and track
|
||||
the resource ID in the cache manifest so re-syncs stay idempotent.
|
||||
- DALL-E images on the assistant side: not observed in this user's data; the
|
||||
code path exists (`source = "model_generated"`) but is untested.
|
||||
|
||||
The block-level schema is already in place — only the file-fetch + rewrite
|
||||
layer needs to be added. See the `image_placeholder` and `file_placeholder`
|
||||
block definitions in `src/blocks.py`.
|
||||
|
||||
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
|
||||
|
||||
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
|
||||
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
|
||||
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
|
||||
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
|
||||
the re-run picked it up.
|
||||
|
||||
Cap how many conversations are downloaded in a single `export` run so the
|
||||
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
|
||||
most importantly on the very first export, which otherwise fetches the
|
||||
entire conversation history in one session. Because every run is resumable
|
||||
(the manifest records each conversation immediately), a capped run simply
|
||||
exports the first N pending conversations and the next run picks up where
|
||||
it left off. This keeps traffic looking like a human-paced session rather
|
||||
than a scraper, reducing the risk of rate limiting or account flags.
|
||||
|
||||
Two complementary pieces:
|
||||
|
||||
1. **Session cap** — `--max-conversations N` flag (and
|
||||
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
|
||||
`export` command, slice the pending list after the cache filter:
|
||||
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
|
||||
counts (reuse the message format from the 429 early-exit path) and
|
||||
remind the user to re-run to continue.
|
||||
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
|
||||
random jitter) slept between per-conversation detail fetches in
|
||||
`BaseProvider`, so even a capped run doesn't fire requests
|
||||
back-to-back. The existing 429 backoff in `_request` stays as the
|
||||
reactive safety net.
|
||||
|
||||
Note: the conversation *listing* (paginated, 100/page) still runs in full
|
||||
each time so the cache comparison works — the cap applies to the heavy
|
||||
per-conversation detail fetches, which dominate request volume.
|
||||
|
||||
This is a stepping stone to the StartOS service: a capped, politely-paced
|
||||
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
|
||||
traffic profile a headless deployment needs.
|
||||
|
||||
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
|
||||
|
||||
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
|
||||
whatever hosts the tool: a user cron line locally, and on the long-term
|
||||
StartOS target the platform's own scheduling. Either way the tool only needs
|
||||
to do one capped, politely-paced `export` + `joplin` run and exit — which it
|
||||
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
|
||||
that chains `export` then `joplin` for a single cron line is a trivial
|
||||
add-on, but the polling loop itself is off the roadmap.
|
||||
|
||||
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
|
||||
|
||||
Not pursuing a headless StartOS service. The local, manually-run CLI is
|
||||
sufficient: the source conversations live in the providers' clouds and can be
|
||||
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
|
||||
so durability is covered without a server in the loop.
|
||||
|
||||
Dropping this also retires the one genuinely hard sub-problem it carried —
|
||||
session-token freshness without a browser. There is no headless context to
|
||||
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
|
||||
extraction remains a dead end regardless — see #2.)
|
||||
|
||||
### REOPENED as a TODO (2026-08-18) — centralization, not durability
|
||||
|
||||
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
|
||||
decision rested on "the source conversations live in the providers' clouds and
|
||||
can be re-downloaded". **That is no longer true of half the providers.**
|
||||
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
|
||||
that exist *only* on the machine that produced them — Codex prunes its rollout
|
||||
files, and neither is recoverable from any cloud. The re-download premise now
|
||||
covers the web providers only.
|
||||
|
||||
The new motivation is consolidation rather than durability: work is split across
|
||||
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
|
||||
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
|
||||
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
|
||||
conversations from everywhere, instead of per-machine islands that only meet
|
||||
inside Joplin.
|
||||
This was dropped on 2026-06-28 and **reopened 2026-08-18**, because the
|
||||
reasoning behind the drop has gone stale. It rested on "the source conversations
|
||||
live in the providers' clouds and can be re-downloaded". That is no longer true
|
||||
of half the providers: `claude-code` and `codex` transcripts exist *only* on the
|
||||
machine that produced them, Codex prunes its rollout files, and neither is
|
||||
recoverable from any cloud. The motivation is also different from the one
|
||||
weighed then — consolidation, not durability.
|
||||
|
||||
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
|
||||
running the exporter locally and keeps doing what it is uniquely able to do —
|
||||
read that machine's local transcripts, and hold the browser session for the web
|
||||
providers. What changes is where the output goes: instead of syncing to Joplin
|
||||
itself, a local run **uploads its conversations to the StartOS storage area**,
|
||||
and the StartOS service owns the Joplin connection for the whole corpus.
|
||||
running the exporter locally and keeps doing what only it can do: read that
|
||||
machine's local transcripts, and hold the browser session for the web providers.
|
||||
What changes is where the output goes. Instead of syncing to Joplin itself, a
|
||||
local run **uploads its conversations to the StartOS storage area**, and the
|
||||
StartOS service owns the Joplin connection for the whole corpus.
|
||||
|
||||
That inverts today's arrangement, where every machine talks to its own Joplin
|
||||
desktop, and it removes two problems we already have:
|
||||
That inverts today's arrangement and removes two problems we already have:
|
||||
|
||||
- **The Joplin-availability race disappears from the clients.** A scheduled run
|
||||
currently has to find Joplin desktop open on that same machine — the
|
||||
2026-08-18 09:02 timer run exported fine and then skipped the sync because
|
||||
Joplin did not start until 09:07. Uploading to a server that is always up has
|
||||
no such window, and `--joplin-optional` stops being load-bearing.
|
||||
- **The Joplin-availability race leaves the clients.** A scheduled run currently
|
||||
has to find Joplin desktop open on that same machine — the 09:02 timer run on
|
||||
2026-08-18 exported fine and then skipped the sync because Joplin did not
|
||||
start until 09:07. A server that is always up has no such window, and
|
||||
`--joplin-optional` stops being load-bearing.
|
||||
- **One Joplin integration instead of N.** Notebook naming, resource upload and
|
||||
note-update logic run once, server-side, against one manifest — rather than
|
||||
each machine independently deciding what a notebook is called and racing to
|
||||
update the same note.
|
||||
note updates run once, server-side, against one manifest — rather than each
|
||||
machine independently deciding what a notebook is called and racing to update
|
||||
the same note.
|
||||
|
||||
What this would need, and what it would *not*:
|
||||
What this needs, and what it does *not*:
|
||||
|
||||
- **Not** a headless web-provider login. The hard sub-problem the original drop
|
||||
retired stays retired: the web providers can keep running interactively on the
|
||||
machine that has the browser, pushing their output to the server. Only the
|
||||
local providers need to run server-side, and they need no tokens at all.
|
||||
retired stays retired: the web providers keep running interactively on the
|
||||
machine that has the browser, and push their output to the server. Only the
|
||||
local providers would run server-side, and they need no tokens at all.
|
||||
- An upload step in the client — the counterpart of today's `joplin` command,
|
||||
pointed at the StartOS service instead of a local Joplin API. Probably a
|
||||
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
|
||||
- Per-machine identity in the corpus, which the exporter currently does not
|
||||
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
|
||||
per-machine label". Centralizing would make that label load-bearing.
|
||||
`push` alongside `sync`, so a scheduled client run stays one line.
|
||||
- Per-machine identity in the corpus, which the exporter does not track today:
|
||||
`claude_code.resolve_roots` deliberately merges multiple roots with "no
|
||||
per-machine label". Centralizing makes that label load-bearing.
|
||||
- Conflict handling for one conversation seen by two machines, and a decision
|
||||
about whether the server or the client owns the cache manifest. It is
|
||||
per-machine today, and that is what makes "already up to date" mean anything.
|
||||
- A story for what the client keeps locally after a successful upload. Exports
|
||||
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
|
||||
rollouts, so the client should probably keep them rather than move them.
|
||||
rollouts, so the client should keep them rather than hand them off.
|
||||
|
||||
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
|
||||
where the archive is actually read. This is a "nice eventually", not a gap.
|
||||
|
||||
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
|
||||
## 2. Split the README into separate documents
|
||||
|
||||
Shipped: `doctor` now adds a "ChatGPT token active" check via
|
||||
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
|
||||
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
|
||||
decode path was removed, and `_fetch_access_token` now fails fast on a set
|
||||
`error` instead of returning a stale token. Tests in
|
||||
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
|
||||
below for the record.
|
||||
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61
|
||||
headings. An H2-only table of contents was added 2026-08-18 and helps, but it
|
||||
treats the symptom: the file is doing at least four unrelated jobs at once.
|
||||
|
||||
Goal: if it's cheap to tell how much longer a token will work, show it on
|
||||
`doctor`. Findings from live recon:
|
||||
Rough shape of a split:
|
||||
|
||||
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
|
||||
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
|
||||
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
|
||||
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
|
||||
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
|
||||
never yields an expiry for the real tokens (falls to the "not decodable"
|
||||
branch).
|
||||
- **`/api/auth/session` exposes an `expires`** (the provider already calls
|
||||
this endpoint in `_fetch_access_token`; the response includes `expires`
|
||||
alongside `accessToken`). Live value observed 2026-06-28:
|
||||
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
|
||||
assumption and the lived weekly-refresh cadence, so it is almost certainly
|
||||
the NextAuth **rolling session window** (re-extended on every call), not
|
||||
the point at which the pasted token actually 401s. Displaying it verbatim
|
||||
would give false confidence.
|
||||
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
|
||||
was actually dead (conversations API → 401), `/api/auth/session` still
|
||||
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
|
||||
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
|
||||
So `expires` is a **rolling session window that rolls forward on every call
|
||||
even for a dead token**; displaying it would actively lie. The same response
|
||||
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
|
||||
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
|
||||
token `error` is absent/null; when the session token is dead NextAuth can't
|
||||
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
|
||||
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
|
||||
health check.
|
||||
- **Design (ready to build):**
|
||||
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
|
||||
falsy and `accessToken` present; on `RefreshAccessTokenError` report
|
||||
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
|
||||
work) and do NOT surface `expires`.
|
||||
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
|
||||
without checking `error`, so it proceeds with a stale token and yields a
|
||||
confusing downstream 401 instead of a clear "refresh your token" message.
|
||||
Check `error` there and fail fast.
|
||||
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
|
||||
| Document | Content today |
|
||||
|----------|---------------|
|
||||
| `README.md` | What it is, install, first run, a pointer to the rest |
|
||||
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
|
||||
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
|
||||
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
|
||||
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
|
||||
|
||||
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
|
||||
Not urgent, because it has a real cost the TOC did not: any existing link into a
|
||||
README section — a bookmark, a note, another repo, a commit message — breaks
|
||||
when that section moves to another file. Worth folding into the next substantial
|
||||
README edit rather than doing as a change of its own.
|
||||
|
||||
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
|
||||
ChatGPT and Claude). It fetches one listing page + one conversation per
|
||||
provider and asserts only the normalizer's load-bearing fields, emitting
|
||||
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
|
||||
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
|
||||
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
|
||||
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
|
||||
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
|
||||
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
|
||||
trail below for the record.
|
||||
Two things to settle when it happens:
|
||||
|
||||
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
|
||||
|
||||
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
|
||||
change shape without notice. The worst failure for a backup tool is *silent*:
|
||||
a response-schema change that makes the exporter skip or mis-parse content
|
||||
without erroring. `doctor` currently checks token validity, reachability, and
|
||||
manifest↔disk integrity — but not "does the provider's response still look
|
||||
like what the parser expects."
|
||||
|
||||
To investigate: a lightweight schema/shape assertion on a known-good sample
|
||||
of each provider's listing + conversation-detail responses (presence and type
|
||||
of the fields the normalizers rely on), surfaced as a `doctor` check or a
|
||||
dedicated canary.
|
||||
|
||||
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
|
||||
|
||||
- **Dependency surface (assert ONLY these — see below for why):**
|
||||
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
|
||||
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
|
||||
`project.name`, `chat_messages[]`.
|
||||
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
|
||||
(list of typed blocks), `created_at`.
|
||||
- **Key finding — full-shape diffing is the wrong design.** Claude's
|
||||
`settings` object is full of volatile internal codenames that churn
|
||||
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
|
||||
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
|
||||
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
|
||||
canary would fire on every UI experiment. The canary MUST target the
|
||||
normalizer's load-bearing fields only, not the whole response. (Aligns with
|
||||
the drop-noise-don't-retain-it principle.)
|
||||
- **Real Claude messages are flat `text`/`sender`** — in this archive every
|
||||
message had a string `text` and NO `content` block list (0 rich blocks
|
||||
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
|
||||
thinking, image, …) is an **unexercised theoretical path**; drift there
|
||||
can't be "caught" by a canary because it never runs on real data — it's a
|
||||
safety net for if Claude ever switches to block content. The canary should
|
||||
assert the flat shape and *warn if `content` ever appears as a list* (that
|
||||
itself is the drift event that would activate the dormant code).
|
||||
- **Possible silent-loss spot (separate from drift):** Claude messages carry
|
||||
`attachments` and `files` arrays (empty in this sample) that the normalizer
|
||||
ignores entirely. If a user ever attaches files in Claude, they'd be
|
||||
dropped without a LossReport entry. Worth a follow-up check.
|
||||
**Live recon — ChatGPT captured 2026-06-28:**
|
||||
|
||||
- **Dependency surface (assert ONLY these):**
|
||||
- listing item: `id`, `title`, `update_time`/`create_time`.
|
||||
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
|
||||
`update_time`, `mapping` (non-empty).
|
||||
- mapping node: `message`, `children` (the tree walk depends on both);
|
||||
message: `author.role`, `author.name`, `content.content_type`,
|
||||
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
|
||||
- **content_type vocabulary observed (all currently handled):** `text`,
|
||||
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
|
||||
`execution_output`, `reasoning_recap`, `user_editable_context`,
|
||||
`tether_browsing_display`. A *new* content_type already degrades gracefully
|
||||
(visible `unknown` block + WARNING + LossReport tally) — so content_type
|
||||
drift is **already non-silent**. The canary just needs to confirm the known
|
||||
set still parses to non-empty blocks.
|
||||
- **The genuinely silent drift risk — `author.name` collapse keys.**
|
||||
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
|
||||
confirms `file_search` is live (and `web.run`/`python` are correctly left
|
||||
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
|
||||
stops** and the archive re-bloats with no error or LossReport entry. This is
|
||||
the top canary target: assert that retrieval-dump tool authors are still
|
||||
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
|
||||
- **Second silent risk — empty `content.parts`.** A `text` message whose
|
||||
`parts` field is renamed/emptied yields zero blocks and is skipped with only
|
||||
a debug log = silent loss. Canary should assert a sampled `text` message
|
||||
produces a non-empty block.
|
||||
|
||||
**Recon complete for both providers. Canary design (ready to build):** a
|
||||
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
|
||||
listing page + one conversation and asserts the dependency-surface fields
|
||||
above by presence+type — NOT full shape (Claude `settings` codenames prove
|
||||
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
|
||||
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
|
||||
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
|
||||
surface as a warning, never a hard error (a backup tool must still run).
|
||||
- **`.gitignore` ignores `*.md` on purpose** — exported conversations are
|
||||
Markdown and may contain private content — and re-includes each doc by name
|
||||
(`!README.md`, `!FUTURE.md`, …). New files under `docs/` will be silently
|
||||
ignored, with no error, until `!docs/*.md` is added. This already bit the
|
||||
creation of `FUTURE-ARCHIVE.md` on 2026-08-18.
|
||||
- Whether `docs/` renders acceptably on the Gitea instance hosting this repo.
|
||||
Relative links between Markdown files do work there, but confirm before
|
||||
splitting rather than after.
|
||||
- Whether anchors in the split files are stable enough to link *between*
|
||||
documents, or whether cross-references should point at file tops only. GFM
|
||||
anchors derive from heading text and break silently on a reword — the same
|
||||
fragility the TOC already carries.
|
||||
|
||||
---
|
||||
|
||||
# Deprioritized — CLOSED as not needed (2026-06-28)
|
||||
# Decided against
|
||||
|
||||
These were considered and intentionally **not** built. Closed, not planned —
|
||||
the tool is feature-complete for its purpose. Kept for reference in case a
|
||||
real need ever revives one: Joplin `--force`, per-conversation cache reset,
|
||||
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
|
||||
output, token-expiry notifications (also moot — see §9), and a search
|
||||
command. (Also closed, outside this list: handling Claude `attachments`/
|
||||
`files`, which the canary will flag if they ever appear in real data.)
|
||||
Recorded so they are not re-proposed. Full reasoning and the recon behind each
|
||||
is in `FUTURE-ARCHIVE.md`.
|
||||
|
||||
## Export `--force` Flag — SHIPPED v0.6.0
|
||||
| Item | Verdict |
|
||||
|------|---------|
|
||||
| Brave/Chromium cookie auto-extraction | **Not viable** (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
|
||||
| In-app watch/polling loop | **Dropped** (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as `sync` plus a systemd timer and a Windows scheduled task. |
|
||||
| Proactive token-expiry notification | **Blocked, and moot** (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and `/api/auth/session`'s `expires` is a rolling window that advances even for a dead token. `doctor`'s `error`-based health check covers the real need. |
|
||||
| Official export-ZIP fallback | **Closed** (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. `BaseProvider` still admits a `FileProvider` if that changes. |
|
||||
| Joplin `--force` flag | Closed as not needed (2026-06-28). |
|
||||
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
|
||||
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
|
||||
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
|
||||
| Search command | Closed as not needed (2026-06-28). `grep`/`ripgrep` over `EXPORT_DIR` covers it. |
|
||||
| Claude `attachments` / `files` handling | Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
|
||||
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
|
||||
|
||||
Implemented 2026-06-12: `export --force` passes `force=True` to
|
||||
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
|
||||
preserves Joplin links across re-exports, so a forced re-render + `joplin`
|
||||
updates existing notes instead of duplicating them.
|
||||
---
|
||||
|
||||
## Joplin `--force` Flag
|
||||
# Completed
|
||||
|
||||
Similarly, add `--force` to the `joplin` command to re-sync all cached
|
||||
conversations to Joplin regardless of whether they've been synced before.
|
||||
Useful after making formatting changes to the Markdown exporter.
|
||||
Detail for each release is in `CHANGELOG.md`.
|
||||
|
||||
Implementation: in `get_joplin_pending()`, return all entries that have a
|
||||
`file_path` when `force=True`, ignoring `joplin_synced_at`.
|
||||
|
||||
## Per-Conversation Cache Reset
|
||||
|
||||
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
|
||||
single conversation without clearing the entire provider cache.
|
||||
|
||||
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
|
||||
delete the entry, then re-run export.
|
||||
|
||||
## Official API Fallback
|
||||
|
||||
If the unofficial internal web API approach breaks, migrate to official export
|
||||
file parsing as a fallback:
|
||||
- ChatGPT: parse `conversations.json` from Settings → Export Data
|
||||
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
|
||||
|
||||
The `BaseProvider` abstract class is intentionally designed so that a
|
||||
`FileProvider` subclass can implement the same interface
|
||||
(`list_conversations`, `get_conversation`, `normalize_conversation`)
|
||||
without any changes to cache, exporters, or CLI code.
|
||||
|
||||
To add this: implement `src/providers/file_chatgpt.py` and
|
||||
`src/providers/file_claude.py`, then add `--input-file` flag to the
|
||||
export command to accept a pre-downloaded export ZIP or JSON.
|
||||
|
||||
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
|
||||
this user needs (project data), so it isn't a real fallback here.
|
||||
|
||||
## Reclassify o1/o3 Reasoning Subparts
|
||||
|
||||
v0.4.0 leaves dict parts inside `text` content_type messages with shape
|
||||
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
|
||||
shape was inferred from a code comment, not captured live). Once a real
|
||||
reasoning conversation is captured, reclassify these as `thinking` blocks.
|
||||
|
||||
## Obsidian Vault Output
|
||||
|
||||
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
|
||||
conversations into an Obsidian vault directory. The current Markdown format
|
||||
is already largely compatible; the main differences are:
|
||||
|
||||
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
|
||||
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
|
||||
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
|
||||
supports both
|
||||
|
||||
Implementation: the existing `MarkdownExporter` output is already valid in
|
||||
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
|
||||
copy files to the vault directory and maintain a flat or nested folder
|
||||
structure matching the user's Obsidian setup. No API needed — just file I/O.
|
||||
|
||||
## Token Expiry Notifications
|
||||
|
||||
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
|
||||
"proactively notify before expiry" idea is blocked by the same finding: a
|
||||
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
|
||||
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
|
||||
window rather than the real refresh cadence). Any heads-up — a `doctor`
|
||||
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
|
||||
resolving the §9 open question of what signal is actually trustworthy.
|
||||
|
||||
## Search Command
|
||||
|
||||
Add a `search` command to full-text search across all exported Markdown files:
|
||||
|
||||
```bash
|
||||
python -m src.main search "kubernetes ingress"
|
||||
python -m src.main search "kubernetes ingress" --provider claude --project devops
|
||||
```
|
||||
|
||||
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
|
||||
conversation title, date, and a snippet. No index needed — Markdown files are
|
||||
small enough to grep directly.
|
||||
- **v0.1.0** — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
||||
- **v0.2.0** — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
||||
- **v0.4.0** — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via `LossReport` and visible `unknown` blocks
|
||||
- **v0.5.0** — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
||||
- **v0.6.0** — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT`); Claude Code session provider; `prune` + manifest integrity; binary content downloads; session limiter and request pacing; `export --force`
|
||||
- **v0.7.0** — `canary` drift detection; real ChatGPT token-health check on `doctor`; removal of the non-viable browser-cookie path
|
||||
- **v0.8.0** — Claude Code coverage reopened: subagent capture as folded `<details>`, repo `[tags]` in titles, its own `AI-ClaudeCode` notebook with self-healing note moves, multi-root scanning (`CLAUDE_CODE_DIR`, `CLAUDE_CONFIG_DIR`)
|
||||
- **v0.9.0** — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms; `sync` with a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution via `gizmo_id` and the `projects` command; a silent data-loss fix in both local providers (`splitlines` breaking on U+0085/U+2028/U+2029)
|
||||
|
||||
@@ -6,7 +6,28 @@ Supports incremental sync — only new or updated conversations are exported on
|
||||
|
||||
---
|
||||
|
||||
## ⚠️ Terms of Service Warning
|
||||
## Contents
|
||||
|
||||
- [Terms of Service Warning](#terms-of-service-warning)
|
||||
- [Installation](#installation)
|
||||
- [First Run: Run Doctor](#first-run-run-doctor)
|
||||
- [Getting Your Session Tokens](#getting-your-session-tokens)
|
||||
- [The `auth` Command](#the-auth-command)
|
||||
- [`.env` Setup](#env-setup)
|
||||
- [ChatGPT Projects](#chatgpt-projects)
|
||||
- [Claude Code Sessions](#claude-code-sessions)
|
||||
- [Codex Sessions](#codex-sessions)
|
||||
- [Scheduling a Daily Run](#scheduling-a-daily-run)
|
||||
- [Output Structure](#output-structure)
|
||||
- [CLI Reference](#cli-reference)
|
||||
- [How the Cache Works](#how-the-cache-works)
|
||||
- [Troubleshooting](#troubleshooting)
|
||||
- [Future Work](#future-work)
|
||||
- [Security Notes](#security-notes)
|
||||
|
||||
---
|
||||
|
||||
## Terms of Service Warning
|
||||
|
||||
**Read this before using this tool.**
|
||||
|
||||
@@ -223,12 +244,26 @@ cp .env.example .env
|
||||
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
|
||||
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
|
||||
|
||||
### Notifications
|
||||
|
||||
| Variable | Default | Description |
|
||||
|----------|---------|-------------|
|
||||
| `NTFY_TOPIC` | — | [ntfy](https://ntfy.sh) topic to push run results to. Unset disables notifications entirely. |
|
||||
| `NTFY_SERVER` | `https://ntfy.sh` | Point at your own host if self-hosting. |
|
||||
| `NTFY_TOKEN` | — | Bearer token, for access-controlled topics. |
|
||||
| `NTFY_NOTIFY` | `always` | `always` notifies on every run, `failure` only when something failed, `off` never. |
|
||||
|
||||
A topic on public ntfy.sh is readable by anyone who knows its name, so
|
||||
notifications carry per-provider counts and a machine name only — never
|
||||
conversation titles. See [Getting notified](#getting-notified).
|
||||
|
||||
### Cache & logging
|
||||
|
||||
| Variable | Default | Description |
|
||||
|----------|---------|-------------|
|
||||
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
|
||||
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
|
||||
| `AI_CHAT_EXPORTER_QUIET_CWD` | — | Set to `1` to silence the launcher's warning when run from outside the repo. Read by the `ai-chat-exporter` wrapper scripts, not by Python; the scheduler installers set it, since they always set the correct working directory. |
|
||||
|
||||
---
|
||||
|
||||
@@ -589,7 +624,59 @@ Reads the local export cache and pushes each exported Markdown file to Joplin as
|
||||
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
|
||||
4. Joplin desktop must be open when you run this command
|
||||
|
||||
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
|
||||
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--project NAME`, `--dry-run`
|
||||
|
||||
### `sync` — Export and sync in one run
|
||||
|
||||
```bash
|
||||
# The whole archive run: export, then push to Joplin
|
||||
ai-chat-exporter sync
|
||||
|
||||
# One provider
|
||||
ai-chat-exporter sync --provider codex
|
||||
|
||||
# Export only; don't touch Joplin
|
||||
ai-chat-exporter sync --skip-joplin
|
||||
|
||||
# Joplin being closed is a warning, not a failure (used by the schedulers)
|
||||
ai-chat-exporter sync --joplin-optional
|
||||
```
|
||||
|
||||
Equivalent to `export` followed by `joplin` with the same `--provider`. Intended
|
||||
for scheduled runs — see [Scheduling a Daily Run](#scheduling-a-daily-run).
|
||||
|
||||
Unlike the individual commands, `sync` sets a **meaningful exit code**: non-zero
|
||||
if any conversation failed to export or any note failed to sync. A provider whose
|
||||
listing call fails outright (an expired web session token being the usual cause)
|
||||
counts its whole batch as failed. A provider that is simply unconfigured, or that
|
||||
had nothing new, is ordinary success. `export` on its own always exits 0, which
|
||||
is fine when you're reading the summary table and useless to a scheduler.
|
||||
|
||||
`--joplin-optional` downgrades an unreachable Joplin to a warning: the export has
|
||||
already captured the local transcripts, and the notes are rebuilt from the cache
|
||||
by the next run that finds Joplin open.
|
||||
|
||||
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--since YYYY-MM-DD`, `--hidden-content [full|placeholder|omit]`, `--max-conversations N`, `--skip-joplin`, `--joplin-optional`, `--notify/--no-notify`, `--dry-run`
|
||||
|
||||
Note this is a deliberate subset of `export`'s options — `--format`, `--output`,
|
||||
`--project`, `--download-media` and `--force` are not passed through. Use
|
||||
`export` directly for those. (`--download-media` still applies from `.env`; the
|
||||
flag is only a per-run override.)
|
||||
|
||||
### `notify` — Push-notification settings and test
|
||||
|
||||
```bash
|
||||
# Show the current settings
|
||||
ai-chat-exporter notify
|
||||
|
||||
# Send a test push to confirm the topic works
|
||||
ai-chat-exporter notify --test
|
||||
```
|
||||
|
||||
Shows the resolved ntfy configuration and which machine name will appear in the
|
||||
title. See [Getting notified](#getting-notified) for what a scheduled run sends.
|
||||
|
||||
Options: `--test`
|
||||
|
||||
### `prune` — Delete stale export files
|
||||
|
||||
@@ -607,6 +694,50 @@ Joplin. Refuses to run when the manifest is empty (e.g. right after
|
||||
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
|
||||
command separately verifies that every manifest entry's file exists on disk.
|
||||
|
||||
### `projects` — Discover ChatGPT project IDs
|
||||
|
||||
```bash
|
||||
# List the projects your conversations belong to
|
||||
ai-chat-exporter projects
|
||||
|
||||
# Also inspect conversations whose listing entry doesn't name a project
|
||||
ai-chat-exporter projects --deep
|
||||
|
||||
# Write the discovered IDs straight into .env
|
||||
ai-chat-exporter projects --write
|
||||
```
|
||||
|
||||
`CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is
|
||||
invisible to the listing pass — conversations that live *only* inside that
|
||||
project are never fetched at all. This reports every project your conversations
|
||||
belong to, marks the ones absent from `.env`, and prints a paste-ready line.
|
||||
|
||||
`--deep` fetches each conversation's detail when the listing doesn't name its
|
||||
project: complete, but one request per conversation, so it's slow.
|
||||
|
||||
Options: `--deep`, `--write`
|
||||
|
||||
### `canary` — Check for provider API drift
|
||||
|
||||
```bash
|
||||
ai-chat-exporter canary
|
||||
ai-chat-exporter canary --provider chatgpt
|
||||
```
|
||||
|
||||
The web providers are undocumented internal APIs that can change shape without
|
||||
notice, and the failure mode is silent — a renamed field means content is
|
||||
quietly dropped rather than an error being raised. The canary fetches one
|
||||
listing page and one conversation per provider and asserts only the fields the
|
||||
normalizer actually depends on.
|
||||
|
||||
Findings are `ERROR` (a load-bearing field is missing or mistyped — the parser
|
||||
will break or silently lose data) or `WARN` (something unfamiliar appeared;
|
||||
worth investigating, not necessarily broken). **Exits non-zero on any ERROR**,
|
||||
so it can be scheduled or run in CI. Local providers have no remote schema and
|
||||
are not probed.
|
||||
|
||||
Options: `--provider [chatgpt|claude|all]`
|
||||
|
||||
### `cache` — Manage the sync manifest
|
||||
|
||||
```bash
|
||||
@@ -706,7 +837,10 @@ No new or updated conversations since your last run. To verify: `ai-chat-exporte
|
||||
|
||||
See `FUTURE.md` for the full roadmap. Current priorities:
|
||||
|
||||
- **Watch/scheduled mode** on the way to a headless StartOS service
|
||||
- **A StartOS service** that centralises every machine's conversations into one
|
||||
corpus and owns the Joplin connection, so each machine only has to upload
|
||||
(`FUTURE.md` §8)
|
||||
- **Splitting this README** into a short overview plus separate documents
|
||||
|
||||
---
|
||||
|
||||
|
||||
+2
-2
@@ -4,8 +4,8 @@ build-backend = "setuptools.build_meta"
|
||||
|
||||
[project]
|
||||
name = "ai-chat-exporter"
|
||||
version = "0.8.0"
|
||||
description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin"
|
||||
version = "0.9.0"
|
||||
description = "Archive ChatGPT, Claude, Claude Code and Codex conversation history to Markdown for personal backup in Joplin"
|
||||
requires-python = ">=3.11"
|
||||
dependencies = [
|
||||
"requests==2.31.0",
|
||||
|
||||
Reference in New Issue
Block a user