release: v0.9.0, and rewrite FUTURE.md around what is actually planned
FUTURE.md had become a 648-line archaeological record: six shipped roadmap items, two dropped, two implemented, an investigation trail for each, and a backlog closed as not needed — with the two genuinely planned items buried at the bottom. It is now 139 lines, planned work first. - Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon behind decisions now recorded in one line each — why Brave cookie extraction is not viable, why the ChatGPT token's expiry cannot be read client-side, how the drift canary was designed — which would be expensive to rediscover. - Roadmap is now two items: the StartOS service (including the 2026-08-18 decision that clients upload to StartOS storage and the service owns the Joplin connection), and the README split. - Everything shipped or dropped is off the roadmap. Decisions not to build survive as a one-line table so they are not re-proposed. - Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0. Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased] section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy notifications, gizmo_id project attribution and the `projects` command, and the splitlines data-loss fix in both local providers. Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose — exported conversations are Markdown and may contain private content — with each doc re-included by name, so a new doc is silently untracked rather than rejected. The archive would not have been committed at all. Recorded in the README-split roadmap item, since `docs/*.md` will hit exactly this. Also refreshes the package description, which still named only ChatGPT and Claude after two local providers were added. Carries the documentation and table-of-contents work from earlier today.
This commit is contained in:
@@ -48,6 +48,13 @@ CLAUDE_SESSION_KEY=
|
|||||||
# touched. To never tag specific repos, list their names here (comma-separated).
|
# touched. To never tag specific repos, list their names here (comma-separated).
|
||||||
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
|
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
|
||||||
|
|
||||||
|
# --- Launcher ---
|
||||||
|
# Read by the ai-chat-exporter wrapper scripts, not by the Python code. The
|
||||||
|
# wrapper warns when run from outside the repo, because cache/ and exports/
|
||||||
|
# resolve against the current directory and the wrong one silently starts a
|
||||||
|
# separate archive. Set to 1 to silence that warning.
|
||||||
|
#AI_CHAT_EXPORTER_QUIET_CWD=1
|
||||||
|
|
||||||
# --- Notifications (ntfy) ---
|
# --- Notifications (ntfy) ---
|
||||||
# Push the result of a run to ntfy so an unattended archive reports back — the
|
# Push the result of a run to ntfy so an unattended archive reports back — the
|
||||||
# log file, the systemd journal and Task Scheduler's exit code are all pull-only.
|
# log file, the systemd journal and Task Scheduler's exit code are all pull-only.
|
||||||
|
|||||||
@@ -22,6 +22,7 @@ exports/
|
|||||||
!tests/fixtures/*.json
|
!tests/fixtures/*.json
|
||||||
!README.md
|
!README.md
|
||||||
!FUTURE.md
|
!FUTURE.md
|
||||||
|
!FUTURE-ARCHIVE.md
|
||||||
!CHANGELOG.md
|
!CHANGELOG.md
|
||||||
|
|
||||||
# Cache and logs
|
# Cache and logs
|
||||||
|
|||||||
+1
-1
@@ -3,7 +3,7 @@
|
|||||||
All notable changes to this project will be documented here.
|
All notable changes to this project will be documented here.
|
||||||
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
||||||
|
|
||||||
## [Unreleased]
|
## [0.9.0] - 2026-08-18
|
||||||
|
|
||||||
### Fixed
|
### Fixed
|
||||||
- **An em dash in a notification title silently dropped the notification.** HTTP header values are latin-1 at best and `requests` raises on anything outside it, so the first real send failed with `'latin-1' codec can't encode character '\u2014'`. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
|
- **An em dash in a notification title silently dropped the notification.** HTTP header values are latin-1 at best and `requests` raises on anything outside it, so the first real send failed with `'latin-1' codec can't encode character '\u2014'`. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
|
||||||
|
|||||||
@@ -0,0 +1,660 @@
|
|||||||
|
# FUTURE.md — archived 2026-08-18 (pre-v0.9.0 cleanup)
|
||||||
|
|
||||||
|
This is the full `FUTURE.md` as it stood before the v0.9.0 cleanup, kept
|
||||||
|
because it carries the investigation trails behind decisions that are now
|
||||||
|
recorded in one line each: why Brave cookie extraction is not viable, why the
|
||||||
|
ChatGPT token's expiry cannot be read client-side, how the drift canary was
|
||||||
|
designed, and the reasoning behind each closed backlog item.
|
||||||
|
|
||||||
|
Nothing here is planned work. The live roadmap is in `FUTURE.md`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Planned Future Work
|
||||||
|
|
||||||
|
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
|
||||||
|
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
|
||||||
|
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
|
||||||
|
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
|
||||||
|
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
|
||||||
|
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
|
||||||
|
>
|
||||||
|
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
|
||||||
|
> active roadmap is empty and the remaining backlog below has been **closed as
|
||||||
|
> not needed** — the tool does what it's needed to do as a local, manually-run
|
||||||
|
> backup CLI. Items are kept for reference only; revisit on demand if a real
|
||||||
|
> need shows up. Nothing here is planned work.
|
||||||
|
|
||||||
|
Items completed in each release are moved to the changelog. Items below the
|
||||||
|
roadmap were designed for but intentionally not implemented. The codebase is
|
||||||
|
structured to make each of these additions straightforward if ever revived.
|
||||||
|
|
||||||
|
**Completed:**
|
||||||
|
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
||||||
|
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
||||||
|
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
|
||||||
|
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
||||||
|
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Roadmap (decided 2026-06-12)
|
||||||
|
|
||||||
|
Priorities reflect the tool's primary purpose — a trustworthy backup so that
|
||||||
|
conversation data is not lost if a provider account is ever closed — plus the
|
||||||
|
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
|
||||||
|
local, manually-run CLI; the headless/StartOS direction was dropped
|
||||||
|
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
|
||||||
|
(manual refresh is sufficient).
|
||||||
|
|
||||||
|
**Now (in order):**
|
||||||
|
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
|
||||||
|
(full-archive `export --force` re-export completed 2026-06-13)
|
||||||
|
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
|
||||||
|
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
|
||||||
|
needs admin + SYSTEM impersonation that AV flags as credential theft, and
|
||||||
|
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
|
||||||
|
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
|
||||||
|
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
|
||||||
|
|
||||||
|
**Soon, but later:**
|
||||||
|
|
||||||
|
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
|
||||||
|
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
|
||||||
|
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
|
||||||
|
manually-run CLI, so no in-app polling loop is needed
|
||||||
|
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
|
||||||
|
sufficient (source convos live in the cloud and can be re-downloaded;
|
||||||
|
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
|
||||||
|
also retires the headless token-freshness problem — manual weekly refresh
|
||||||
|
is fine.
|
||||||
|
|
||||||
|
**Active:**
|
||||||
|
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
|
||||||
|
via the `/api/auth/session` `error` field (not `expires`). See §9.
|
||||||
|
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
|
||||||
|
`canary` command. See §10.
|
||||||
|
|
||||||
|
**Deprioritized** (entries kept at the bottom of this file; revisit on
|
||||||
|
demand): `--force` flags, per-conversation cache reset, official export-ZIP
|
||||||
|
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
|
||||||
|
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
|
||||||
|
scope — no significant usage to archive.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
|
||||||
|
|
||||||
|
**Implemented 2026-06-12** as designed below, with one scoping correction
|
||||||
|
from live recon: the retrieval dumps are NOT flagged
|
||||||
|
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
|
||||||
|
the discriminator (the hidden flag only marks Custom Instructions and small
|
||||||
|
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
|
||||||
|
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
|
||||||
|
completed 2026-06-13.
|
||||||
|
|
||||||
|
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
|
||||||
|
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
|
||||||
|
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
|
||||||
|
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
|
||||||
|
the user's own attached documents, the same files dumped dozens of times
|
||||||
|
per conversation. These messages were invisible in the ChatGPT web UI;
|
||||||
|
they appear in exports because v0.4.0 lifted the role filter to fix silent
|
||||||
|
data loss. Custom Instructions hidden-context blocks are a minor secondary
|
||||||
|
case (~2KB, once per conversation) — the originally planned
|
||||||
|
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
|
||||||
|
file contains zero hidden-context blocks.
|
||||||
|
|
||||||
|
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
|
||||||
|
|
||||||
|
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
|
||||||
|
`placeholder`, plus a `--hidden-content` CLI override on `export`.
|
||||||
|
- `placeholder` renders affected messages as one line with type and size:
|
||||||
|
`> 🔧 Tool output (file_search, 24KB) — omitted
|
||||||
|
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
|
||||||
|
roughly halves; worst files shrink ~90%.
|
||||||
|
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
|
||||||
|
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
|
||||||
|
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
|
||||||
|
decision must happen in the provider, not the renderer); (b) messages
|
||||||
|
flagged `is_visually_hidden_from_conversation`, including
|
||||||
|
`user_editable_context` / `model_editable_context` (Custom
|
||||||
|
Instructions) — subsumes the old suppress-hidden-context idea.
|
||||||
|
- **Scope — keep at full size:** code-execution `tool_result` blocks and
|
||||||
|
web-search results; those are usually content the user wants.
|
||||||
|
- Count collapsed messages in the post-export summary so the omission
|
||||||
|
stays visible (mirror the LossReport presentation, but as intentional
|
||||||
|
policy, not loss).
|
||||||
|
|
||||||
|
Re-export workflow after shipping: `cache --clear` + `export` (same as
|
||||||
|
the v0.4.0 migration).
|
||||||
|
|
||||||
|
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
|
||||||
|
|
||||||
|
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
|
||||||
|
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
|
||||||
|
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
|
||||||
|
different mechanism (see below).
|
||||||
|
|
||||||
|
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
|
||||||
|
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
|
||||||
|
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
|
||||||
|
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
|
||||||
|
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
|
||||||
|
symptom hit in the field:
|
||||||
|
|
||||||
|
```
|
||||||
|
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
|
||||||
|
```
|
||||||
|
|
||||||
|
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
|
||||||
|
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
|
||||||
|
rights** and behavior that AV/EDR flags as infostealer activity. The one
|
||||||
|
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
|
||||||
|
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
|
||||||
|
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
|
||||||
|
this, so no off-disk reader is a reliable, non-invasive fit.
|
||||||
|
|
||||||
|
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
|
||||||
|
(launch the browser with `--remote-debugging-port`, read cookies via
|
||||||
|
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
|
||||||
|
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
|
||||||
|
hand. With the headless/StartOS direction dropped (#8), manual DevTools
|
||||||
|
refresh is the accepted approach — no automated extraction is needed.
|
||||||
|
|
||||||
|
## 3. Claude Code Session Provider — SHIPPED v0.6.0
|
||||||
|
|
||||||
|
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
|
||||||
|
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
|
||||||
|
during implementation: `isSidechain` records are subagent transcripts (skipped),
|
||||||
|
`isMeta` marks harness-generated user records (skipped), and listing/normalized
|
||||||
|
`updated_at` must both use file mtime or the cache would re-export every
|
||||||
|
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
|
||||||
|
|
||||||
|
Archive local Claude Code session transcripts. No tokens, no rate limits,
|
||||||
|
no ToS risk — the data is already on disk but lives in a single JSONL per
|
||||||
|
session that Claude Code may clean up, and it contains deliverables
|
||||||
|
(reviews, plans, analyses) that exist nowhere else.
|
||||||
|
|
||||||
|
Decisions (2026-06-12):
|
||||||
|
- **Rendering: prose-only.** Keep user prompts and assistant text
|
||||||
|
(including full deliverable write-ups); collapse tool activity to
|
||||||
|
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
|
||||||
|
86KB — omitted`); **exclude thinking blocks**.
|
||||||
|
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
|
||||||
|
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
|
||||||
|
in v0.5.0).
|
||||||
|
|
||||||
|
Data facts (measured 2026-06-12):
|
||||||
|
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
|
||||||
|
Currently 29 sessions, 18.8MB total, largest 3.8MB.
|
||||||
|
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
|
||||||
|
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
|
||||||
|
activity is what makes these exports readable.
|
||||||
|
- Record types: `user` / `assistant` (Anthropic-style `message.content`
|
||||||
|
block arrays) plus harness records: `ai-title` (use for note title and
|
||||||
|
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
|
||||||
|
`permission-mode`, `system` (skip). Strip harness noise from user
|
||||||
|
messages (`<local-command-caveat>`, `<command-name>` blocks).
|
||||||
|
|
||||||
|
Implementation shape: new `src/providers/claude_code.py` implementing the
|
||||||
|
`BaseProvider` interface — `list_conversations` scans project dirs,
|
||||||
|
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
|
||||||
|
existing block schema (content is already block-shaped: text / tool_use /
|
||||||
|
tool_result / thinking). Incremental sync via file mtime/size recorded in
|
||||||
|
the existing manifest. Project name derives from the munged cwd dirname.
|
||||||
|
|
||||||
|
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
|
||||||
|
|
||||||
|
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
|
||||||
|
(refuses to prune right after `cache --clear`). First live run removed 420
|
||||||
|
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
|
||||||
|
manifest↔disk integrity (293/293 after the run).
|
||||||
|
|
||||||
|
A backup is only trustworthy if the on-disk tree matches the manifest.
|
||||||
|
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
|
||||||
|
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
|
||||||
|
v0.4.1) sit alongside current exports and would double-sync into Joplin.
|
||||||
|
|
||||||
|
- `prune` command: delete export files not referenced by the manifest.
|
||||||
|
`--dry-run` (default off, but always print the list before deleting)
|
||||||
|
shows what would be removed and why (old layout / orphan / unknown).
|
||||||
|
- `doctor` extension: verify every manifest entry's `file_path` exists on
|
||||||
|
disk; report missing files (re-export candidates) and unreferenced files
|
||||||
|
(prune candidates).
|
||||||
|
|
||||||
|
## 5. Binary Content Downloads — SHIPPED v0.6.0
|
||||||
|
|
||||||
|
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
|
||||||
|
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
|
||||||
|
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
|
||||||
|
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
|
||||||
|
`download_url`; a second GET yields the bytes (works for user uploads;
|
||||||
|
older AI-generated images 404 — expired server-side, handled gracefully).
|
||||||
|
Asset refs come in three shapes — `sediment://file_…`,
|
||||||
|
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
|
||||||
|
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
|
||||||
|
images-only is the default and audio is opt-in via `all`.
|
||||||
|
|
||||||
|
Original notes below.
|
||||||
|
|
||||||
|
**Priority note (2026-06-12): "later, but soon" — under the
|
||||||
|
backup-if-account-closes goal, embedded images are part of the data that
|
||||||
|
would be lost; placeholders alone don't preserve them.**
|
||||||
|
|
||||||
|
v0.4.0 ships placeholders for images and audio assets but does not download
|
||||||
|
the binary content. The `_safe_fence`-wrapped placeholders include the asset
|
||||||
|
reference (`sediment://...` or `file-service://...`), MIME type, size, and
|
||||||
|
duration where available; the actual bytes are not preserved.
|
||||||
|
|
||||||
|
Next steps:
|
||||||
|
- Download attached images alongside the Markdown export, save under a
|
||||||
|
`media/` sibling directory with a stable filename derived from the asset
|
||||||
|
reference.
|
||||||
|
- Replace `image_placeholder` rendering with an inline ``
|
||||||
|
reference once the file is on disk.
|
||||||
|
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
|
||||||
|
rewrite the rendered Markdown to use `:/resourceId` references, and track
|
||||||
|
the resource ID in the cache manifest so re-syncs stay idempotent.
|
||||||
|
- DALL-E images on the assistant side: not observed in this user's data; the
|
||||||
|
code path exists (`source = "model_generated"`) but is untested.
|
||||||
|
|
||||||
|
The block-level schema is already in place — only the file-fetch + rewrite
|
||||||
|
layer needs to be added. See the `image_placeholder` and `file_placeholder`
|
||||||
|
block definitions in `src/blocks.py`.
|
||||||
|
|
||||||
|
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
|
||||||
|
|
||||||
|
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
|
||||||
|
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
|
||||||
|
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
|
||||||
|
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
|
||||||
|
the re-run picked it up.
|
||||||
|
|
||||||
|
Cap how many conversations are downloaded in a single `export` run so the
|
||||||
|
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
|
||||||
|
most importantly on the very first export, which otherwise fetches the
|
||||||
|
entire conversation history in one session. Because every run is resumable
|
||||||
|
(the manifest records each conversation immediately), a capped run simply
|
||||||
|
exports the first N pending conversations and the next run picks up where
|
||||||
|
it left off. This keeps traffic looking like a human-paced session rather
|
||||||
|
than a scraper, reducing the risk of rate limiting or account flags.
|
||||||
|
|
||||||
|
Two complementary pieces:
|
||||||
|
|
||||||
|
1. **Session cap** — `--max-conversations N` flag (and
|
||||||
|
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
|
||||||
|
`export` command, slice the pending list after the cache filter:
|
||||||
|
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
|
||||||
|
counts (reuse the message format from the 429 early-exit path) and
|
||||||
|
remind the user to re-run to continue.
|
||||||
|
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
|
||||||
|
random jitter) slept between per-conversation detail fetches in
|
||||||
|
`BaseProvider`, so even a capped run doesn't fire requests
|
||||||
|
back-to-back. The existing 429 backoff in `_request` stays as the
|
||||||
|
reactive safety net.
|
||||||
|
|
||||||
|
Note: the conversation *listing* (paginated, 100/page) still runs in full
|
||||||
|
each time so the cache comparison works — the cap applies to the heavy
|
||||||
|
per-conversation detail fetches, which dominate request volume.
|
||||||
|
|
||||||
|
This is a stepping stone to the StartOS service: a capped, politely-paced
|
||||||
|
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
|
||||||
|
traffic profile a headless deployment needs.
|
||||||
|
|
||||||
|
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
|
||||||
|
|
||||||
|
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
|
||||||
|
whatever hosts the tool: a user cron line locally, and on the long-term
|
||||||
|
StartOS target the platform's own scheduling. Either way the tool only needs
|
||||||
|
to do one capped, politely-paced `export` + `joplin` run and exit — which it
|
||||||
|
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
|
||||||
|
that chains `export` then `joplin` for a single cron line is a trivial
|
||||||
|
add-on, but the polling loop itself is off the roadmap.
|
||||||
|
|
||||||
|
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
|
||||||
|
|
||||||
|
Not pursuing a headless StartOS service. The local, manually-run CLI is
|
||||||
|
sufficient: the source conversations live in the providers' clouds and can be
|
||||||
|
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
|
||||||
|
so durability is covered without a server in the loop.
|
||||||
|
|
||||||
|
Dropping this also retires the one genuinely hard sub-problem it carried —
|
||||||
|
session-token freshness without a browser. There is no headless context to
|
||||||
|
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
|
||||||
|
extraction remains a dead end regardless — see #2.)
|
||||||
|
|
||||||
|
### REOPENED as a TODO (2026-08-18) — centralization, not durability
|
||||||
|
|
||||||
|
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
|
||||||
|
decision rested on "the source conversations live in the providers' clouds and
|
||||||
|
can be re-downloaded". **That is no longer true of half the providers.**
|
||||||
|
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
|
||||||
|
that exist *only* on the machine that produced them — Codex prunes its rollout
|
||||||
|
files, and neither is recoverable from any cloud. The re-download premise now
|
||||||
|
covers the web providers only.
|
||||||
|
|
||||||
|
The new motivation is consolidation rather than durability: work is split across
|
||||||
|
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
|
||||||
|
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
|
||||||
|
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
|
||||||
|
conversations from everywhere, instead of per-machine islands that only meet
|
||||||
|
inside Joplin.
|
||||||
|
|
||||||
|
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
|
||||||
|
running the exporter locally and keeps doing what it is uniquely able to do —
|
||||||
|
read that machine's local transcripts, and hold the browser session for the web
|
||||||
|
providers. What changes is where the output goes: instead of syncing to Joplin
|
||||||
|
itself, a local run **uploads its conversations to the StartOS storage area**,
|
||||||
|
and the StartOS service owns the Joplin connection for the whole corpus.
|
||||||
|
|
||||||
|
That inverts today's arrangement, where every machine talks to its own Joplin
|
||||||
|
desktop, and it removes two problems we already have:
|
||||||
|
|
||||||
|
- **The Joplin-availability race disappears from the clients.** A scheduled run
|
||||||
|
currently has to find Joplin desktop open on that same machine — the
|
||||||
|
2026-08-18 09:02 timer run exported fine and then skipped the sync because
|
||||||
|
Joplin did not start until 09:07. Uploading to a server that is always up has
|
||||||
|
no such window, and `--joplin-optional` stops being load-bearing.
|
||||||
|
- **One Joplin integration instead of N.** Notebook naming, resource upload and
|
||||||
|
note-update logic run once, server-side, against one manifest — rather than
|
||||||
|
each machine independently deciding what a notebook is called and racing to
|
||||||
|
update the same note.
|
||||||
|
|
||||||
|
What this would need, and what it would *not*:
|
||||||
|
|
||||||
|
- **Not** a headless web-provider login. The hard sub-problem the original drop
|
||||||
|
retired stays retired: the web providers can keep running interactively on the
|
||||||
|
machine that has the browser, pushing their output to the server. Only the
|
||||||
|
local providers need to run server-side, and they need no tokens at all.
|
||||||
|
- An upload step in the client — the counterpart of today's `joplin` command,
|
||||||
|
pointed at the StartOS service instead of a local Joplin API. Probably a
|
||||||
|
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
|
||||||
|
- Per-machine identity in the corpus, which the exporter currently does not
|
||||||
|
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
|
||||||
|
per-machine label". Centralizing would make that label load-bearing.
|
||||||
|
- Conflict handling for one conversation seen by two machines, and a decision
|
||||||
|
about whether the server or the client owns the cache manifest. It is
|
||||||
|
per-machine today, and that is what makes "already up to date" mean anything.
|
||||||
|
- A story for what the client keeps locally after a successful upload. Exports
|
||||||
|
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
|
||||||
|
rollouts, so the client should probably keep them rather than move them.
|
||||||
|
|
||||||
|
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
|
||||||
|
where the archive is actually read. This is a "nice eventually", not a gap.
|
||||||
|
|
||||||
|
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
|
||||||
|
|
||||||
|
Shipped: `doctor` now adds a "ChatGPT token active" check via
|
||||||
|
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
|
||||||
|
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
|
||||||
|
decode path was removed, and `_fetch_access_token` now fails fast on a set
|
||||||
|
`error` instead of returning a stale token. Tests in
|
||||||
|
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
|
||||||
|
below for the record.
|
||||||
|
|
||||||
|
Goal: if it's cheap to tell how much longer a token will work, show it on
|
||||||
|
`doctor`. Findings from live recon:
|
||||||
|
|
||||||
|
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
|
||||||
|
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
|
||||||
|
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
|
||||||
|
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
|
||||||
|
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
|
||||||
|
never yields an expiry for the real tokens (falls to the "not decodable"
|
||||||
|
branch).
|
||||||
|
- **`/api/auth/session` exposes an `expires`** (the provider already calls
|
||||||
|
this endpoint in `_fetch_access_token`; the response includes `expires`
|
||||||
|
alongside `accessToken`). Live value observed 2026-06-28:
|
||||||
|
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
|
||||||
|
assumption and the lived weekly-refresh cadence, so it is almost certainly
|
||||||
|
the NextAuth **rolling session window** (re-extended on every call), not
|
||||||
|
the point at which the pasted token actually 401s. Displaying it verbatim
|
||||||
|
would give false confidence.
|
||||||
|
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
|
||||||
|
was actually dead (conversations API → 401), `/api/auth/session` still
|
||||||
|
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
|
||||||
|
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
|
||||||
|
So `expires` is a **rolling session window that rolls forward on every call
|
||||||
|
even for a dead token**; displaying it would actively lie. The same response
|
||||||
|
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
|
||||||
|
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
|
||||||
|
token `error` is absent/null; when the session token is dead NextAuth can't
|
||||||
|
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
|
||||||
|
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
|
||||||
|
health check.
|
||||||
|
- **Design (ready to build):**
|
||||||
|
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
|
||||||
|
falsy and `accessToken` present; on `RefreshAccessTokenError` report
|
||||||
|
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
|
||||||
|
work) and do NOT surface `expires`.
|
||||||
|
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
|
||||||
|
without checking `error`, so it proceeds with a stale token and yields a
|
||||||
|
confusing downstream 401 instead of a clear "refresh your token" message.
|
||||||
|
Check `error` there and fail fast.
|
||||||
|
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
|
||||||
|
|
||||||
|
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
|
||||||
|
|
||||||
|
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
|
||||||
|
ChatGPT and Claude). It fetches one listing page + one conversation per
|
||||||
|
provider and asserts only the normalizer's load-bearing fields, emitting
|
||||||
|
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
|
||||||
|
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
|
||||||
|
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
|
||||||
|
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
|
||||||
|
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
|
||||||
|
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
|
||||||
|
trail below for the record.
|
||||||
|
|
||||||
|
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
|
||||||
|
|
||||||
|
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
|
||||||
|
change shape without notice. The worst failure for a backup tool is *silent*:
|
||||||
|
a response-schema change that makes the exporter skip or mis-parse content
|
||||||
|
without erroring. `doctor` currently checks token validity, reachability, and
|
||||||
|
manifest↔disk integrity — but not "does the provider's response still look
|
||||||
|
like what the parser expects."
|
||||||
|
|
||||||
|
To investigate: a lightweight schema/shape assertion on a known-good sample
|
||||||
|
of each provider's listing + conversation-detail responses (presence and type
|
||||||
|
of the fields the normalizers rely on), surfaced as a `doctor` check or a
|
||||||
|
dedicated canary.
|
||||||
|
|
||||||
|
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
|
||||||
|
|
||||||
|
- **Dependency surface (assert ONLY these — see below for why):**
|
||||||
|
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
|
||||||
|
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
|
||||||
|
`project.name`, `chat_messages[]`.
|
||||||
|
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
|
||||||
|
(list of typed blocks), `created_at`.
|
||||||
|
- **Key finding — full-shape diffing is the wrong design.** Claude's
|
||||||
|
`settings` object is full of volatile internal codenames that churn
|
||||||
|
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
|
||||||
|
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
|
||||||
|
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
|
||||||
|
canary would fire on every UI experiment. The canary MUST target the
|
||||||
|
normalizer's load-bearing fields only, not the whole response. (Aligns with
|
||||||
|
the drop-noise-don't-retain-it principle.)
|
||||||
|
- **Real Claude messages are flat `text`/`sender`** — in this archive every
|
||||||
|
message had a string `text` and NO `content` block list (0 rich blocks
|
||||||
|
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
|
||||||
|
thinking, image, …) is an **unexercised theoretical path**; drift there
|
||||||
|
can't be "caught" by a canary because it never runs on real data — it's a
|
||||||
|
safety net for if Claude ever switches to block content. The canary should
|
||||||
|
assert the flat shape and *warn if `content` ever appears as a list* (that
|
||||||
|
itself is the drift event that would activate the dormant code).
|
||||||
|
- **Possible silent-loss spot (separate from drift):** Claude messages carry
|
||||||
|
`attachments` and `files` arrays (empty in this sample) that the normalizer
|
||||||
|
ignores entirely. If a user ever attaches files in Claude, they'd be
|
||||||
|
dropped without a LossReport entry. Worth a follow-up check.
|
||||||
|
**Live recon — ChatGPT captured 2026-06-28:**
|
||||||
|
|
||||||
|
- **Dependency surface (assert ONLY these):**
|
||||||
|
- listing item: `id`, `title`, `update_time`/`create_time`.
|
||||||
|
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
|
||||||
|
`update_time`, `mapping` (non-empty).
|
||||||
|
- mapping node: `message`, `children` (the tree walk depends on both);
|
||||||
|
message: `author.role`, `author.name`, `content.content_type`,
|
||||||
|
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
|
||||||
|
- **content_type vocabulary observed (all currently handled):** `text`,
|
||||||
|
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
|
||||||
|
`execution_output`, `reasoning_recap`, `user_editable_context`,
|
||||||
|
`tether_browsing_display`. A *new* content_type already degrades gracefully
|
||||||
|
(visible `unknown` block + WARNING + LossReport tally) — so content_type
|
||||||
|
drift is **already non-silent**. The canary just needs to confirm the known
|
||||||
|
set still parses to non-empty blocks.
|
||||||
|
- **The genuinely silent drift risk — `author.name` collapse keys.**
|
||||||
|
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
|
||||||
|
confirms `file_search` is live (and `web.run`/`python` are correctly left
|
||||||
|
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
|
||||||
|
stops** and the archive re-bloats with no error or LossReport entry. This is
|
||||||
|
the top canary target: assert that retrieval-dump tool authors are still
|
||||||
|
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
|
||||||
|
- **Second silent risk — empty `content.parts`.** A `text` message whose
|
||||||
|
`parts` field is renamed/emptied yields zero blocks and is skipped with only
|
||||||
|
a debug log = silent loss. Canary should assert a sampled `text` message
|
||||||
|
produces a non-empty block.
|
||||||
|
|
||||||
|
**Recon complete for both providers. Canary design (ready to build):** a
|
||||||
|
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
|
||||||
|
listing page + one conversation and asserts the dependency-surface fields
|
||||||
|
above by presence+type — NOT full shape (Claude `settings` codenames prove
|
||||||
|
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
|
||||||
|
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
|
||||||
|
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
|
||||||
|
surface as a warning, never a hard error (a backup tool must still run).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Deprioritized — CLOSED as not needed (2026-06-28)
|
||||||
|
|
||||||
|
These were considered and intentionally **not** built. Closed, not planned —
|
||||||
|
the tool is feature-complete for its purpose. Kept for reference in case a
|
||||||
|
real need ever revives one: Joplin `--force`, per-conversation cache reset,
|
||||||
|
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
|
||||||
|
output, token-expiry notifications (also moot — see §9), and a search
|
||||||
|
command. (Also closed, outside this list: handling Claude `attachments`/
|
||||||
|
`files`, which the canary will flag if they ever appear in real data.)
|
||||||
|
|
||||||
|
## Export `--force` Flag — SHIPPED v0.6.0
|
||||||
|
|
||||||
|
Implemented 2026-06-12: `export --force` passes `force=True` to
|
||||||
|
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
|
||||||
|
preserves Joplin links across re-exports, so a forced re-render + `joplin`
|
||||||
|
updates existing notes instead of duplicating them.
|
||||||
|
|
||||||
|
## Joplin `--force` Flag
|
||||||
|
|
||||||
|
Similarly, add `--force` to the `joplin` command to re-sync all cached
|
||||||
|
conversations to Joplin regardless of whether they've been synced before.
|
||||||
|
Useful after making formatting changes to the Markdown exporter.
|
||||||
|
|
||||||
|
Implementation: in `get_joplin_pending()`, return all entries that have a
|
||||||
|
`file_path` when `force=True`, ignoring `joplin_synced_at`.
|
||||||
|
|
||||||
|
## Per-Conversation Cache Reset
|
||||||
|
|
||||||
|
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
|
||||||
|
single conversation without clearing the entire provider cache.
|
||||||
|
|
||||||
|
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
|
||||||
|
delete the entry, then re-run export.
|
||||||
|
|
||||||
|
## Official API Fallback
|
||||||
|
|
||||||
|
If the unofficial internal web API approach breaks, migrate to official export
|
||||||
|
file parsing as a fallback:
|
||||||
|
- ChatGPT: parse `conversations.json` from Settings → Export Data
|
||||||
|
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
|
||||||
|
|
||||||
|
The `BaseProvider` abstract class is intentionally designed so that a
|
||||||
|
`FileProvider` subclass can implement the same interface
|
||||||
|
(`list_conversations`, `get_conversation`, `normalize_conversation`)
|
||||||
|
without any changes to cache, exporters, or CLI code.
|
||||||
|
|
||||||
|
To add this: implement `src/providers/file_chatgpt.py` and
|
||||||
|
`src/providers/file_claude.py`, then add `--input-file` flag to the
|
||||||
|
export command to accept a pre-downloaded export ZIP or JSON.
|
||||||
|
|
||||||
|
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
|
||||||
|
this user needs (project data), so it isn't a real fallback here.
|
||||||
|
|
||||||
|
## Reclassify o1/o3 Reasoning Subparts
|
||||||
|
|
||||||
|
v0.4.0 leaves dict parts inside `text` content_type messages with shape
|
||||||
|
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
|
||||||
|
shape was inferred from a code comment, not captured live). Once a real
|
||||||
|
reasoning conversation is captured, reclassify these as `thinking` blocks.
|
||||||
|
|
||||||
|
## Obsidian Vault Output
|
||||||
|
|
||||||
|
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
|
||||||
|
conversations into an Obsidian vault directory. The current Markdown format
|
||||||
|
is already largely compatible; the main differences are:
|
||||||
|
|
||||||
|
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
|
||||||
|
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
|
||||||
|
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
|
||||||
|
supports both
|
||||||
|
|
||||||
|
Implementation: the existing `MarkdownExporter` output is already valid in
|
||||||
|
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
|
||||||
|
copy files to the vault directory and maintain a flat or nested folder
|
||||||
|
structure matching the user's Obsidian setup. No API needed — just file I/O.
|
||||||
|
|
||||||
|
## Token Expiry Notifications
|
||||||
|
|
||||||
|
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
|
||||||
|
"proactively notify before expiry" idea is blocked by the same finding: a
|
||||||
|
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
|
||||||
|
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
|
||||||
|
window rather than the real refresh cadence). Any heads-up — a `doctor`
|
||||||
|
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
|
||||||
|
resolving the §9 open question of what signal is actually trustworthy.
|
||||||
|
|
||||||
|
## Search Command
|
||||||
|
|
||||||
|
Add a `search` command to full-text search across all exported Markdown files:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m src.main search "kubernetes ingress"
|
||||||
|
python -m src.main search "kubernetes ingress" --provider claude --project devops
|
||||||
|
```
|
||||||
|
|
||||||
|
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
|
||||||
|
conversation title, date, and a snippet. No index needed — Markdown files are
|
||||||
|
small enough to grep directly.
|
||||||
|
|
||||||
|
## Split the README into Separate Documents
|
||||||
|
|
||||||
|
**TODO (2026-08-18).** The README is 830 lines / 5,562 words / 37 KB — about a
|
||||||
|
25-minute read, with 61 headings. An H2-only table of contents was added the
|
||||||
|
same day and helps navigation, but it treats the symptom: the file is doing at
|
||||||
|
least four unrelated jobs at once.
|
||||||
|
|
||||||
|
Rough shape of a split:
|
||||||
|
|
||||||
|
| Document | Content today |
|
||||||
|
|----------|---------------|
|
||||||
|
| `README.md` | What it is, install, first run, a pointer to the rest |
|
||||||
|
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
|
||||||
|
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
|
||||||
|
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
|
||||||
|
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
|
||||||
|
|
||||||
|
Not done yet, and not urgent, because it has a real cost the TOC does not: any
|
||||||
|
existing link into a README section (a bookmark, a note, another repo, a commit
|
||||||
|
message) breaks when that section moves to another file. Worth doing when the
|
||||||
|
README next needs substantial editing anyway, rather than as a change of its own.
|
||||||
|
|
||||||
|
Two things to decide when it happens:
|
||||||
|
|
||||||
|
- Whether `docs/` renders acceptably on the Gitea instance that hosts this repo
|
||||||
|
(relative links between Markdown files do work there, but worth confirming
|
||||||
|
before splitting rather than after).
|
||||||
|
- Whether the anchors in the split files stay stable enough to link *between*
|
||||||
|
documents, or whether cross-references should point at file tops only. GFM
|
||||||
|
anchors are derived from heading text, so they break silently on a reword —
|
||||||
|
the same fragility the TOC already carries.
|
||||||
@@ -1,616 +1,144 @@
|
|||||||
# Planned Future Work
|
# Planned Future Work
|
||||||
|
|
||||||
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
|
> **Status 2026-08-18 (v0.9.0).** The tool archives four providers — two web
|
||||||
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
|
> (`chatgpt`, `claude`) and two local agent-transcript (`claude-code`, `codex`)
|
||||||
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
|
> — on a schedule, on Linux and Windows, reporting results by push notification.
|
||||||
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
|
> Two items below are genuinely planned. Everything else has shipped or been
|
||||||
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
|
> decided against.
|
||||||
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
|
|
||||||
>
|
|
||||||
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
|
|
||||||
> active roadmap is empty and the remaining backlog below has been **closed as
|
|
||||||
> not needed** — the tool does what it's needed to do as a local, manually-run
|
|
||||||
> backup CLI. Items are kept for reference only; revisit on demand if a real
|
|
||||||
> need shows up. Nothing here is planned work.
|
|
||||||
|
|
||||||
Items completed in each release are moved to the changelog. Items below the
|
Completed work moves to the changelog; this file holds only what is *not* built
|
||||||
roadmap were designed for but intentionally not implemented. The codebase is
|
yet. Decisions not to build something are recorded at the bottom in one line
|
||||||
structured to make each of these additions straightforward if ever revived.
|
each, so they are not re-proposed — the full investigation trails behind them
|
||||||
|
are in `FUTURE-ARCHIVE.md`.
|
||||||
**Completed:**
|
|
||||||
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
|
||||||
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
|
||||||
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
|
|
||||||
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
|
||||||
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# Roadmap (decided 2026-06-12)
|
# Roadmap
|
||||||
|
|
||||||
Priorities reflect the tool's primary purpose — a trustworthy backup so that
|
## 1. StartOS Service — one corpus, not per-machine islands
|
||||||
conversation data is not lost if a provider account is ever closed — plus the
|
|
||||||
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
|
|
||||||
local, manually-run CLI; the headless/StartOS direction was dropped
|
|
||||||
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
|
|
||||||
(manual refresh is sufficient).
|
|
||||||
|
|
||||||
**Now (in order):**
|
Each machine currently archives to its own `exports/` and its own Joplin. Work
|
||||||
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
|
is split across boxes — coding sessions (`claude-code`, `codex`) on the Linux
|
||||||
(full-archive `export --force` re-export completed 2026-06-13)
|
machine, web chats (`chatgpt`, `claude`) on the Windows one — so there is no
|
||||||
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
|
single place where all conversations exist together.
|
||||||
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
|
|
||||||
needs admin + SYSTEM impersonation that AV flags as credential theft, and
|
|
||||||
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
|
|
||||||
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
|
|
||||||
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
|
|
||||||
|
|
||||||
**Soon, but later:**
|
This was dropped on 2026-06-28 and **reopened 2026-08-18**, because the
|
||||||
|
reasoning behind the drop has gone stale. It rested on "the source conversations
|
||||||
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
|
live in the providers' clouds and can be re-downloaded". That is no longer true
|
||||||
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
|
of half the providers: `claude-code` and `codex` transcripts exist *only* on the
|
||||||
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
|
machine that produced them, Codex prunes its rollout files, and neither is
|
||||||
manually-run CLI, so no in-app polling loop is needed
|
recoverable from any cloud. The motivation is also different from the one
|
||||||
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
|
weighed then — consolidation, not durability.
|
||||||
sufficient (source convos live in the cloud and can be re-downloaded;
|
|
||||||
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
|
|
||||||
also retires the headless token-freshness problem — manual weekly refresh
|
|
||||||
is fine.
|
|
||||||
|
|
||||||
**Active:**
|
|
||||||
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
|
|
||||||
via the `/api/auth/session` `error` field (not `expires`). See §9.
|
|
||||||
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
|
|
||||||
`canary` command. See §10.
|
|
||||||
|
|
||||||
**Deprioritized** (entries kept at the bottom of this file; revisit on
|
|
||||||
demand): `--force` flags, per-conversation cache reset, official export-ZIP
|
|
||||||
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
|
|
||||||
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
|
|
||||||
scope — no significant usage to archive.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
|
|
||||||
|
|
||||||
**Implemented 2026-06-12** as designed below, with one scoping correction
|
|
||||||
from live recon: the retrieval dumps are NOT flagged
|
|
||||||
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
|
|
||||||
the discriminator (the hidden flag only marks Custom Instructions and small
|
|
||||||
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
|
|
||||||
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
|
|
||||||
completed 2026-06-13.
|
|
||||||
|
|
||||||
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
|
|
||||||
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
|
|
||||||
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
|
|
||||||
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
|
|
||||||
the user's own attached documents, the same files dumped dozens of times
|
|
||||||
per conversation. These messages were invisible in the ChatGPT web UI;
|
|
||||||
they appear in exports because v0.4.0 lifted the role filter to fix silent
|
|
||||||
data loss. Custom Instructions hidden-context blocks are a minor secondary
|
|
||||||
case (~2KB, once per conversation) — the originally planned
|
|
||||||
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
|
|
||||||
file contains zero hidden-context blocks.
|
|
||||||
|
|
||||||
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
|
|
||||||
|
|
||||||
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
|
|
||||||
`placeholder`, plus a `--hidden-content` CLI override on `export`.
|
|
||||||
- `placeholder` renders affected messages as one line with type and size:
|
|
||||||
`> 🔧 Tool output (file_search, 24KB) — omitted
|
|
||||||
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
|
|
||||||
roughly halves; worst files shrink ~90%.
|
|
||||||
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
|
|
||||||
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
|
|
||||||
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
|
|
||||||
decision must happen in the provider, not the renderer); (b) messages
|
|
||||||
flagged `is_visually_hidden_from_conversation`, including
|
|
||||||
`user_editable_context` / `model_editable_context` (Custom
|
|
||||||
Instructions) — subsumes the old suppress-hidden-context idea.
|
|
||||||
- **Scope — keep at full size:** code-execution `tool_result` blocks and
|
|
||||||
web-search results; those are usually content the user wants.
|
|
||||||
- Count collapsed messages in the post-export summary so the omission
|
|
||||||
stays visible (mirror the LossReport presentation, but as intentional
|
|
||||||
policy, not loss).
|
|
||||||
|
|
||||||
Re-export workflow after shipping: `cache --clear` + `export` (same as
|
|
||||||
the v0.4.0 migration).
|
|
||||||
|
|
||||||
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
|
|
||||||
|
|
||||||
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
|
|
||||||
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
|
|
||||||
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
|
|
||||||
different mechanism (see below).
|
|
||||||
|
|
||||||
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
|
|
||||||
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
|
|
||||||
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
|
|
||||||
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
|
|
||||||
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
|
|
||||||
symptom hit in the field:
|
|
||||||
|
|
||||||
```
|
|
||||||
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
|
|
||||||
```
|
|
||||||
|
|
||||||
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
|
|
||||||
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
|
|
||||||
rights** and behavior that AV/EDR flags as infostealer activity. The one
|
|
||||||
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
|
|
||||||
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
|
|
||||||
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
|
|
||||||
this, so no off-disk reader is a reliable, non-invasive fit.
|
|
||||||
|
|
||||||
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
|
|
||||||
(launch the browser with `--remote-debugging-port`, read cookies via
|
|
||||||
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
|
|
||||||
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
|
|
||||||
hand. With the headless/StartOS direction dropped (#8), manual DevTools
|
|
||||||
refresh is the accepted approach — no automated extraction is needed.
|
|
||||||
|
|
||||||
## 3. Claude Code Session Provider — SHIPPED v0.6.0
|
|
||||||
|
|
||||||
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
|
|
||||||
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
|
|
||||||
during implementation: `isSidechain` records are subagent transcripts (skipped),
|
|
||||||
`isMeta` marks harness-generated user records (skipped), and listing/normalized
|
|
||||||
`updated_at` must both use file mtime or the cache would re-export every
|
|
||||||
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
|
|
||||||
|
|
||||||
Archive local Claude Code session transcripts. No tokens, no rate limits,
|
|
||||||
no ToS risk — the data is already on disk but lives in a single JSONL per
|
|
||||||
session that Claude Code may clean up, and it contains deliverables
|
|
||||||
(reviews, plans, analyses) that exist nowhere else.
|
|
||||||
|
|
||||||
Decisions (2026-06-12):
|
|
||||||
- **Rendering: prose-only.** Keep user prompts and assistant text
|
|
||||||
(including full deliverable write-ups); collapse tool activity to
|
|
||||||
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
|
|
||||||
86KB — omitted`); **exclude thinking blocks**.
|
|
||||||
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
|
|
||||||
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
|
|
||||||
in v0.5.0).
|
|
||||||
|
|
||||||
Data facts (measured 2026-06-12):
|
|
||||||
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
|
|
||||||
Currently 29 sessions, 18.8MB total, largest 3.8MB.
|
|
||||||
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
|
|
||||||
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
|
|
||||||
activity is what makes these exports readable.
|
|
||||||
- Record types: `user` / `assistant` (Anthropic-style `message.content`
|
|
||||||
block arrays) plus harness records: `ai-title` (use for note title and
|
|
||||||
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
|
|
||||||
`permission-mode`, `system` (skip). Strip harness noise from user
|
|
||||||
messages (`<local-command-caveat>`, `<command-name>` blocks).
|
|
||||||
|
|
||||||
Implementation shape: new `src/providers/claude_code.py` implementing the
|
|
||||||
`BaseProvider` interface — `list_conversations` scans project dirs,
|
|
||||||
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
|
|
||||||
existing block schema (content is already block-shaped: text / tool_use /
|
|
||||||
tool_result / thinking). Incremental sync via file mtime/size recorded in
|
|
||||||
the existing manifest. Project name derives from the munged cwd dirname.
|
|
||||||
|
|
||||||
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
|
|
||||||
|
|
||||||
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
|
|
||||||
(refuses to prune right after `cache --clear`). First live run removed 420
|
|
||||||
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
|
|
||||||
manifest↔disk integrity (293/293 after the run).
|
|
||||||
|
|
||||||
A backup is only trustworthy if the on-disk tree matches the manifest.
|
|
||||||
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
|
|
||||||
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
|
|
||||||
v0.4.1) sit alongside current exports and would double-sync into Joplin.
|
|
||||||
|
|
||||||
- `prune` command: delete export files not referenced by the manifest.
|
|
||||||
`--dry-run` (default off, but always print the list before deleting)
|
|
||||||
shows what would be removed and why (old layout / orphan / unknown).
|
|
||||||
- `doctor` extension: verify every manifest entry's `file_path` exists on
|
|
||||||
disk; report missing files (re-export candidates) and unreferenced files
|
|
||||||
(prune candidates).
|
|
||||||
|
|
||||||
## 5. Binary Content Downloads — SHIPPED v0.6.0
|
|
||||||
|
|
||||||
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
|
|
||||||
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
|
|
||||||
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
|
|
||||||
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
|
|
||||||
`download_url`; a second GET yields the bytes (works for user uploads;
|
|
||||||
older AI-generated images 404 — expired server-side, handled gracefully).
|
|
||||||
Asset refs come in three shapes — `sediment://file_…`,
|
|
||||||
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
|
|
||||||
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
|
|
||||||
images-only is the default and audio is opt-in via `all`.
|
|
||||||
|
|
||||||
Original notes below.
|
|
||||||
|
|
||||||
**Priority note (2026-06-12): "later, but soon" — under the
|
|
||||||
backup-if-account-closes goal, embedded images are part of the data that
|
|
||||||
would be lost; placeholders alone don't preserve them.**
|
|
||||||
|
|
||||||
v0.4.0 ships placeholders for images and audio assets but does not download
|
|
||||||
the binary content. The `_safe_fence`-wrapped placeholders include the asset
|
|
||||||
reference (`sediment://...` or `file-service://...`), MIME type, size, and
|
|
||||||
duration where available; the actual bytes are not preserved.
|
|
||||||
|
|
||||||
Next steps:
|
|
||||||
- Download attached images alongside the Markdown export, save under a
|
|
||||||
`media/` sibling directory with a stable filename derived from the asset
|
|
||||||
reference.
|
|
||||||
- Replace `image_placeholder` rendering with an inline ``
|
|
||||||
reference once the file is on disk.
|
|
||||||
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
|
|
||||||
rewrite the rendered Markdown to use `:/resourceId` references, and track
|
|
||||||
the resource ID in the cache manifest so re-syncs stay idempotent.
|
|
||||||
- DALL-E images on the assistant side: not observed in this user's data; the
|
|
||||||
code path exists (`source = "model_generated"`) but is untested.
|
|
||||||
|
|
||||||
The block-level schema is already in place — only the file-fetch + rewrite
|
|
||||||
layer needs to be added. See the `image_placeholder` and `file_placeholder`
|
|
||||||
block definitions in `src/blocks.py`.
|
|
||||||
|
|
||||||
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
|
|
||||||
|
|
||||||
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
|
|
||||||
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
|
|
||||||
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
|
|
||||||
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
|
|
||||||
the re-run picked it up.
|
|
||||||
|
|
||||||
Cap how many conversations are downloaded in a single `export` run so the
|
|
||||||
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
|
|
||||||
most importantly on the very first export, which otherwise fetches the
|
|
||||||
entire conversation history in one session. Because every run is resumable
|
|
||||||
(the manifest records each conversation immediately), a capped run simply
|
|
||||||
exports the first N pending conversations and the next run picks up where
|
|
||||||
it left off. This keeps traffic looking like a human-paced session rather
|
|
||||||
than a scraper, reducing the risk of rate limiting or account flags.
|
|
||||||
|
|
||||||
Two complementary pieces:
|
|
||||||
|
|
||||||
1. **Session cap** — `--max-conversations N` flag (and
|
|
||||||
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
|
|
||||||
`export` command, slice the pending list after the cache filter:
|
|
||||||
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
|
|
||||||
counts (reuse the message format from the 429 early-exit path) and
|
|
||||||
remind the user to re-run to continue.
|
|
||||||
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
|
|
||||||
random jitter) slept between per-conversation detail fetches in
|
|
||||||
`BaseProvider`, so even a capped run doesn't fire requests
|
|
||||||
back-to-back. The existing 429 backoff in `_request` stays as the
|
|
||||||
reactive safety net.
|
|
||||||
|
|
||||||
Note: the conversation *listing* (paginated, 100/page) still runs in full
|
|
||||||
each time so the cache comparison works — the cap applies to the heavy
|
|
||||||
per-conversation detail fetches, which dominate request volume.
|
|
||||||
|
|
||||||
This is a stepping stone to the StartOS service: a capped, politely-paced
|
|
||||||
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
|
|
||||||
traffic profile a headless deployment needs.
|
|
||||||
|
|
||||||
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
|
|
||||||
|
|
||||||
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
|
|
||||||
whatever hosts the tool: a user cron line locally, and on the long-term
|
|
||||||
StartOS target the platform's own scheduling. Either way the tool only needs
|
|
||||||
to do one capped, politely-paced `export` + `joplin` run and exit — which it
|
|
||||||
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
|
|
||||||
that chains `export` then `joplin` for a single cron line is a trivial
|
|
||||||
add-on, but the polling loop itself is off the roadmap.
|
|
||||||
|
|
||||||
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
|
|
||||||
|
|
||||||
Not pursuing a headless StartOS service. The local, manually-run CLI is
|
|
||||||
sufficient: the source conversations live in the providers' clouds and can be
|
|
||||||
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
|
|
||||||
so durability is covered without a server in the loop.
|
|
||||||
|
|
||||||
Dropping this also retires the one genuinely hard sub-problem it carried —
|
|
||||||
session-token freshness without a browser. There is no headless context to
|
|
||||||
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
|
|
||||||
extraction remains a dead end regardless — see #2.)
|
|
||||||
|
|
||||||
### REOPENED as a TODO (2026-08-18) — centralization, not durability
|
|
||||||
|
|
||||||
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
|
|
||||||
decision rested on "the source conversations live in the providers' clouds and
|
|
||||||
can be re-downloaded". **That is no longer true of half the providers.**
|
|
||||||
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
|
|
||||||
that exist *only* on the machine that produced them — Codex prunes its rollout
|
|
||||||
files, and neither is recoverable from any cloud. The re-download premise now
|
|
||||||
covers the web providers only.
|
|
||||||
|
|
||||||
The new motivation is consolidation rather than durability: work is split across
|
|
||||||
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
|
|
||||||
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
|
|
||||||
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
|
|
||||||
conversations from everywhere, instead of per-machine islands that only meet
|
|
||||||
inside Joplin.
|
|
||||||
|
|
||||||
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
|
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
|
||||||
running the exporter locally and keeps doing what it is uniquely able to do —
|
running the exporter locally and keeps doing what only it can do: read that
|
||||||
read that machine's local transcripts, and hold the browser session for the web
|
machine's local transcripts, and hold the browser session for the web providers.
|
||||||
providers. What changes is where the output goes: instead of syncing to Joplin
|
What changes is where the output goes. Instead of syncing to Joplin itself, a
|
||||||
itself, a local run **uploads its conversations to the StartOS storage area**,
|
local run **uploads its conversations to the StartOS storage area**, and the
|
||||||
and the StartOS service owns the Joplin connection for the whole corpus.
|
StartOS service owns the Joplin connection for the whole corpus.
|
||||||
|
|
||||||
That inverts today's arrangement, where every machine talks to its own Joplin
|
That inverts today's arrangement and removes two problems we already have:
|
||||||
desktop, and it removes two problems we already have:
|
|
||||||
|
|
||||||
- **The Joplin-availability race disappears from the clients.** A scheduled run
|
- **The Joplin-availability race leaves the clients.** A scheduled run currently
|
||||||
currently has to find Joplin desktop open on that same machine — the
|
has to find Joplin desktop open on that same machine — the 09:02 timer run on
|
||||||
2026-08-18 09:02 timer run exported fine and then skipped the sync because
|
2026-08-18 exported fine and then skipped the sync because Joplin did not
|
||||||
Joplin did not start until 09:07. Uploading to a server that is always up has
|
start until 09:07. A server that is always up has no such window, and
|
||||||
no such window, and `--joplin-optional` stops being load-bearing.
|
`--joplin-optional` stops being load-bearing.
|
||||||
- **One Joplin integration instead of N.** Notebook naming, resource upload and
|
- **One Joplin integration instead of N.** Notebook naming, resource upload and
|
||||||
note-update logic run once, server-side, against one manifest — rather than
|
note updates run once, server-side, against one manifest — rather than each
|
||||||
each machine independently deciding what a notebook is called and racing to
|
machine independently deciding what a notebook is called and racing to update
|
||||||
update the same note.
|
the same note.
|
||||||
|
|
||||||
What this would need, and what it would *not*:
|
What this needs, and what it does *not*:
|
||||||
|
|
||||||
- **Not** a headless web-provider login. The hard sub-problem the original drop
|
- **Not** a headless web-provider login. The hard sub-problem the original drop
|
||||||
retired stays retired: the web providers can keep running interactively on the
|
retired stays retired: the web providers keep running interactively on the
|
||||||
machine that has the browser, pushing their output to the server. Only the
|
machine that has the browser, and push their output to the server. Only the
|
||||||
local providers need to run server-side, and they need no tokens at all.
|
local providers would run server-side, and they need no tokens at all.
|
||||||
- An upload step in the client — the counterpart of today's `joplin` command,
|
- An upload step in the client — the counterpart of today's `joplin` command,
|
||||||
pointed at the StartOS service instead of a local Joplin API. Probably a
|
pointed at the StartOS service instead of a local Joplin API. Probably a
|
||||||
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
|
`push` alongside `sync`, so a scheduled client run stays one line.
|
||||||
- Per-machine identity in the corpus, which the exporter currently does not
|
- Per-machine identity in the corpus, which the exporter does not track today:
|
||||||
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
|
`claude_code.resolve_roots` deliberately merges multiple roots with "no
|
||||||
per-machine label". Centralizing would make that label load-bearing.
|
per-machine label". Centralizing makes that label load-bearing.
|
||||||
- Conflict handling for one conversation seen by two machines, and a decision
|
- Conflict handling for one conversation seen by two machines, and a decision
|
||||||
about whether the server or the client owns the cache manifest. It is
|
about whether the server or the client owns the cache manifest. It is
|
||||||
per-machine today, and that is what makes "already up to date" mean anything.
|
per-machine today, and that is what makes "already up to date" mean anything.
|
||||||
- A story for what the client keeps locally after a successful upload. Exports
|
- A story for what the client keeps locally after a successful upload. Exports
|
||||||
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
|
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
|
||||||
rollouts, so the client should probably keep them rather than move them.
|
rollouts, so the client should keep them rather than hand them off.
|
||||||
|
|
||||||
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
|
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
|
||||||
where the archive is actually read. This is a "nice eventually", not a gap.
|
where the archive is actually read. This is a "nice eventually", not a gap.
|
||||||
|
|
||||||
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
|
## 2. Split the README into separate documents
|
||||||
|
|
||||||
Shipped: `doctor` now adds a "ChatGPT token active" check via
|
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61
|
||||||
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
|
headings. An H2-only table of contents was added 2026-08-18 and helps, but it
|
||||||
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
|
treats the symptom: the file is doing at least four unrelated jobs at once.
|
||||||
decode path was removed, and `_fetch_access_token` now fails fast on a set
|
|
||||||
`error` instead of returning a stale token. Tests in
|
|
||||||
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
|
|
||||||
below for the record.
|
|
||||||
|
|
||||||
Goal: if it's cheap to tell how much longer a token will work, show it on
|
Rough shape of a split:
|
||||||
`doctor`. Findings from live recon:
|
|
||||||
|
|
||||||
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
|
| Document | Content today |
|
||||||
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
|
|----------|---------------|
|
||||||
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
|
| `README.md` | What it is, install, first run, a pointer to the rest |
|
||||||
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
|
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
|
||||||
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
|
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
|
||||||
never yields an expiry for the real tokens (falls to the "not decodable"
|
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
|
||||||
branch).
|
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
|
||||||
- **`/api/auth/session` exposes an `expires`** (the provider already calls
|
|
||||||
this endpoint in `_fetch_access_token`; the response includes `expires`
|
|
||||||
alongside `accessToken`). Live value observed 2026-06-28:
|
|
||||||
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
|
|
||||||
assumption and the lived weekly-refresh cadence, so it is almost certainly
|
|
||||||
the NextAuth **rolling session window** (re-extended on every call), not
|
|
||||||
the point at which the pasted token actually 401s. Displaying it verbatim
|
|
||||||
would give false confidence.
|
|
||||||
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
|
|
||||||
was actually dead (conversations API → 401), `/api/auth/session` still
|
|
||||||
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
|
|
||||||
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
|
|
||||||
So `expires` is a **rolling session window that rolls forward on every call
|
|
||||||
even for a dead token**; displaying it would actively lie. The same response
|
|
||||||
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
|
|
||||||
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
|
|
||||||
token `error` is absent/null; when the session token is dead NextAuth can't
|
|
||||||
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
|
|
||||||
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
|
|
||||||
health check.
|
|
||||||
- **Design (ready to build):**
|
|
||||||
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
|
|
||||||
falsy and `accessToken` present; on `RefreshAccessTokenError` report
|
|
||||||
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
|
|
||||||
work) and do NOT surface `expires`.
|
|
||||||
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
|
|
||||||
without checking `error`, so it proceeds with a stale token and yields a
|
|
||||||
confusing downstream 401 instead of a clear "refresh your token" message.
|
|
||||||
Check `error` there and fail fast.
|
|
||||||
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
|
|
||||||
|
|
||||||
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
|
Not urgent, because it has a real cost the TOC did not: any existing link into a
|
||||||
|
README section — a bookmark, a note, another repo, a commit message — breaks
|
||||||
|
when that section moves to another file. Worth folding into the next substantial
|
||||||
|
README edit rather than doing as a change of its own.
|
||||||
|
|
||||||
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
|
Two things to settle when it happens:
|
||||||
ChatGPT and Claude). It fetches one listing page + one conversation per
|
|
||||||
provider and asserts only the normalizer's load-bearing fields, emitting
|
|
||||||
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
|
|
||||||
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
|
|
||||||
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
|
|
||||||
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
|
|
||||||
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
|
|
||||||
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
|
|
||||||
trail below for the record.
|
|
||||||
|
|
||||||
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
|
- **`.gitignore` ignores `*.md` on purpose** — exported conversations are
|
||||||
|
Markdown and may contain private content — and re-includes each doc by name
|
||||||
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
|
(`!README.md`, `!FUTURE.md`, …). New files under `docs/` will be silently
|
||||||
change shape without notice. The worst failure for a backup tool is *silent*:
|
ignored, with no error, until `!docs/*.md` is added. This already bit the
|
||||||
a response-schema change that makes the exporter skip or mis-parse content
|
creation of `FUTURE-ARCHIVE.md` on 2026-08-18.
|
||||||
without erroring. `doctor` currently checks token validity, reachability, and
|
- Whether `docs/` renders acceptably on the Gitea instance hosting this repo.
|
||||||
manifest↔disk integrity — but not "does the provider's response still look
|
Relative links between Markdown files do work there, but confirm before
|
||||||
like what the parser expects."
|
splitting rather than after.
|
||||||
|
- Whether anchors in the split files are stable enough to link *between*
|
||||||
To investigate: a lightweight schema/shape assertion on a known-good sample
|
documents, or whether cross-references should point at file tops only. GFM
|
||||||
of each provider's listing + conversation-detail responses (presence and type
|
anchors derive from heading text and break silently on a reword — the same
|
||||||
of the fields the normalizers rely on), surfaced as a `doctor` check or a
|
fragility the TOC already carries.
|
||||||
dedicated canary.
|
|
||||||
|
|
||||||
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
|
|
||||||
|
|
||||||
- **Dependency surface (assert ONLY these — see below for why):**
|
|
||||||
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
|
|
||||||
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
|
|
||||||
`project.name`, `chat_messages[]`.
|
|
||||||
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
|
|
||||||
(list of typed blocks), `created_at`.
|
|
||||||
- **Key finding — full-shape diffing is the wrong design.** Claude's
|
|
||||||
`settings` object is full of volatile internal codenames that churn
|
|
||||||
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
|
|
||||||
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
|
|
||||||
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
|
|
||||||
canary would fire on every UI experiment. The canary MUST target the
|
|
||||||
normalizer's load-bearing fields only, not the whole response. (Aligns with
|
|
||||||
the drop-noise-don't-retain-it principle.)
|
|
||||||
- **Real Claude messages are flat `text`/`sender`** — in this archive every
|
|
||||||
message had a string `text` and NO `content` block list (0 rich blocks
|
|
||||||
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
|
|
||||||
thinking, image, …) is an **unexercised theoretical path**; drift there
|
|
||||||
can't be "caught" by a canary because it never runs on real data — it's a
|
|
||||||
safety net for if Claude ever switches to block content. The canary should
|
|
||||||
assert the flat shape and *warn if `content` ever appears as a list* (that
|
|
||||||
itself is the drift event that would activate the dormant code).
|
|
||||||
- **Possible silent-loss spot (separate from drift):** Claude messages carry
|
|
||||||
`attachments` and `files` arrays (empty in this sample) that the normalizer
|
|
||||||
ignores entirely. If a user ever attaches files in Claude, they'd be
|
|
||||||
dropped without a LossReport entry. Worth a follow-up check.
|
|
||||||
**Live recon — ChatGPT captured 2026-06-28:**
|
|
||||||
|
|
||||||
- **Dependency surface (assert ONLY these):**
|
|
||||||
- listing item: `id`, `title`, `update_time`/`create_time`.
|
|
||||||
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
|
|
||||||
`update_time`, `mapping` (non-empty).
|
|
||||||
- mapping node: `message`, `children` (the tree walk depends on both);
|
|
||||||
message: `author.role`, `author.name`, `content.content_type`,
|
|
||||||
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
|
|
||||||
- **content_type vocabulary observed (all currently handled):** `text`,
|
|
||||||
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
|
|
||||||
`execution_output`, `reasoning_recap`, `user_editable_context`,
|
|
||||||
`tether_browsing_display`. A *new* content_type already degrades gracefully
|
|
||||||
(visible `unknown` block + WARNING + LossReport tally) — so content_type
|
|
||||||
drift is **already non-silent**. The canary just needs to confirm the known
|
|
||||||
set still parses to non-empty blocks.
|
|
||||||
- **The genuinely silent drift risk — `author.name` collapse keys.**
|
|
||||||
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
|
|
||||||
confirms `file_search` is live (and `web.run`/`python` are correctly left
|
|
||||||
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
|
|
||||||
stops** and the archive re-bloats with no error or LossReport entry. This is
|
|
||||||
the top canary target: assert that retrieval-dump tool authors are still
|
|
||||||
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
|
|
||||||
- **Second silent risk — empty `content.parts`.** A `text` message whose
|
|
||||||
`parts` field is renamed/emptied yields zero blocks and is skipped with only
|
|
||||||
a debug log = silent loss. Canary should assert a sampled `text` message
|
|
||||||
produces a non-empty block.
|
|
||||||
|
|
||||||
**Recon complete for both providers. Canary design (ready to build):** a
|
|
||||||
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
|
|
||||||
listing page + one conversation and asserts the dependency-surface fields
|
|
||||||
above by presence+type — NOT full shape (Claude `settings` codenames prove
|
|
||||||
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
|
|
||||||
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
|
|
||||||
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
|
|
||||||
surface as a warning, never a hard error (a backup tool must still run).
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# Deprioritized — CLOSED as not needed (2026-06-28)
|
# Decided against
|
||||||
|
|
||||||
These were considered and intentionally **not** built. Closed, not planned —
|
Recorded so they are not re-proposed. Full reasoning and the recon behind each
|
||||||
the tool is feature-complete for its purpose. Kept for reference in case a
|
is in `FUTURE-ARCHIVE.md`.
|
||||||
real need ever revives one: Joplin `--force`, per-conversation cache reset,
|
|
||||||
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
|
|
||||||
output, token-expiry notifications (also moot — see §9), and a search
|
|
||||||
command. (Also closed, outside this list: handling Claude `attachments`/
|
|
||||||
`files`, which the canary will flag if they ever appear in real data.)
|
|
||||||
|
|
||||||
## Export `--force` Flag — SHIPPED v0.6.0
|
| Item | Verdict |
|
||||||
|
|------|---------|
|
||||||
|
| Brave/Chromium cookie auto-extraction | **Not viable** (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
|
||||||
|
| In-app watch/polling loop | **Dropped** (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as `sync` plus a systemd timer and a Windows scheduled task. |
|
||||||
|
| Proactive token-expiry notification | **Blocked, and moot** (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and `/api/auth/session`'s `expires` is a rolling window that advances even for a dead token. `doctor`'s `error`-based health check covers the real need. |
|
||||||
|
| Official export-ZIP fallback | **Closed** (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. `BaseProvider` still admits a `FileProvider` if that changes. |
|
||||||
|
| Joplin `--force` flag | Closed as not needed (2026-06-28). |
|
||||||
|
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
|
||||||
|
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
|
||||||
|
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
|
||||||
|
| Search command | Closed as not needed (2026-06-28). `grep`/`ripgrep` over `EXPORT_DIR` covers it. |
|
||||||
|
| Claude `attachments` / `files` handling | Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
|
||||||
|
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
|
||||||
|
|
||||||
Implemented 2026-06-12: `export --force` passes `force=True` to
|
---
|
||||||
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
|
|
||||||
preserves Joplin links across re-exports, so a forced re-render + `joplin`
|
|
||||||
updates existing notes instead of duplicating them.
|
|
||||||
|
|
||||||
## Joplin `--force` Flag
|
# Completed
|
||||||
|
|
||||||
Similarly, add `--force` to the `joplin` command to re-sync all cached
|
Detail for each release is in `CHANGELOG.md`.
|
||||||
conversations to Joplin regardless of whether they've been synced before.
|
|
||||||
Useful after making formatting changes to the Markdown exporter.
|
|
||||||
|
|
||||||
Implementation: in `get_joplin_pending()`, return all entries that have a
|
- **v0.1.0** — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
|
||||||
`file_path` when `force=True`, ignoring `joplin_synced_at`.
|
- **v0.2.0** — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
|
||||||
|
- **v0.4.0** — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via `LossReport` and visible `unknown` blocks
|
||||||
## Per-Conversation Cache Reset
|
- **v0.5.0** — Nested Joplin notebooks, date-prefixed note titles, flat year folders
|
||||||
|
- **v0.6.0** — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT`); Claude Code session provider; `prune` + manifest integrity; binary content downloads; session limiter and request pacing; `export --force`
|
||||||
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
|
- **v0.7.0** — `canary` drift detection; real ChatGPT token-health check on `doctor`; removal of the non-viable browser-cookie path
|
||||||
single conversation without clearing the entire provider cache.
|
- **v0.8.0** — Claude Code coverage reopened: subagent capture as folded `<details>`, repo `[tags]` in titles, its own `AI-ClaudeCode` notebook with self-healing note moves, multi-root scanning (`CLAUDE_CODE_DIR`, `CLAUDE_CONFIG_DIR`)
|
||||||
|
- **v0.9.0** — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms; `sync` with a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution via `gizmo_id` and the `projects` command; a silent data-loss fix in both local providers (`splitlines` breaking on U+0085/U+2028/U+2029)
|
||||||
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
|
|
||||||
delete the entry, then re-run export.
|
|
||||||
|
|
||||||
## Official API Fallback
|
|
||||||
|
|
||||||
If the unofficial internal web API approach breaks, migrate to official export
|
|
||||||
file parsing as a fallback:
|
|
||||||
- ChatGPT: parse `conversations.json` from Settings → Export Data
|
|
||||||
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
|
|
||||||
|
|
||||||
The `BaseProvider` abstract class is intentionally designed so that a
|
|
||||||
`FileProvider` subclass can implement the same interface
|
|
||||||
(`list_conversations`, `get_conversation`, `normalize_conversation`)
|
|
||||||
without any changes to cache, exporters, or CLI code.
|
|
||||||
|
|
||||||
To add this: implement `src/providers/file_chatgpt.py` and
|
|
||||||
`src/providers/file_claude.py`, then add `--input-file` flag to the
|
|
||||||
export command to accept a pre-downloaded export ZIP or JSON.
|
|
||||||
|
|
||||||
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
|
|
||||||
this user needs (project data), so it isn't a real fallback here.
|
|
||||||
|
|
||||||
## Reclassify o1/o3 Reasoning Subparts
|
|
||||||
|
|
||||||
v0.4.0 leaves dict parts inside `text` content_type messages with shape
|
|
||||||
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
|
|
||||||
shape was inferred from a code comment, not captured live). Once a real
|
|
||||||
reasoning conversation is captured, reclassify these as `thinking` blocks.
|
|
||||||
|
|
||||||
## Obsidian Vault Output
|
|
||||||
|
|
||||||
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
|
|
||||||
conversations into an Obsidian vault directory. The current Markdown format
|
|
||||||
is already largely compatible; the main differences are:
|
|
||||||
|
|
||||||
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
|
|
||||||
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
|
|
||||||
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
|
|
||||||
supports both
|
|
||||||
|
|
||||||
Implementation: the existing `MarkdownExporter` output is already valid in
|
|
||||||
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
|
|
||||||
copy files to the vault directory and maintain a flat or nested folder
|
|
||||||
structure matching the user's Obsidian setup. No API needed — just file I/O.
|
|
||||||
|
|
||||||
## Token Expiry Notifications
|
|
||||||
|
|
||||||
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
|
|
||||||
"proactively notify before expiry" idea is blocked by the same finding: a
|
|
||||||
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
|
|
||||||
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
|
|
||||||
window rather than the real refresh cadence). Any heads-up — a `doctor`
|
|
||||||
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
|
|
||||||
resolving the §9 open question of what signal is actually trustworthy.
|
|
||||||
|
|
||||||
## Search Command
|
|
||||||
|
|
||||||
Add a `search` command to full-text search across all exported Markdown files:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
python -m src.main search "kubernetes ingress"
|
|
||||||
python -m src.main search "kubernetes ingress" --provider claude --project devops
|
|
||||||
```
|
|
||||||
|
|
||||||
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
|
|
||||||
conversation title, date, and a snippet. No index needed — Markdown files are
|
|
||||||
small enough to grep directly.
|
|
||||||
|
|||||||
@@ -6,7 +6,28 @@ Supports incremental sync — only new or updated conversations are exported on
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## ⚠️ Terms of Service Warning
|
## Contents
|
||||||
|
|
||||||
|
- [Terms of Service Warning](#terms-of-service-warning)
|
||||||
|
- [Installation](#installation)
|
||||||
|
- [First Run: Run Doctor](#first-run-run-doctor)
|
||||||
|
- [Getting Your Session Tokens](#getting-your-session-tokens)
|
||||||
|
- [The `auth` Command](#the-auth-command)
|
||||||
|
- [`.env` Setup](#env-setup)
|
||||||
|
- [ChatGPT Projects](#chatgpt-projects)
|
||||||
|
- [Claude Code Sessions](#claude-code-sessions)
|
||||||
|
- [Codex Sessions](#codex-sessions)
|
||||||
|
- [Scheduling a Daily Run](#scheduling-a-daily-run)
|
||||||
|
- [Output Structure](#output-structure)
|
||||||
|
- [CLI Reference](#cli-reference)
|
||||||
|
- [How the Cache Works](#how-the-cache-works)
|
||||||
|
- [Troubleshooting](#troubleshooting)
|
||||||
|
- [Future Work](#future-work)
|
||||||
|
- [Security Notes](#security-notes)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Terms of Service Warning
|
||||||
|
|
||||||
**Read this before using this tool.**
|
**Read this before using this tool.**
|
||||||
|
|
||||||
@@ -223,12 +244,26 @@ cp .env.example .env
|
|||||||
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
|
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
|
||||||
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
|
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
|
||||||
|
|
||||||
|
### Notifications
|
||||||
|
|
||||||
|
| Variable | Default | Description |
|
||||||
|
|----------|---------|-------------|
|
||||||
|
| `NTFY_TOPIC` | — | [ntfy](https://ntfy.sh) topic to push run results to. Unset disables notifications entirely. |
|
||||||
|
| `NTFY_SERVER` | `https://ntfy.sh` | Point at your own host if self-hosting. |
|
||||||
|
| `NTFY_TOKEN` | — | Bearer token, for access-controlled topics. |
|
||||||
|
| `NTFY_NOTIFY` | `always` | `always` notifies on every run, `failure` only when something failed, `off` never. |
|
||||||
|
|
||||||
|
A topic on public ntfy.sh is readable by anyone who knows its name, so
|
||||||
|
notifications carry per-provider counts and a machine name only — never
|
||||||
|
conversation titles. See [Getting notified](#getting-notified).
|
||||||
|
|
||||||
### Cache & logging
|
### Cache & logging
|
||||||
|
|
||||||
| Variable | Default | Description |
|
| Variable | Default | Description |
|
||||||
|----------|---------|-------------|
|
|----------|---------|-------------|
|
||||||
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
|
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
|
||||||
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
|
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
|
||||||
|
| `AI_CHAT_EXPORTER_QUIET_CWD` | — | Set to `1` to silence the launcher's warning when run from outside the repo. Read by the `ai-chat-exporter` wrapper scripts, not by Python; the scheduler installers set it, since they always set the correct working directory. |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -589,7 +624,59 @@ Reads the local export cache and pushes each exported Markdown file to Joplin as
|
|||||||
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
|
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
|
||||||
4. Joplin desktop must be open when you run this command
|
4. Joplin desktop must be open when you run this command
|
||||||
|
|
||||||
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
|
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--project NAME`, `--dry-run`
|
||||||
|
|
||||||
|
### `sync` — Export and sync in one run
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# The whole archive run: export, then push to Joplin
|
||||||
|
ai-chat-exporter sync
|
||||||
|
|
||||||
|
# One provider
|
||||||
|
ai-chat-exporter sync --provider codex
|
||||||
|
|
||||||
|
# Export only; don't touch Joplin
|
||||||
|
ai-chat-exporter sync --skip-joplin
|
||||||
|
|
||||||
|
# Joplin being closed is a warning, not a failure (used by the schedulers)
|
||||||
|
ai-chat-exporter sync --joplin-optional
|
||||||
|
```
|
||||||
|
|
||||||
|
Equivalent to `export` followed by `joplin` with the same `--provider`. Intended
|
||||||
|
for scheduled runs — see [Scheduling a Daily Run](#scheduling-a-daily-run).
|
||||||
|
|
||||||
|
Unlike the individual commands, `sync` sets a **meaningful exit code**: non-zero
|
||||||
|
if any conversation failed to export or any note failed to sync. A provider whose
|
||||||
|
listing call fails outright (an expired web session token being the usual cause)
|
||||||
|
counts its whole batch as failed. A provider that is simply unconfigured, or that
|
||||||
|
had nothing new, is ordinary success. `export` on its own always exits 0, which
|
||||||
|
is fine when you're reading the summary table and useless to a scheduler.
|
||||||
|
|
||||||
|
`--joplin-optional` downgrades an unreachable Joplin to a warning: the export has
|
||||||
|
already captured the local transcripts, and the notes are rebuilt from the cache
|
||||||
|
by the next run that finds Joplin open.
|
||||||
|
|
||||||
|
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--since YYYY-MM-DD`, `--hidden-content [full|placeholder|omit]`, `--max-conversations N`, `--skip-joplin`, `--joplin-optional`, `--notify/--no-notify`, `--dry-run`
|
||||||
|
|
||||||
|
Note this is a deliberate subset of `export`'s options — `--format`, `--output`,
|
||||||
|
`--project`, `--download-media` and `--force` are not passed through. Use
|
||||||
|
`export` directly for those. (`--download-media` still applies from `.env`; the
|
||||||
|
flag is only a per-run override.)
|
||||||
|
|
||||||
|
### `notify` — Push-notification settings and test
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Show the current settings
|
||||||
|
ai-chat-exporter notify
|
||||||
|
|
||||||
|
# Send a test push to confirm the topic works
|
||||||
|
ai-chat-exporter notify --test
|
||||||
|
```
|
||||||
|
|
||||||
|
Shows the resolved ntfy configuration and which machine name will appear in the
|
||||||
|
title. See [Getting notified](#getting-notified) for what a scheduled run sends.
|
||||||
|
|
||||||
|
Options: `--test`
|
||||||
|
|
||||||
### `prune` — Delete stale export files
|
### `prune` — Delete stale export files
|
||||||
|
|
||||||
@@ -607,6 +694,50 @@ Joplin. Refuses to run when the manifest is empty (e.g. right after
|
|||||||
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
|
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
|
||||||
command separately verifies that every manifest entry's file exists on disk.
|
command separately verifies that every manifest entry's file exists on disk.
|
||||||
|
|
||||||
|
### `projects` — Discover ChatGPT project IDs
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# List the projects your conversations belong to
|
||||||
|
ai-chat-exporter projects
|
||||||
|
|
||||||
|
# Also inspect conversations whose listing entry doesn't name a project
|
||||||
|
ai-chat-exporter projects --deep
|
||||||
|
|
||||||
|
# Write the discovered IDs straight into .env
|
||||||
|
ai-chat-exporter projects --write
|
||||||
|
```
|
||||||
|
|
||||||
|
`CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is
|
||||||
|
invisible to the listing pass — conversations that live *only* inside that
|
||||||
|
project are never fetched at all. This reports every project your conversations
|
||||||
|
belong to, marks the ones absent from `.env`, and prints a paste-ready line.
|
||||||
|
|
||||||
|
`--deep` fetches each conversation's detail when the listing doesn't name its
|
||||||
|
project: complete, but one request per conversation, so it's slow.
|
||||||
|
|
||||||
|
Options: `--deep`, `--write`
|
||||||
|
|
||||||
|
### `canary` — Check for provider API drift
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ai-chat-exporter canary
|
||||||
|
ai-chat-exporter canary --provider chatgpt
|
||||||
|
```
|
||||||
|
|
||||||
|
The web providers are undocumented internal APIs that can change shape without
|
||||||
|
notice, and the failure mode is silent — a renamed field means content is
|
||||||
|
quietly dropped rather than an error being raised. The canary fetches one
|
||||||
|
listing page and one conversation per provider and asserts only the fields the
|
||||||
|
normalizer actually depends on.
|
||||||
|
|
||||||
|
Findings are `ERROR` (a load-bearing field is missing or mistyped — the parser
|
||||||
|
will break or silently lose data) or `WARN` (something unfamiliar appeared;
|
||||||
|
worth investigating, not necessarily broken). **Exits non-zero on any ERROR**,
|
||||||
|
so it can be scheduled or run in CI. Local providers have no remote schema and
|
||||||
|
are not probed.
|
||||||
|
|
||||||
|
Options: `--provider [chatgpt|claude|all]`
|
||||||
|
|
||||||
### `cache` — Manage the sync manifest
|
### `cache` — Manage the sync manifest
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -706,7 +837,10 @@ No new or updated conversations since your last run. To verify: `ai-chat-exporte
|
|||||||
|
|
||||||
See `FUTURE.md` for the full roadmap. Current priorities:
|
See `FUTURE.md` for the full roadmap. Current priorities:
|
||||||
|
|
||||||
- **Watch/scheduled mode** on the way to a headless StartOS service
|
- **A StartOS service** that centralises every machine's conversations into one
|
||||||
|
corpus and owns the Joplin connection, so each machine only has to upload
|
||||||
|
(`FUTURE.md` §8)
|
||||||
|
- **Splitting this README** into a short overview plus separate documents
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
+2
-2
@@ -4,8 +4,8 @@ build-backend = "setuptools.build_meta"
|
|||||||
|
|
||||||
[project]
|
[project]
|
||||||
name = "ai-chat-exporter"
|
name = "ai-chat-exporter"
|
||||||
version = "0.8.0"
|
version = "0.9.0"
|
||||||
description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin"
|
description = "Archive ChatGPT, Claude, Claude Code and Codex conversation history to Markdown for personal backup in Joplin"
|
||||||
requires-python = ">=3.11"
|
requires-python = ">=3.11"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"requests==2.31.0",
|
"requests==2.31.0",
|
||||||
|
|||||||
Reference in New Issue
Block a user