The export reads ChatGPT's and Claude's undocumented internal web APIs, which can change shape without notice; the worst failure for a backup tool is a silent one (skipped/mis-parsed content with no error). Add a `canary` command + BaseProvider.check_drift() (overridden by ChatGPT/Claude) that fetches one listing page + one conversation per provider and asserts only the normalizer's load-bearing fields — not the full response shape, which churns harmlessly. The top silent risk it guards is a renamed retrieval-tool author bypassing the hidden-content collapse. ERROR findings exit non-zero; WARN findings are surfaced but non-fatal so a backup run is never blocked. Also retires the in-app watch mode and headless StartOS direction from the roadmap (tool stays a local, manually-run CLI) and updates docs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
11 KiB
11 KiB
Changelog
All notable changes to this project will be documented here. Format follows Keep a Changelog.
[0.7.0] - 2026-06-28
Added
canarycommand — probes the ChatGPT/Claude web APIs for schema drift against the fields the parser actually depends on (one listing page + one conversation per provider), not the full response shape. Flags the silent-failure risks for a backup tool: a renamed retrieval-tool author bypassing the hidden-content collapse, a newcontent_type, drifted message fields, or non-empty Claudeattachments/filesthe normalizer ignores. ERROR findings exit non-zero; WARN findings are surfaced but non-fatal. SeeFUTURE.md§10.doctornow reports a real ChatGPT token-health check ("ChatGPT token active") based on the/api/auth/sessionerrorfield. ChatGPT session tokens are JWEs whoseexpis encrypted and unreadable client-side; the previous decode path could never yield an expiry. The honest signal iserror == "RefreshAccessTokenError"(verified live), which means the session token is dead even thoughexpires/accessTokenstill echo stale values. SeeFUTURE.md§9.
Fixed
ChatGPTProvider._fetch_access_tokennow checks the/api/auth/sessionerrorfield and fails fast with a clear "refresh your token" message. Previously it returned the staleaccessTokenpresent on a dead session, producing a confusing downstream 401 instead of an actionable auth error.
Removed
auth --from-browserand thebrowser-cookie3dependency (src/browser_tokens.py). Browser cookie auto-extraction is not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level layer that can't be decrypted off disk without Administrator rights and AV-flagged SYSTEM impersonation, and fails on Brave specifically (symptom:Unable to get key for cookie decryption). Auth is manual (DevTools wizard) only. SeeFUTURE.md§2 for the full rationale.
[0.6.0] - 2026-06-12
First tagged release. The project was developed through several internal milestones (0.1.0–0.5.0) that were never published; their changes are consolidated here.
Added
Core export & sync
- ChatGPT and Claude export via their internal web APIs, with a local cache/manifest for incremental, resumable sync (only new or updated conversations are fetched; each is written to the manifest immediately).
- Markdown and JSON exporters. Markdown rendering happens at exporter-write time (providers produce typed blocks; exporters render them), which keeps the door open for future Obsidian/HTML output.
- CLI:
export,list,cache,doctor,auth,joplin,prune. --projectfilter onexport/list/joplin(case-insensitive substring, ornone).- ChatGPT Projects support via
CHATGPT_PROJECT_IDS(project conversations are fetched separately from the default listing).
Joplin integration
joplincommand syncs exported Markdown to Joplin as notes; notebooks are created automatically and nested per provider+project. Re-running is safe — the Joplin note ID is stored in the manifest, so notes are updated, not duplicated.JOPLIN_API_TOKEN,JOPLIN_API_URL,JOPLIN_REQUEST_TIMEOUTconfig, with actionable error messages on timeout/connection failure.
Rich content (typed blocks)
- Messages carry an ordered
blockslist: text, code, thinking, tool_use, tool_result, citation, image_placeholder, file_placeholder, unknown. - ChatGPT voice mode:
audio_transcriptionparts render as text; audio asset pointers render as📎 File attachedplaceholders with size/duration. Custom Instructions (user_editable_context/model_editable_context) now appear (were silently dropped) with a> ℹ️ Hidden contextmarker. - ChatGPT
execution_output,system_error, andtether_browsing_displayrender astool_resultblocks (withtool_name,is_error, optionalsummary); transient browse spinners skip silently. - Defensive Claude block extraction (text/thinking/tool_use/tool_result/image, including nested-block flattening).
_safe_fencepicks a backtick fence longer than any run in the content, so embedded triple-backticks can't corrupt rendering (verified live in Joplin).- Data-loss visibility: a
LossReportsummary at the end of everyexportrun breaks downunknown blocksandextraction failuresby raw type;unknownblocks render as visible> ⚠️ Unsupported contentwith the raw type and observed keys.
Invisible-content collapse (EXPORTER_HIDDEN_CONTENT=full|placeholder|omit, default placeholder; export --hidden-content)
- Messages that were invisible in the ChatGPT web UI are collapsed to one-line placeholders instead of exported in full. Two triggers: (a) tool-role retrieval dumps identified by
author.name(file_search,myfiles_browser) — ChatGPT re-injects the full text of attached/project files on every tool run, measured at 86% of all content bytes across the three largest conversations and not hidden-flagged; (b) messages flaggedis_visually_hidden_from_conversation(Custom Instructions). Code-execution results, web-search output, and all dialogue are untouched. - New
collapsedblock renders as> 🔧 **Tool output** — file_search (15.0 KB) — omitted (EXPORTER_HIDDEN_CONTENT=full to keep)(ℹ️ Hidden context variant for hidden-flagged messages). TheLossReportgains acollapsed by policysection (count per origin + total KB) so the omission stays visible.
Claude Code session provider (--provider claude-code)
- Archives local agent transcripts from
~/.claude/projects/*/*.jsonl(override withCLAUDE_CODE_DIR) — no tokens, no rate limits, no ToS exposure. Prose-only by default per the hidden-content policy: dialogue and deliverables kept; tool traffic grouped into one placeholder per activity run (> 🔧 Tool output — 14 calls: Read ×9, Bash ×2 (86KB) — omitted); thinking dropped (counted in the summary); subagent (isSidechain) and harness (isMeta, snapshots, command tags) records stripped. Titles from the lastai-titlerecord; project from the working-directory basename. Joplin notebooks nest under theAI-Claudeparent. Measured: 19MB of JSONL → 804KB of Markdown across 25 sessions.
Binary/image downloads (EXPORTER_DOWNLOAD_MEDIA=images|all|off, default images; export --download-media)
- Downloads ChatGPT conversation assets into a
media/folder beside each export and inlines them — images become realembeds, files become local links. Two-hop fetch via/backend-api/files/{id}/download→ signed URL (verified live).allalso pulls voice-mode audio (≈162MB of clips in this archive, transcripts already in text — hence the images-only default). Idempotent (skips assets already on disk), atomic writes,600perms. Expired/missing assets (old generated images 404) keep their placeholder and are counted asmedia failed— never fatal. - Joplin resource upload: the
joplincommand uploads downloaded media as resources and rewritesmedia/…links to:/resourceId, so images render inside notes (verified live: resources created with correct mime/size and linked to the note). Resource IDs are tracked per-conversation in the manifest, so re-syncs reuse them instead of duplicating.
Token setup & throughput
auth --from-browser [brave|chrome|chromium|edge|firefox]: extracts ChatGPT session-token cookies and the ClaudesessionKeystraight from a local browser's cookie store (viabrowser-cookie3; defaults to Brave), validates each against the live API, and writes.env. Tokens that fail validation never overwrite a working entry; clear fallback when the browser is absent or the cookie DB is locked.- Session limiter:
--max-conversations N/MAX_CONVERSATIONS_PER_RUNcap downloads per run (per provider); capped-out conversations are reported as "deferred" and resumed next run. - Polite pacing:
REQUEST_DELAY(default 1.0s, ±25% jitter,0disables) between consecutive API requests, with the existing 429 backoff as the reactive net.
Re-rendering
export --forcere-exports every conversation even if cached and unchanged, so the whole archive can be re-rendered after a formatting/feature change withoutcache --clear. Runs as a tracked campaign (a stamped start time in the manifest): each run re-renders the least-recently-exported conversations, the "still to go" count shrinks toward zero, finished providers do no further work, and the campaign auto-completes (reporting "Force re-render complete"). Combines with--max-conversationsto spread the work across runs.mark_exportednow preservesjoplin_note_id/joplin_synced_at/joplin_resourcesacross re-exports (refreshingexported_at), so a re-export followed byjoplinupdates the existing notes instead of creating duplicates.- Fix:
.envis now loaded at the very start of every command, soCACHE_DIRandLOG_FILEare honored (previously the cache silently used./cacheregardless ofCACHE_DIR).
Archive hygiene
prunedeletes export files not referenced by the manifest (old-layout trees, no-ID orphans), with listing, confirmation,--dry-run,-y, and empty-directory sweep. Refuses to run when the manifest references no files (socache --clear+prunecan't wipe the archive). First live run removed 420 stale files (9.4 MB).doctorverifies manifest ↔ disk integrity: every recordedfile_pathmust exist on disk.
Changed
- ChatGPT role filter that dropped
tool/systemmessages is lifted; all roles route through normal extraction (truly empty messages skip via the empty-content guard). BaseProvider.normalize_conversationaccepts an optionalLossReportparameter.ChatGPTProvider.normalize_conversationreadsconversation_idas a fallback forid(live detail responses useconversation_id; fixtures useid).ConfigvalidatesEXPORTER_HIDDEN_CONTENT,EXPORTER_DOWNLOAD_MEDIA,MAX_CONVERSATIONS_PER_RUN, andREQUEST_DELAY, and logs the active values at startup.- New dependency:
browser-cookie3==0.20.1.
Fixed
- Custom Instructions (
user_editable_context/model_editable_context) were silently dropped from every conversation (parts-vs-direct-fields mismatch).
Migration
- JSON exports: messages contain typed
blocksand may omit the legacycontentfield — external consumers should preferblocks. - To re-render existing exports with the current rendering and collapse policy:
python -m src.main cache --clearthenpython -m src.main export, followed byjoplinto update notes (and upload media resources). Per-conversation message counts may increase as previously-dropped Custom Instructions, image-only turns, and tool-only turns now appear.
Test suite
- 264 tests, all passing.