Reading gizmo_id during normalization fixed attribution, but it cannot answer "which projects am I missing?": normalization only runs on conversations being exported, and a normal run skips everything already cached. Discovering the gaps would have meant --force re-exporting the whole archive. `ai-chat-exporter projects` does it directly. It lists conversations, collects the g-p- ids they belong to, resolves display names, and prints a table marking which are absent from .env plus a paste-ready CHATGPT_PROJECT_IDS line. --write applies it; --deep falls back to one detail request per conversation when the listing does not carry gizmo_id (unverified which shape this account returns, so the command reports which path it took rather than assuming). This matters beyond tidiness: attribution is now self-correcting, but the listing pass still needs the ids. Conversations that live only inside a project never appear in the default listing, so an unlisted project's chats are not merely misfiled — they are never fetched. 6 CLI tests: reporting, the already-configured case, the g-p- guard, --deep, the hint when --deep is needed, and --write. 324 pass.
18 KiB
Changelog
All notable changes to this project will be documented here. Format follows Keep a Changelog.
[Unreleased]
Fixed
- Deleted uploads are no longer reported as permission errors. ChatGPT's
/backend-api/files/{id}/downloadanswers a missing asset with403 {"detail":"Forbidden"}, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned404 {"detail":"File not found"}on/files/{id}, while assets that downloaded fine returned 200 on both in the same session, andChatGPT-Account-Idmade no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted asexpired-or-missing. A 403 on an asset that does still exist is left alone asforbidden— that one would be a real problem. - 4xx errors now report why.
_make_requestended non-retryable statuses withraise_for_status(), whose curl_cffi message isHTTP Error {code}: {reason}— and HTTP/2 carries no reason phrase, so a refused request logged as bareHTTP Error 403:and the response body (the only explanation the provider gives) was discarded. The body'sdetail/error/messageis now carried into theProviderError, redacted and truncated. This is what made the media 403s onGET /backend-api/files/{id}/downloadundiagnosable. redact_secretsmissed compound key names. It matched keys exactly, soaccess_token,api_key, andsession-tokenpassed through un-redacted into debug-logged response bodies; matching now applies per word ("keywords", "monkey", "tokenizer" stay intact).tests/test_config.py::TestSessionLimiterConfig::test_defaultsdepended on the developer's.env.load_config()callsload_dotenv(override=False), which re-populated the variable the test had just deleted — so it passed only on a machine with no.env. The test now stubs dotenv discovery.
Added
projectscommand — discover the project IDs your config is missing.CHATGPT_PROJECT_IDSis maintained by hand, and a project missing from it is invisible to the listing pass, so its conversations are never fetched. The command reports every project your conversations belong to, marks which are absent from.env, and prints a paste-ready line (--writeapplies it). It reads project ids from the conversation listing when they are there and falls back to--deep, one detail request per conversation, when they are not.- Project attribution now reads the conversation's own
gizmo_id. Previously the project name came only fromCHATGPT_PROJECT_IDS, so a conversation in a project you had not listed exported intono-project/even though its payload names its project. The detail response carriesgizmo_id, so it is used as a fallback after the listing annotation and the project map — attribution stays correct without maintaining a list, and moving a chat into a new project no longer silently misfiles it. Onlyg-p-ids count: a custom GPT is not a project and must not become a folder. Each unconfigured project is reported once per run, naming the id to add, because the listing pass still needsCHATGPT_PROJECT_IDS— conversations that live only inside a project never appear in the default listing.
Changed
- Media download failures are bucketed as
forbidden(403 — the file record survives) separately fromdownload-error, so the run summary distinguishes it fromexpired-or-missing(404).
Notes
- ChatGPT media 403s: investigated and closed (2026-08-17). 19 images across 7 conversations would not download. They are unrecoverable, and not because of anything the exporter or the export schedule did. Findings, recorded so this is not re-litigated: uploads do not expire (36/36 sampled images from 2025-09 through 2026-08 are still live, so export cadence is not a factor); the failures split into 7 records that 404 outright and 12 that report
state: "ready"with alibrary_file_id;/files/{id}/downloadis the endpoint that mints the signedestuary/contentURL the web UI fetches, and it refuses the survivors with a bare 403 regardless of headers,Authorization,gizmo_id,conversation_id, Referer, or namespace, while the Library id is rejected asfile_not_found. Decisively, those same images render blank in ChatGPT's own UI — nothing is being withheld from the exporter. All the survivors were created 2026-07-14 within minutes of each other yet appear in conversations predating that date, pointing at a Library migration that kept the metadata and lost the bytes.
[0.8.0] - 2026-07-06
Focused on the Claude Code provider after the tool's on-disk layout changed and Claude Code sessions proved hard to find in Joplin.
Added
- Subagent capture. Claude Code now stores subagent (Task-tool) transcripts as separate
<session>/subagents/agent-*.jsonlfiles (with anagent-*.meta.jsonsidecar). These were previously invisible to the exporter — all delegated work (reviews, research, plans) was lost. Each subagent is now folded into its parent session inline at theTask/Agentcall that spawned it, rendered as a collapsible<details>block labeled with itsagentType/description. The subagent's own tool traffic is collapsed under the sameEXPORTER_HIDDEN_CONTENTpolicy as the main dialogue; nesting is handled at any depth viatoolUseIdmatching. - Repo tags in Claude Code titles. Sessions launched from a workspace root all land in one folder-named notebook, so titles now carry the repos each session touched, e.g.
Resume StartWRT project work [start-technologies]— visible in the note list and matched by Joplin search. A file's repo is the git repository it lives in (nearest ancestor with a.git), resolved from tool_use paths anywhere in the filesystem — so cross-workspace work is captured and non-repo noise (config dirs, one-off files, reference dirs) is excluded because it isn't a git repo. Frequency-ordered, capped at 3. Escape hatch:CLAUDE_CODE_REPO_TAG_IGNORE(comma-separated repo names). Tags reflect the current git layout, so a since-deleted/moved repo drops from the tag on re-export. - Multiple projects roots.
CLAUDE_CODE_DIRnow accepts anos.pathsep-separated list of roots (a single path stays backward compatible), andCLAUDE_CONFIG_DIR'sprojects/tree is scanned automatically when set. Sessions from all roots are merged by launch-folder; a session UUID present in two roots keeps the newer-mtime copy.
Changed
- Claude Code gets its own top-level Joplin notebook,
AI-ClaudeCode(was nested underAI-Claudealongside Claude web chats, which made dev sessions hard to find). Existing Claude Code notes self-heal into the new notebook on the nextjoplinrun —update_notenow setsparent_id, so a changed provider→notebook mapping relocates notes in place instead of duplicating them. After migrating, the emptiedAI-Claude/{Myworkspace,Services,…}notebooks can be deleted by hand (Joplin does not auto-remove empty folders).
Notes
- The sibling
<session>/tool-results/*.txtsidecars (externalized large tool outputs) are intentionally not captured — tool_result content is collapsed under the default policy anyway.
[0.7.0] - 2026-06-28
Added
canarycommand — probes the ChatGPT/Claude web APIs for schema drift against the fields the parser actually depends on (one listing page + one conversation per provider), not the full response shape. Flags the silent-failure risks for a backup tool: a renamed retrieval-tool author bypassing the hidden-content collapse, a newcontent_type, drifted message fields, or non-empty Claudeattachments/filesthe normalizer ignores. ERROR findings exit non-zero; WARN findings are surfaced but non-fatal. SeeFUTURE.md§10.doctornow reports a real ChatGPT token-health check ("ChatGPT token active") based on the/api/auth/sessionerrorfield. ChatGPT session tokens are JWEs whoseexpis encrypted and unreadable client-side; the previous decode path could never yield an expiry. The honest signal iserror == "RefreshAccessTokenError"(verified live), which means the session token is dead even thoughexpires/accessTokenstill echo stale values. SeeFUTURE.md§9.
Fixed
ChatGPTProvider._fetch_access_tokennow checks the/api/auth/sessionerrorfield and fails fast with a clear "refresh your token" message. Previously it returned the staleaccessTokenpresent on a dead session, producing a confusing downstream 401 instead of an actionable auth error.
Removed
auth --from-browserand thebrowser-cookie3dependency (src/browser_tokens.py). Browser cookie auto-extraction is not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level layer that can't be decrypted off disk without Administrator rights and AV-flagged SYSTEM impersonation, and fails on Brave specifically (symptom:Unable to get key for cookie decryption). Auth is manual (DevTools wizard) only. SeeFUTURE.md§2 for the full rationale.
[0.6.0] - 2026-06-12
First tagged release. The project was developed through several internal milestones (0.1.0–0.5.0) that were never published; their changes are consolidated here.
Added
Core export & sync
- ChatGPT and Claude export via their internal web APIs, with a local cache/manifest for incremental, resumable sync (only new or updated conversations are fetched; each is written to the manifest immediately).
- Markdown and JSON exporters. Markdown rendering happens at exporter-write time (providers produce typed blocks; exporters render them), which keeps the door open for future Obsidian/HTML output.
- CLI:
export,list,cache,doctor,auth,joplin,prune. --projectfilter onexport/list/joplin(case-insensitive substring, ornone).- ChatGPT Projects support via
CHATGPT_PROJECT_IDS(project conversations are fetched separately from the default listing).
Joplin integration
joplincommand syncs exported Markdown to Joplin as notes; notebooks are created automatically and nested per provider+project. Re-running is safe — the Joplin note ID is stored in the manifest, so notes are updated, not duplicated.JOPLIN_API_TOKEN,JOPLIN_API_URL,JOPLIN_REQUEST_TIMEOUTconfig, with actionable error messages on timeout/connection failure.
Rich content (typed blocks)
- Messages carry an ordered
blockslist: text, code, thinking, tool_use, tool_result, citation, image_placeholder, file_placeholder, unknown. - ChatGPT voice mode:
audio_transcriptionparts render as text; audio asset pointers render as📎 File attachedplaceholders with size/duration. Custom Instructions (user_editable_context/model_editable_context) now appear (were silently dropped) with a> ℹ️ Hidden contextmarker. - ChatGPT
execution_output,system_error, andtether_browsing_displayrender astool_resultblocks (withtool_name,is_error, optionalsummary); transient browse spinners skip silently. - Defensive Claude block extraction (text/thinking/tool_use/tool_result/image, including nested-block flattening).
_safe_fencepicks a backtick fence longer than any run in the content, so embedded triple-backticks can't corrupt rendering (verified live in Joplin).- Data-loss visibility: a
LossReportsummary at the end of everyexportrun breaks downunknown blocksandextraction failuresby raw type;unknownblocks render as visible> ⚠️ Unsupported contentwith the raw type and observed keys.
Invisible-content collapse (EXPORTER_HIDDEN_CONTENT=full|placeholder|omit, default placeholder; export --hidden-content)
- Messages that were invisible in the ChatGPT web UI are collapsed to one-line placeholders instead of exported in full. Two triggers: (a) tool-role retrieval dumps identified by
author.name(file_search,myfiles_browser) — ChatGPT re-injects the full text of attached/project files on every tool run, measured at 86% of all content bytes across the three largest conversations and not hidden-flagged; (b) messages flaggedis_visually_hidden_from_conversation(Custom Instructions). Code-execution results, web-search output, and all dialogue are untouched. - New
collapsedblock renders as> 🔧 **Tool output** — file_search (15.0 KB) — omitted (EXPORTER_HIDDEN_CONTENT=full to keep)(ℹ️ Hidden context variant for hidden-flagged messages). TheLossReportgains acollapsed by policysection (count per origin + total KB) so the omission stays visible.
Claude Code session provider (--provider claude-code)
- Archives local agent transcripts from
~/.claude/projects/*/*.jsonl(override withCLAUDE_CODE_DIR) — no tokens, no rate limits, no ToS exposure. Prose-only by default per the hidden-content policy: dialogue and deliverables kept; tool traffic grouped into one placeholder per activity run (> 🔧 Tool output — 14 calls: Read ×9, Bash ×2 (86KB) — omitted); thinking dropped (counted in the summary); subagent (isSidechain) and harness (isMeta, snapshots, command tags) records stripped. Titles from the lastai-titlerecord; project from the working-directory basename. Joplin notebooks nest under theAI-Claudeparent. Measured: 19MB of JSONL → 804KB of Markdown across 25 sessions.
Binary/image downloads (EXPORTER_DOWNLOAD_MEDIA=images|all|off, default images; export --download-media)
- Downloads ChatGPT conversation assets into a
media/folder beside each export and inlines them — images become realembeds, files become local links. Two-hop fetch via/backend-api/files/{id}/download→ signed URL (verified live).allalso pulls voice-mode audio (≈162MB of clips in this archive, transcripts already in text — hence the images-only default). Idempotent (skips assets already on disk), atomic writes,600perms. Expired/missing assets (old generated images 404) keep their placeholder and are counted asmedia failed— never fatal. - Joplin resource upload: the
joplincommand uploads downloaded media as resources and rewritesmedia/…links to:/resourceId, so images render inside notes (verified live: resources created with correct mime/size and linked to the note). Resource IDs are tracked per-conversation in the manifest, so re-syncs reuse them instead of duplicating.
Token setup & throughput
auth --from-browser [brave|chrome|chromium|edge|firefox]: extracts ChatGPT session-token cookies and the ClaudesessionKeystraight from a local browser's cookie store (viabrowser-cookie3; defaults to Brave), validates each against the live API, and writes.env. Tokens that fail validation never overwrite a working entry; clear fallback when the browser is absent or the cookie DB is locked.- Session limiter:
--max-conversations N/MAX_CONVERSATIONS_PER_RUNcap downloads per run (per provider); capped-out conversations are reported as "deferred" and resumed next run. - Polite pacing:
REQUEST_DELAY(default 1.0s, ±25% jitter,0disables) between consecutive API requests, with the existing 429 backoff as the reactive net.
Re-rendering
export --forcere-exports every conversation even if cached and unchanged, so the whole archive can be re-rendered after a formatting/feature change withoutcache --clear. Runs as a tracked campaign (a stamped start time in the manifest): each run re-renders the least-recently-exported conversations, the "still to go" count shrinks toward zero, finished providers do no further work, and the campaign auto-completes (reporting "Force re-render complete"). Combines with--max-conversationsto spread the work across runs.mark_exportednow preservesjoplin_note_id/joplin_synced_at/joplin_resourcesacross re-exports (refreshingexported_at), so a re-export followed byjoplinupdates the existing notes instead of creating duplicates.- Fix:
.envis now loaded at the very start of every command, soCACHE_DIRandLOG_FILEare honored (previously the cache silently used./cacheregardless ofCACHE_DIR).
Archive hygiene
prunedeletes export files not referenced by the manifest (old-layout trees, no-ID orphans), with listing, confirmation,--dry-run,-y, and empty-directory sweep. Refuses to run when the manifest references no files (socache --clear+prunecan't wipe the archive). First live run removed 420 stale files (9.4 MB).doctorverifies manifest ↔ disk integrity: every recordedfile_pathmust exist on disk.
Changed
- ChatGPT role filter that dropped
tool/systemmessages is lifted; all roles route through normal extraction (truly empty messages skip via the empty-content guard). BaseProvider.normalize_conversationaccepts an optionalLossReportparameter.ChatGPTProvider.normalize_conversationreadsconversation_idas a fallback forid(live detail responses useconversation_id; fixtures useid).ConfigvalidatesEXPORTER_HIDDEN_CONTENT,EXPORTER_DOWNLOAD_MEDIA,MAX_CONVERSATIONS_PER_RUN, andREQUEST_DELAY, and logs the active values at startup.- New dependency:
browser-cookie3==0.20.1.
Fixed
- Custom Instructions (
user_editable_context/model_editable_context) were silently dropped from every conversation (parts-vs-direct-fields mismatch).
Migration
- JSON exports: messages contain typed
blocksand may omit the legacycontentfield — external consumers should preferblocks. - To re-render existing exports with the current rendering and collapse policy:
python -m src.main cache --clearthenpython -m src.main export, followed byjoplinto update notes (and upload media resources). Per-conversation message counts may increase as previously-dropped Custom Instructions, image-only turns, and tool-only turns now appear.
Test suite
- 264 tests, all passing.