Files
AIChatExporter/CHANGELOG.md
T
JesseMarkowitzandClaude Opus 5.5 64068bb19b fix: fork subagents recursed forever in the Claude Code export
A fork subagent's transcript opens with a copy of the parent turn that spawned it, its own Agent call included; folding that call re-entered the same fork until RecursionError, failing every daily sync since 2026-09-24. Skip a spawn call while its own subagent is being expanded, and strip the <fork-boilerplate> preamble.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
2026-10-05 07:14:16 -04:00

32 KiB
Raw Blame History

Changelog

All notable changes to this project will be documented here. Format follows Keep a Changelog.

[Unreleased]

Fixed

  • Every scheduled Claude Code sync since 2026-09-24 died with RecursionError. Claude Code's fork subagents write a transcript that opens with a copy of the parent turn that spawned them — the fork's own Agent call included. _extract_messages folds a subagent inline whenever it meets its spawn call, so it met that copy inside the fork, folded the same fork again, and recursed until Python's limit. One fork anywhere in ~/.claude/projects was enough to fail the whole provider; seven existed across three sessions, and the Codex half of the run, unaffected, hid the cause behind a generic exit 1.

    The recursive pass now carries the spawn ids being expanded around it and skips a tool_use whose id is among them — the enclosing subagent block already stands for that call. Because the set accumulates, a longer cycle (A spawns B, whose transcript re-spawns A) stops too. The fork's <fork-boilerplate> preamble — generic worker rules the harness prepends — is now stripped with the other harness tags, so a fork block opens with its actual directive.

    TestSubagentFold.test_fork_containing_its_own_spawn_call reproduces the on-disk shape (a fork-context-ref record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original RecursionError against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block.

  • An expired Claude session key reported a raw JSON dump instead of how to fix it. _make_request routed only 401 to the auth handler (src/providers/base.py), and claude.ai does not use 401 — an invalid or expired sessionKey comes back as 403 permission_error with details.error_code = account_session_invalid. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it."

    Measured live 2026-09-20 against GET /api/organizations: a valid key returns 200, while an expired key, a deliberately malformed key and no cookie at all return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated."

    Auth detection is now a provider decision rather than a hardcoded status. BaseProvider._is_auth_failure(response) defaults to 401 and ClaudeProvider overrides it to add 403 matched on account_session_invalid, not on the bare status — so a genuine permission error, which carries a different code, is still reported as itself rather than being mislabelled an expired key. _handle_401 is renamed _handle_auth_failure and takes the response, because a handler named for one status that must handle two is how this stayed hidden; its messages now state the status actually observed instead of asserting "401 Unauthorized". ChatGPT is unaffected: it does not override the default, so the deleted-asset 403 path is untouched.

    Seven regression tests cover the split (TestAuthFailureDetection), including the two that matter most: a Claude 403 with a different error code must not be treated as an auth failure, and a ChatGPT 403 must not either. The docs that repeated the wrong premise — README.md's expiry table and "When Tokens Expire" section, the auth wizard's on-screen note, and the ClaudeProvider docstring — are corrected in the same change.

  • The docs claimed both ChatGPT cookie chunks were required; they are not. README.md stated flatly that "ChatGPT splits large session tokens across two cookies to stay under the browser's 4KB cookie limit. Both are required," and .env.example documented only the chunked layout — so a machine whose session token happens to fit in a single __Secure-next-auth.session-token cookie looked broken, with the user hunting for a .1 that does not exist. Chrome splits a cookie only above ~4KB, so the layout varies by session size and the same account can be chunked on one machine and not on another.

    The code was already correct: CHATGPT_SESSION_TOKEN_1 is optional (src/providers/chatgpt.py:162) and the auth wizard already told you to paste a lone cookie into .0 and leave .1 blank. Only the reference docs were wrong, and they are the ones read when setting up a new machine.

    Measured 2026-09-20 against /api/auth/session, reassembling a real 4089-byte token to test each naming: chunked .0+.1 → 200 with an accessToken; the whole value under the unchunked name → 200 with an accessToken; the whole value under .0 alone → 200 with an accessToken. The server reassembles a complete value sent under .0, so both layouts authenticate as the code already assumed. A partial .0 with its .1 omitted is the one combination that fails, and it fails silently — HTTP 200 with no accessToken rather than an error — which is now documented in both files alongside the correction.

  • The test suite sent real push notifications to the developer's phone. TestSyncCommand invokes the actual sync command, which calls load_config(), which calls load_dotenv() — so the real .env was loaded and its live NTFY_TOPIC used for the POST. Every pytest run fired three or four pushes, including a fabricated "codex: 3 conversation(s) failed to export" straight out of a fixture, which is worse than noise: it reports a failure that never happened. Nothing appeared in cache/logs/exporter.log to explain it, because every test invocation passes --no-log-file.

    A tests/conftest.py autouse fixture now neutralises the environment for every test: NTFY_TOPIC/NTFY_TOKEN are emptied, and NTFY_SERVER and JOPLIN_API_URL are pointed at a closed local port, so a stray topic cannot reach the internet and a test cannot write notes into a real Joplin instance. The values are emptied rather than deleted — load_dotenv(override=False) skips only keys already present, so deleting one lets .env put it back. Verified by instrumenting requests across a full run: zero outbound requests, where the same instrumentation without the fixture records POSTs to the live ntfy topic.

[0.9.0] - 2026-08-18

Fixed

  • An em dash in a notification title silently dropped the notification. HTTP header values are latin-1 at best and requests raises on anything outside it, so the first real send failed with 'latin-1' codec can't encode character '\u2014'. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
  • The terms-of-service gate exited 0 without a terminal. click.prompt raises Abort on a closed stdin, which the handler treated as a user Ctrl-C and exited 0 — so a scheduled run on a machine that had never acknowledged the notice would report success having archived nothing. Non-interactive invocations now exit 1 with an explanation of how to clear the gate once by hand. Found by running the new systemd unit rather than by reading the code.
  • A single U+0085 in a transcript silently dropped a whole record. Both local providers split session files with str.splitlines(), which breaks not just on \n but on U+0085 (NEL), U+2028 and U+2029 — all of which are legal inside a JSON string and are written literally by Codex (Rust does not escape non-ASCII). One NEL in captured command output shredded one record into unparseable fragments; the parser logged "skipped 3 unparseable line(s)" and lost the record. Found while exporting a real rollout. Both providers now split on \n only, and both have regression tests that write their fixtures with ensure_ascii=False — with json.dumps' default the hazardous characters are escaped and the bug cannot reproduce.
  • Deleted uploads are no longer reported as permission errors. ChatGPT's /backend-api/files/{id}/download answers a missing asset with 403 {"detail":"Forbidden"}, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned 404 {"detail":"File not found"} on /files/{id}, while assets that downloaded fine returned 200 on both in the same session, and ChatGPT-Account-Id made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as expired-or-missing. A 403 on an asset that does still exist is left alone as forbidden — that one would be a real problem.
  • 4xx errors now report why. _make_request ended non-retryable statuses with raise_for_status(), whose curl_cffi message is HTTP Error {code}: {reason} — and HTTP/2 carries no reason phrase, so a refused request logged as bare HTTP Error 403: and the response body (the only explanation the provider gives) was discarded. The body's detail/error/message is now carried into the ProviderError, redacted and truncated. This is what made the media 403s on GET /backend-api/files/{id}/download undiagnosable.
  • redact_secrets missed compound key names. It matched keys exactly, so access_token, api_key, and session-token passed through un-redacted into debug-logged response bodies; matching now applies per word ("keywords", "monkey", "tokenizer" stay intact).
  • tests/test_config.py::TestSessionLimiterConfig::test_defaults depended on the developer's .env. load_config() calls load_dotenv(override=False), which re-populated the variable the test had just deleted — so it passed only on a machine with no .env. The test now stubs dotenv discovery.

Added

  • ntfy push notifications for unattended runs (NTFY_TOPIC). A scheduled archive reports only to places you have to go and look at — the log file, the systemd journal, Task Scheduler's exit code — so a run that quietly failed every morning would stay quiet. sync now pushes its result: success carries per-provider counts at low/default priority, failure carries the reason at high priority with an alert tag, so the two are distinguishable at a glance on a phone. ai-chat-exporter notify shows the settings and --test sends a test push. NTFY_NOTIFY=failure limits it to failures, off disables, --notify/--no-notify override per run. Never fatal: an unreachable ntfy logs a warning and the run still reports its real exit code.

    The payload is counts and a machine name only, never conversation titles — a topic on public ntfy.sh is readable by anyone who knows its name, and a title is the first line of what you asked. The machine name is included because several machines archive into one topic, where "2 exported" means nothing on its own. NTFY_SERVER points at a self-hosted instance and NTFY_TOKEN authenticates against an access-controlled topic.

  • ai-chat-exporter / ai-chat-exporter.cmd launchers — no virtualenv ceremony. cd into the repo and run; the wrapper creates .venv, installs dependencies on first run, and reinstalls when pyproject.toml changes. A fresh clone goes from nothing to a working command in one step (measured: ~10s), on Linux/macOS and Windows alike, which matters for a tool meant to run on several machines. cmd.exe searches the current directory before PATH, so Windows needs no .\ prefix. The working directory is deliberately not changed — .env, cache/ and exports/ still resolve against it, which is what lets one checkout archive different machines into different places — but the wrapper now warns when you run it from elsewhere, because a different cache/manifest.json silently starts a second archive rather than failing.

  • sync command — export then joplin in one invocation, with a real exit code. Intended for schedulers (and the "trivial add-on" FUTURE.md §7 anticipated): it exits non-zero if any conversation failed to export or any note failed to sync, so a scheduled run that achieved nothing is distinguishable from one that had nothing to do. --skip-joplin exports only; --joplin-optional downgrades an unreachable Joplin to a warning, since the export has already captured the local transcripts and the notes rebuild from the cache on the next run that finds Joplin up.

  • Daily scheduling for both platforms. scheduling/install-systemd-timer.sh (systemd user timer, Persistent=true so a machine that was off catches up at boot) and scheduling/Register-AiChatSyncTask.ps1 (per-user Task Scheduler entry, -StartWhenAvailable). --provider is repeatable in both, because the right set differs per machine: the local providers need no credentials and always work unattended, while a web provider whose session token has expired would fail the job every single day and train you to ignore it.

  • Codex CLI provider (--provider codex). Archives local Codex agent transcripts from ~/.codex/sessions/**/rollout-*.jsonl — local-only, like claude-code: no tokens, no rate limits, no ToS exposure. Sessions land in their own top-level AI-Codex Joplin notebook, with the same prose-only default, repo tags (CODEX_REPO_TAG_IGNORE) and multi-root scanning (CODEX_DIR, plus $CODEX_HOME/sessions).

    Codex writes each session twice in one file and the choice between the two layers is the whole design. response_item records are the model-facing wire format, where a tool call arrives as JavaScript (tools.exec_command({...})) because Codex's exec tool is code-mode; event_msg/item_completed records are Codex's own typed items, already decoded into CommandExecution/FileChange/Extension with argv, cwd, exit code and output as fields. Measured over 7 sessions on 0.147.0 (2026-08-18), the typed layer is 1:1 with the raw layer for prose (91 AgentMessage ↔ 91 assistant messages, sharing ids) and additionally omits every piece of harness plumbing — all 51 developer-role messages plus the 7 # AGENTS.md instructions… and 1 <environment_context> injections — which the Claude Code provider has to strip by regex. So the typed layer is parsed for content.

    Its one gap is that it records what ran, not what was attempted: 26 of 180 exec_command calls produced no item (14 sandbox launch failures, 6 user aborts, ~5 still running at turn end, 1 failure). The raw layer is therefore read for a call count only, and placeholders report the shortfall — 3 calls: exec_command ×3 (+2 did not complete) — instead of silently under-reporting. wait calls are process polls, not attempts, and are excluded.

    Reasoning is not exportable from Codex. All 345 reasoning records carry encrypted_content, with summary empty in the raw layer and summary_text/raw_content empty in the typed layer, in every session. It is always dropped and counted; unlike Claude Code, EXPORTER_HIDDEN_CONTENT=full cannot surface it. The sidecar SQLite databases (state_5.sqlite, thread_history_1.sqlite) are deliberately not read: thread_history_projection_state tracks a byte offset into the rollout file, so the JSONL is canonical and the DB derived, and its title column is just the first user message truncated.

    Codex Cloud is out of scope, verified rather than assumed. Cloud tasks are reachable at chatgpt.com/backend-api/api/codex/tasks{,/list} — the same host and /backend-api root the ChatGPT provider already uses — but local CLI sessions are never uploaded there, so it is not an alternative source for these transcripts. codex cloud list confirmed the account holds no cloud tasks. The provider makes no network calls.

  • projects command — discover the project IDs your config is missing. CHATGPT_PROJECT_IDS is maintained by hand, and a project missing from it is invisible to the listing pass, so its conversations are never fetched. The command reports every project your conversations belong to, marks which are absent from .env, and prints a paste-ready line (--write applies it). It reads project ids from the conversation listing when they are there and falls back to --deep, one detail request per conversation, when they are not.

  • Project attribution now reads the conversation's own gizmo_id. Previously the project name came only from CHATGPT_PROJECT_IDS, so a conversation in a project you had not listed exported into no-project/ even though its payload names its project. The detail response carries gizmo_id, so it is used as a fallback after the listing annotation and the project map — attribution stays correct without maintaining a list, and moving a chat into a new project no longer silently misfiles it. Only g-p- ids count: a custom GPT is not a project and must not become a folder. Each unconfigured project is reported once per run, naming the id to add, because the listing pass still needs CHATGPT_PROJECT_IDS — conversations that live only inside a project never appear in the default listing.

Changed

  • Media download failures are bucketed as forbidden (403 — the file record survives) separately from download-error, so the run summary distinguishes it from expired-or-missing (404).

Notes

  • ChatGPT media 403s: investigated and closed (2026-08-17). 19 images across 7 conversations would not download. They are unrecoverable, and not because of anything the exporter or the export schedule did. Findings, recorded so this is not re-litigated: uploads do not expire (36/36 sampled images from 2025-09 through 2026-08 are still live, so export cadence is not a factor); the failures split into 7 records that 404 outright and 12 that report state: "ready" with a library_file_id; /files/{id}/download is the endpoint that mints the signed estuary/content URL the web UI fetches, and it refuses the survivors with a bare 403 regardless of headers, Authorization, gizmo_id, conversation_id, Referer, or namespace, while the Library id is rejected as file_not_found. Decisively, those same images render blank in ChatGPT's own UI — nothing is being withheld from the exporter. All the survivors were created 2026-07-14 within minutes of each other yet appear in conversations predating that date, pointing at a Library migration that kept the metadata and lost the bytes.

[0.8.0] - 2026-07-06

Focused on the Claude Code provider after the tool's on-disk layout changed and Claude Code sessions proved hard to find in Joplin.

Added

  • Subagent capture. Claude Code now stores subagent (Task-tool) transcripts as separate <session>/subagents/agent-*.jsonl files (with an agent-*.meta.json sidecar). These were previously invisible to the exporter — all delegated work (reviews, research, plans) was lost. Each subagent is now folded into its parent session inline at the Task/Agent call that spawned it, rendered as a collapsible <details> block labeled with its agentType/description. The subagent's own tool traffic is collapsed under the same EXPORTER_HIDDEN_CONTENT policy as the main dialogue; nesting is handled at any depth via toolUseId matching.
  • Repo tags in Claude Code titles. Sessions launched from a workspace root all land in one folder-named notebook, so titles now carry the repos each session touched, e.g. Resume StartWRT project work [start-technologies] — visible in the note list and matched by Joplin search. A file's repo is the git repository it lives in (nearest ancestor with a .git), resolved from tool_use paths anywhere in the filesystem — so cross-workspace work is captured and non-repo noise (config dirs, one-off files, reference dirs) is excluded because it isn't a git repo. Frequency-ordered, capped at 3. Escape hatch: CLAUDE_CODE_REPO_TAG_IGNORE (comma-separated repo names). Tags reflect the current git layout, so a since-deleted/moved repo drops from the tag on re-export.
  • Multiple projects roots. CLAUDE_CODE_DIR now accepts an os.pathsep-separated list of roots (a single path stays backward compatible), and CLAUDE_CONFIG_DIR's projects/ tree is scanned automatically when set. Sessions from all roots are merged by launch-folder; a session UUID present in two roots keeps the newer-mtime copy.

Changed

  • Claude Code gets its own top-level Joplin notebook, AI-ClaudeCode (was nested under AI-Claude alongside Claude web chats, which made dev sessions hard to find). Existing Claude Code notes self-heal into the new notebook on the next joplin run — update_note now sets parent_id, so a changed provider→notebook mapping relocates notes in place instead of duplicating them. After migrating, the emptied AI-Claude/{Myworkspace,Services,…} notebooks can be deleted by hand (Joplin does not auto-remove empty folders).

Notes

  • The sibling <session>/tool-results/*.txt sidecars (externalized large tool outputs) are intentionally not captured — tool_result content is collapsed under the default policy anyway.

[0.7.0] - 2026-06-28

Added

  • canary command — probes the ChatGPT/Claude web APIs for schema drift against the fields the parser actually depends on (one listing page + one conversation per provider), not the full response shape. Flags the silent-failure risks for a backup tool: a renamed retrieval-tool author bypassing the hidden-content collapse, a new content_type, drifted message fields, or non-empty Claude attachments/files the normalizer ignores. ERROR findings exit non-zero; WARN findings are surfaced but non-fatal. See FUTURE.md §10.
  • doctor now reports a real ChatGPT token-health check ("ChatGPT token active") based on the /api/auth/session error field. ChatGPT session tokens are JWEs whose exp is encrypted and unreadable client-side; the previous decode path could never yield an expiry. The honest signal is error == "RefreshAccessTokenError" (verified live), which means the session token is dead even though expires/accessToken still echo stale values. See FUTURE.md §9.

Fixed

  • ChatGPTProvider._fetch_access_token now checks the /api/auth/session error field and fails fast with a clear "refresh your token" message. Previously it returned the stale accessToken present on a dead session, producing a confusing downstream 401 instead of an actionable auth error.

Removed

  • auth --from-browser and the browser-cookie3 dependency (src/browser_tokens.py). Browser cookie auto-extraction is not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level layer that can't be decrypted off disk without Administrator rights and AV-flagged SYSTEM impersonation, and fails on Brave specifically (symptom: Unable to get key for cookie decryption). Auth is manual (DevTools wizard) only. See FUTURE.md §2 for the full rationale.

[0.6.0] - 2026-06-12

First tagged release. The project was developed through several internal milestones (0.1.0–0.5.0) that were never published; their changes are consolidated here.

Added

Core export & sync

  • ChatGPT and Claude export via their internal web APIs, with a local cache/manifest for incremental, resumable sync (only new or updated conversations are fetched; each is written to the manifest immediately).
  • Markdown and JSON exporters. Markdown rendering happens at exporter-write time (providers produce typed blocks; exporters render them), which keeps the door open for future Obsidian/HTML output.
  • CLI: export, list, cache, doctor, auth, joplin, prune.
  • --project filter on export/list/joplin (case-insensitive substring, or none).
  • ChatGPT Projects support via CHATGPT_PROJECT_IDS (project conversations are fetched separately from the default listing).

Joplin integration

  • joplin command syncs exported Markdown to Joplin as notes; notebooks are created automatically and nested per provider+project. Re-running is safe — the Joplin note ID is stored in the manifest, so notes are updated, not duplicated.
  • JOPLIN_API_TOKEN, JOPLIN_API_URL, JOPLIN_REQUEST_TIMEOUT config, with actionable error messages on timeout/connection failure.

Rich content (typed blocks)

  • Messages carry an ordered blocks list: text, code, thinking, tool_use, tool_result, citation, image_placeholder, file_placeholder, unknown.
  • ChatGPT voice mode: audio_transcription parts render as text; audio asset pointers render as 📎 File attached placeholders with size/duration. Custom Instructions (user_editable_context/model_editable_context) now appear (were silently dropped) with a > ℹ️ Hidden context marker.
  • ChatGPT execution_output, system_error, and tether_browsing_display render as tool_result blocks (with tool_name, is_error, optional summary); transient browse spinners skip silently.
  • Defensive Claude block extraction (text/thinking/tool_use/tool_result/image, including nested-block flattening).
  • _safe_fence picks a backtick fence longer than any run in the content, so embedded triple-backticks can't corrupt rendering (verified live in Joplin).
  • Data-loss visibility: a LossReport summary at the end of every export run breaks down unknown blocks and extraction failures by raw type; unknown blocks render as visible > ⚠️ Unsupported content with the raw type and observed keys.

Invisible-content collapse (EXPORTER_HIDDEN_CONTENT=full|placeholder|omit, default placeholder; export --hidden-content)

  • Messages that were invisible in the ChatGPT web UI are collapsed to one-line placeholders instead of exported in full. Two triggers: (a) tool-role retrieval dumps identified by author.name (file_search, myfiles_browser) — ChatGPT re-injects the full text of attached/project files on every tool run, measured at 86% of all content bytes across the three largest conversations and not hidden-flagged; (b) messages flagged is_visually_hidden_from_conversation (Custom Instructions). Code-execution results, web-search output, and all dialogue are untouched.
  • New collapsed block renders as > 🔧 **Tool output** — file_search (15.0 KB) — omitted (EXPORTER_HIDDEN_CONTENT=full to keep) (ℹ️ Hidden context variant for hidden-flagged messages). The LossReport gains a collapsed by policy section (count per origin + total KB) so the omission stays visible.

Claude Code session provider (--provider claude-code)

  • Archives local agent transcripts from ~/.claude/projects/*/*.jsonl (override with CLAUDE_CODE_DIR) — no tokens, no rate limits, no ToS exposure. Prose-only by default per the hidden-content policy: dialogue and deliverables kept; tool traffic grouped into one placeholder per activity run (> 🔧 Tool output — 14 calls: Read ×9, Bash ×2 (86KB) — omitted); thinking dropped (counted in the summary); subagent (isSidechain) and harness (isMeta, snapshots, command tags) records stripped. Titles from the last ai-title record; project from the working-directory basename. Joplin notebooks nest under the AI-Claude parent. Measured: 19MB of JSONL → 804KB of Markdown across 25 sessions.

Binary/image downloads (EXPORTER_DOWNLOAD_MEDIA=images|all|off, default images; export --download-media)

  • Downloads ChatGPT conversation assets into a media/ folder beside each export and inlines them — images become real ![](media/…) embeds, files become local links. Two-hop fetch via /backend-api/files/{id}/download → signed URL (verified live). all also pulls voice-mode audio (≈162MB of clips in this archive, transcripts already in text — hence the images-only default). Idempotent (skips assets already on disk), atomic writes, 600 perms. Expired/missing assets (old generated images 404) keep their placeholder and are counted as media failed — never fatal.
  • Joplin resource upload: the joplin command uploads downloaded media as resources and rewrites media/… links to :/resourceId, so images render inside notes (verified live: resources created with correct mime/size and linked to the note). Resource IDs are tracked per-conversation in the manifest, so re-syncs reuse them instead of duplicating.

Token setup & throughput

  • auth --from-browser [brave|chrome|chromium|edge|firefox]: extracts ChatGPT session-token cookies and the Claude sessionKey straight from a local browser's cookie store (via browser-cookie3; defaults to Brave), validates each against the live API, and writes .env. Tokens that fail validation never overwrite a working entry; clear fallback when the browser is absent or the cookie DB is locked.
  • Session limiter: --max-conversations N / MAX_CONVERSATIONS_PER_RUN cap downloads per run (per provider); capped-out conversations are reported as "deferred" and resumed next run.
  • Polite pacing: REQUEST_DELAY (default 1.0s, ±25% jitter, 0 disables) between consecutive API requests, with the existing 429 backoff as the reactive net.

Re-rendering

  • export --force re-exports every conversation even if cached and unchanged, so the whole archive can be re-rendered after a formatting/feature change without cache --clear. Runs as a tracked campaign (a stamped start time in the manifest): each run re-renders the least-recently-exported conversations, the "still to go" count shrinks toward zero, finished providers do no further work, and the campaign auto-completes (reporting "Force re-render complete"). Combines with --max-conversations to spread the work across runs.
  • mark_exported now preserves joplin_note_id / joplin_synced_at / joplin_resources across re-exports (refreshing exported_at), so a re-export followed by joplin updates the existing notes instead of creating duplicates.
  • Fix: .env is now loaded at the very start of every command, so CACHE_DIR and LOG_FILE are honored (previously the cache silently used ./cache regardless of CACHE_DIR).

Archive hygiene

  • prune deletes export files not referenced by the manifest (old-layout trees, no-ID orphans), with listing, confirmation, --dry-run, -y, and empty-directory sweep. Refuses to run when the manifest references no files (so cache --clear + prune can't wipe the archive). First live run removed 420 stale files (9.4 MB).
  • doctor verifies manifest ↔ disk integrity: every recorded file_path must exist on disk.

Changed

  • ChatGPT role filter that dropped tool/system messages is lifted; all roles route through normal extraction (truly empty messages skip via the empty-content guard).
  • BaseProvider.normalize_conversation accepts an optional LossReport parameter.
  • ChatGPTProvider.normalize_conversation reads conversation_id as a fallback for id (live detail responses use conversation_id; fixtures use id).
  • Config validates EXPORTER_HIDDEN_CONTENT, EXPORTER_DOWNLOAD_MEDIA, MAX_CONVERSATIONS_PER_RUN, and REQUEST_DELAY, and logs the active values at startup.
  • New dependency: browser-cookie3==0.20.1.

Fixed

  • Custom Instructions (user_editable_context/model_editable_context) were silently dropped from every conversation (parts-vs-direct-fields mismatch).

Migration

  • JSON exports: messages contain typed blocks and may omit the legacy content field — external consumers should prefer blocks.
  • To re-render existing exports with the current rendering and collapse policy: python -m src.main cache --clear then python -m src.main export, followed by joplin to update notes (and upload media resources). Per-conversation message counts may increase as previously-dropped Custom Instructions, image-only turns, and tool-only turns now appear.

Test suite

  • 264 tests, all passing.