Expiry is now ruled out: 36/36 sampled images from 2025-09 through
2026-08 are still live server-side, so nothing dies of age and export
cadence is not the variable. The loss is per-asset — 2026-07-09 kept 38
images and lost 12 in one conversation.
Next hypothesis: those images belong to messages that were edited or
regenerated. ChatGPT stores a conversation as a tree and editing forks it,
stranding the superseded messages on an abandoned branch. The exporter
walks every node (chatgpt.py:963), so it exports those branches too — and
an attachment unreachable from live history is a natural GC target.
branch_check.py fetches the raw conversation, derives the live branch by
walking parent links up from current_node, and cross-tabulates every image
against (on the live branch?) x (still downloadable?). If the lost images
are all off-branch and nothing on the live branch is missing, this is not
data loss at all — it is attachments to messages that were replaced.
analyze_media_age.py claimed age was ruled out because an image from
2025-11 was saved while one from 2026-08 was not. That conclusion does not
follow. "Saved" means some earlier run downloaded it, not that ChatGPT
still holds it — an image captured in June looks saved forever after, even
if it died in July. A month with no losses shows the exports were timely,
not that the assets survived.
- probe_survival_by_month.py: sample images already on disk, grouped by
their conversation's month, and ask /files/{id} whether each still
exists today. Old months still 200 → no expiry, and cadence did not save
them. Old months now 404 → uploads do expire and cadence is the whole
ballgame.
- analyze_media_age.py: stop asserting the unsupported verdict; say what
the number does and does not show, and point at the probe.
One signal there is immune to the confound and survives: 2026-07-09 kept
38 images and lost 12. Same conversation, same day, opposite outcomes —
no retention policy does that, so at least part of this is per-asset.
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
( when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.
Offline, no token, no API calls.
The improved error reporting landed, and the answer it produced is
{"detail":"Forbidden"} — generic, no reason. The cause has to be narrowed
by experiment instead, so collect the experiments in one script:
- classify the failed assets offline from the exported placeholders
(user_upload vs model_generated) — no API call needed
- probe a known-good asset in the same session as a control, to rule the
session in or out
- retry the 403 with ChatGPT-Account-Id, which the exporter never sends
and which workspace-scoped resources can require
- compare /files/{id} against /files/{id}/download
Temporary: delete once the cause is known, or fold into `doctor` if the
check earns a permanent home.
Media downloads logged "HTTP Error 403:" with no reason. That string is
curl_cffi's raise_for_status() format, "HTTP Error {code}: {reason}", and
HTTP/2 carries no reason phrase — so the message said nothing, and
_make_request threw the response body away. The provider's JSON `detail`
is the only explanation available for a refused asset.
- base._make_request: end non-retryable statuses with a ProviderError
carrying the body's detail/error/message (redacted, truncated to 300
chars) instead of a bare raise_for_status().
- media: bucket 403 as `forbidden` in the run summary, separately from
`download-error` — "the asset is gone" and "we were refused" are
different problems.
- utils.redact_secrets: match secret key names per word. Exact matching
let access_token, api_key, and session-token through into logged
bodies; "keywords"/"monkey"/"tokenizer" stay intact.
- tests/test_config.py: test_defaults depended on the absence of a local
.env — load_config() calls load_dotenv(override=False), which restored
the variable the test had just deleted. Stub dotenv discovery.
305 tests pass.
- Subagents: fold Task-tool transcripts (subagents/*.jsonl) inline as
collapsible <details> blocks at the spawn point; their own tool traffic
collapses under the same policy; nesting handled via toolUseId matching
- Joplin: Claude Code gets its own top-level AI-ClaudeCode notebook;
update_note sets parent_id so notes self-heal/relocate on re-sync
- Repo [tags] in titles via git-root detection (nearest .git ancestor),
home-wide and cross-workspace; CLAUDE_CODE_REPO_TAG_IGNORE escape hatch
- Multi-root scanning: CLAUDE_CODE_DIR as os.pathsep list + CLAUDE_CONFIG_DIR;
merged by launch-folder, newer-mtime wins on UUID collision
- 298 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Mark FUTURE.md as feature-complete as of v0.7.0: active roadmap is empty and
the deprioritized backlog (Joplin --force, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reclassification, Obsidian output,
token-expiry notifications, search) is closed as not needed, kept for
reference only. No further work planned.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The export reads ChatGPT's and Claude's undocumented internal web APIs,
which can change shape without notice; the worst failure for a backup tool
is a silent one (skipped/mis-parsed content with no error). Add a `canary`
command + BaseProvider.check_drift() (overridden by ChatGPT/Claude) that
fetches one listing page + one conversation per provider and asserts only
the normalizer's load-bearing fields — not the full response shape, which
churns harmlessly. The top silent risk it guards is a renamed retrieval-tool
author bypassing the hidden-content collapse. ERROR findings exit non-zero;
WARN findings are surfaced but non-fatal so a backup run is never blocked.
Also retires the in-app watch mode and headless StartOS direction from the
roadmap (tool stays a local, manually-run CLI) and updates docs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ChatGPT session tokens are JWEs whose `exp` is encrypted and unreadable
client-side, so the old JWT-decode path could never yield an expiry, and the
`/api/auth/session` `expires` is a misleading rolling window (it advances on
every call even for a dead token). The honest signal is the `error` field:
`RefreshAccessTokenError` means the session token is dead while `expires` and
a stale `accessToken` are still echoed (verified live).
- Add ChatGPTProvider.session_health() and a "ChatGPT token active" doctor
check based on it; drop the dead JWE/exp decode path in doctor and auth.
- Fix _fetch_access_token to fail fast on a set `error` instead of returning
the stale accessToken (which produced a confusing downstream 401).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Browser cookie auto-extraction is not viable: modern Chromium App-Bound
Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level
layer that cannot be decrypted off disk without admin rights and AV-flagged
SYSTEM impersonation, and fails on Brave specifically. Drop the
`--from-browser` flag, `_auth_from_browser`, `src/browser_tokens.py`, its
test, and the browser-cookie3 dependency. Auth is manual (DevTools wizard)
only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
mark_exported() was discarding created_at from the metadata dict because
it wasn't in the hardcoded stored-key list, so the joplin sync always
saw an empty date and omitted the prefix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Joplin notebooks now use a two-level hierarchy: AI-ChatGPT / <project> and
AI-Claude / <project> instead of a single flat title. Note titles are prefixed
with the conversation created_at date (YYYY-MM-DD). Export folders collapse
provider/project/year into a single provider/project.year directory.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
First real-data export against v0.4.0 surfaced 66 unknown blocks across
three content types — captured live and added.
Added:
- execution_output (Code Interpreter / container.exec / python tool
output) → tool_result block. output=content.text,
tool_name=author.name, is_error=metadata.aggregate_result.status,
summary=metadata.reasoning_title
- system_error → error tool_result with tool_name=author.name
- tether_browsing_display: spinner placeholders (empty result+summary)
skip silently with DEBUG log; defensive populated-case branch maps
to tool_result (untested in real data)
- tool_result block schema: optional `summary` field rendered as
italic line between header and fence
- tool_result rendering: tool_name appears in header when present
(e.g. `📤 Result: container.exec`); existing tool_name=None calls
unchanged
- _ROLE_LABELS["tool"] = ("🔧 Tool", "tool")
Fixed:
- chatgpt.normalize_conversation reads `conversation_id` as fallback
for `id`. Live API uses conversation_id; fixtures use id.
Pre-fix: empty id in YAML frontmatter and missing context in
WARNING logs.
Tests: 11 new (192 total, 0 failures). Fixture extended with 4
tool-output cases (execution_output success, empty execution_output
that should skip, system_error, tether_browsing_display spinner).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Extracts per-message content into a typed `blocks` list (text, code,
thinking, tool_use, tool_result, image_placeholder, file_placeholder,
unknown) and renders them at exporter write time. Voice transcripts,
Custom Instructions, and image references now appear in exports
instead of being silently dropped.
Foundation:
- src/blocks.py: pure block constructors, _safe_fence (fence-corruption
defense, verified live in Joplin), _blockquote_prefix, render
- src/loss_report.py: per-run tally surfaced as INFO summary at end of
export so silently-dropped data becomes visible
Providers:
- ChatGPT: dispatch on content_type produces typed blocks; voice shapes
(audio_transcription, audio_asset_pointer, real_time_user_audio_video_
asset_pointer) locked from live DevTools capture; Custom Instructions
bug fix (parts-vs-direct-fields); role filter lifted; hidden-context
marker driven by is_visually_hidden_from_conversation flag
- Claude: defensive dispatch for text/thinking/tool_use/tool_result/image
with recursive nested-block flattening; untested against real rich-
content data — fix-forward in v0.4.1
Exporter:
- Markdown renders from blocks at write time via render_blocks_to_markdown;
backward-compat fallback to content for any pre-v0.4.0 cached data
Tests:
- 27 new tests across providers, exporters, CLI; fixtures rebuilt with
real-shape ChatGPT voice + Custom Instructions cases
- 181/181 pass
Behavior changes (intentional):
- JSON output omits content; consumers should read blocks
- Per-conversation message counts increase (Custom Instructions, image-
only, tool-only messages now appear)
- Existing exports not auto-re-rendered; users wanting fresh output run
cache --clear then export
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Support __Secure-next-auth.session-token.0/.1 split cookies; ChatGPT
now issues tokens that exceed the 4KB per-cookie limit and must be
sent as two named chunks or the auth endpoint returns no accessToken.
Add CHATGPT_SESSION_TOKEN_1 env var; update auth wizard instructions.
- Fix Claude conversations exported to wrong directory when project name
is present in the listing but absent from the detail endpoint response.
Explicitly propagate "project" alongside _-prefixed annotation keys.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude's list endpoint returns conversations with a `name` field rather
than `title`, so every Claude row was falling through to "Untitled".
Also set no_wrap + ellipsis overflow and tune column widths so the table
renders one row per conversation in Windows Command Prompt (80 cols).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Core features:
- Add `joplin` command: syncs exported Markdown to Joplin via local REST API
- Notebooks auto-created per provider+project (e.g. "ChatGPT - My Project")
- Idempotent: notes updated (not duplicated) on re-run; note ID tracked in manifest
- Add `--project` filter to `export` and `list` commands (substring or 'none')
- Add ChatGPT Projects support via CHATGPT_PROJECT_IDS env var
Config:
- Add JOPLIN_API_TOKEN, JOPLIN_API_URL, JOPLIN_REQUEST_TIMEOUT
- Version now read from importlib.metadata (single source of truth: pyproject.toml)
- Bump version to 0.2.0
Quality:
- Explicit Timeout handling in JoplinClient with actionable error messages
- token validation (validate_token) separate from connectivity (ping)
- Remove debug_auth.py, debug_claude.py, and untracked .har file
- Add *.har to .gitignore (may contain auth cookies/session tokens)
- Update README, CHANGELOG, FUTURE.md to reflect v0.2.0
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
claude.ai has the same Cloudflare TLS fingerprinting protection as
chatgpt.com. Apply the same fix: curl_cffi impersonate=chrome120,
remove base class User-Agent to avoid JA3/UA mismatch.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
curl_cffi sets a User-Agent consistent with its JA3 TLS fingerprint.
BaseProvider's custom UA (Chrome/121) conflicted with the chrome120
TLS fingerprint, causing Cloudflare to flag the request as a bot.
Removing the UA from session headers lets curl_cffi manage its own.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
chatgpt.com uses Cloudflare's TLS fingerprinting (JA3/JA4) which
blocks Python requests regardless of cookies. curl_cffi impersonates
Chrome's exact TLS handshake, making requests indistinguishable from
a real browser at the transport layer.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Using self._session.cookies.set() ensures the cookie is sent correctly
by the requests session on all calls, including /api/auth/session.
Also add sec-fetch-* headers required by chatgpt.com.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The __Secure-next-auth.session-token cannot be used directly as a Bearer
token. It must first be exchanged via GET /api/auth/session (with the token
sent as a Cookie) to obtain a short-lived accessToken. This accessToken is
then used as the Authorization: Bearer header for all backend-api calls.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Doctor was reading env vars before loading .env, so tokens set in .env
were invisible. ChatGPT now uses JWE (encrypted JWT) tokens which
PyJWT cannot decode without the server key — treat decode failure as
"token set, expiry unknown" rather than a FAIL.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>