sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
A fork subagent's transcript opens with a copy of the parent turn that spawned it, its own Agent call included; folding that call re-entered the same fork until RecursionError, failing every daily sync since 2026-09-24. Skip a spawn call while its own subagent is being expanded, and strip the <fork-boilerplate> preamble.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
claude.ai answers an invalid or expired sessionKey with 403 account_session_invalid, never 401, so the refresh-your-cookie message could not fire for Claude. Auth detection is now a provider decision (_is_auth_failure); Claude matches the 403 on its error code so a real permission error still reports as itself. README and .env.example no longer claim both ChatGPT cookie chunks are required.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
FUTURE.md had become a 648-line archaeological record: six shipped roadmap
items, two dropped, two implemented, an investigation trail for each, and a
backlog closed as not needed — with the two genuinely planned items buried
at the bottom. It is now 139 lines, planned work first.
- Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon
behind decisions now recorded in one line each — why Brave cookie
extraction is not viable, why the ChatGPT token's expiry cannot be read
client-side, how the drift canary was designed — which would be expensive
to rediscover.
- Roadmap is now two items: the StartOS service (including the 2026-08-18
decision that clients upload to StartOS storage and the service owns the
Joplin connection), and the README split.
- Everything shipped or dropped is off the roadmap. Decisions not to build
survive as a one-line table so they are not re-proposed.
- Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0.
Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased]
section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the
launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy
notifications, gizmo_id project attribution and the `projects` command, and
the splitlines data-loss fix in both local providers.
Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose —
exported conversations are Markdown and may contain private content — with
each doc re-included by name, so a new doc is silently untracked rather than
rejected. The archive would not have been committed at all. Recorded in the
README-split roadmap item, since `docs/*.md` will hit exactly this.
Also refreshes the package description, which still named only ChatGPT and
Claude after two local providers were added.
Carries the documentation and table-of-contents work from earlier today.
Reading gizmo_id during normalization fixed attribution, but it cannot
answer "which projects am I missing?": normalization only runs on
conversations being exported, and a normal run skips everything already
cached. Discovering the gaps would have meant --force re-exporting the
whole archive.
`ai-chat-exporter projects` does it directly. It lists conversations,
collects the g-p- ids they belong to, resolves display names, and prints a
table marking which are absent from .env plus a paste-ready
CHATGPT_PROJECT_IDS line. --write applies it; --deep falls back to one
detail request per conversation when the listing does not carry gizmo_id
(unverified which shape this account returns, so the command reports which
path it took rather than assuming).
This matters beyond tidiness: attribution is now self-correcting, but the
listing pass still needs the ids. Conversations that live only inside a
project never appear in the default listing, so an unlisted project's chats
are not merely misfiled — they are never fetched.
6 CLI tests: reporting, the already-configured case, the g-p- guard, --deep,
the hint when --deep is needed, and --write. 324 pass.
A conversation in a project absent from CHATGPT_PROJECT_IDS exported into
no-project/ even though its own payload names the project. Found while
investigating the media 403s: bpi-f3-case-options sits in
g-p-6a4edf5160848191927b05da20f49151, which is not among the 13 configured
ids, so it filed under no-project.2026.
Both existing sources are bounded by what the user configured. The detail
response is not: it carries gizmo_id. Use it as a third fallback, after the
listing annotation and the project map, and cache the result into the map.
Attribution now stays correct with no list to maintain, and moving a chat
into a new project stops silently misfiling it.
Only g-p- ids are treated as projects — a custom GPT is not a project and
must not become a folder.
CHATGPT_PROJECT_IDS still matters for the listing pass: conversations that
live only inside a project never appear in the default listing, so an
unconfigured project's chats can be missed entirely. Each one is now
reported once per run with the id to add, which turns "some chats are
missing" into a line to paste.
8 tests cover precedence, the g-p- guard, map caching and the once-per-run
report. 318 pass.
The render check settled it. Scrolling the whole conversation found sets of
images that do NOT display in ChatGPT — a set of 6 and a set of 4 — and they
are exactly our failures: the 6 are the records that 404 outright, the 4 are
the refused attachment batch we dumped (image.png, image(1).png,
image(2).png + 1). Every set that displays downloaded fine.
So the 403 was never OpenAI withholding something. It is the same failure
their own UI hits. All 19 are unrecoverable.
What the investigation established, now recorded in media.py and the
changelog so it is not repeated:
- uploads do not expire (36/36 sampled, 2025-09 → 2026-08, still live), so
export cadence was never the variable
- 7 records 404; 12 report state=ready with a library_file_id
- /files/{id}/download mints the signed estuary/content URL the UI fetches,
and refuses the survivors regardless of headers, Authorization, gizmo_id,
conversation_id, Referer or namespace; the Library id is file_not_found
- all survivors were created 2026-07-14 within minutes of each other yet
appear in conversations predating that date — a Library migration that
kept the metadata and lost the bytes
Dropping the nine one-off probes; their findings live in the code comments
and changelog, and git history has the scripts if they are ever wanted.
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
( when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.
Offline, no token, no API calls.
Media downloads logged "HTTP Error 403:" with no reason. That string is
curl_cffi's raise_for_status() format, "HTTP Error {code}: {reason}", and
HTTP/2 carries no reason phrase — so the message said nothing, and
_make_request threw the response body away. The provider's JSON `detail`
is the only explanation available for a refused asset.
- base._make_request: end non-retryable statuses with a ProviderError
carrying the body's detail/error/message (redacted, truncated to 300
chars) instead of a bare raise_for_status().
- media: bucket 403 as `forbidden` in the run summary, separately from
`download-error` — "the asset is gone" and "we were refused" are
different problems.
- utils.redact_secrets: match secret key names per word. Exact matching
let access_token, api_key, and session-token through into logged
bodies; "keywords"/"monkey"/"tokenizer" stay intact.
- tests/test_config.py: test_defaults depended on the absence of a local
.env — load_config() calls load_dotenv(override=False), which restored
the variable the test had just deleted. Stub dotenv discovery.
305 tests pass.
- Subagents: fold Task-tool transcripts (subagents/*.jsonl) inline as
collapsible <details> blocks at the spawn point; their own tool traffic
collapses under the same policy; nesting handled via toolUseId matching
- Joplin: Claude Code gets its own top-level AI-ClaudeCode notebook;
update_note sets parent_id so notes self-heal/relocate on re-sync
- Repo [tags] in titles via git-root detection (nearest .git ancestor),
home-wide and cross-workspace; CLAUDE_CODE_REPO_TAG_IGNORE escape hatch
- Multi-root scanning: CLAUDE_CODE_DIR as os.pathsep list + CLAUDE_CONFIG_DIR;
merged by launch-folder, newer-mtime wins on UUID collision
- 298 tests passing
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The export reads ChatGPT's and Claude's undocumented internal web APIs,
which can change shape without notice; the worst failure for a backup tool
is a silent one (skipped/mis-parsed content with no error). Add a `canary`
command + BaseProvider.check_drift() (overridden by ChatGPT/Claude) that
fetches one listing page + one conversation per provider and asserts only
the normalizer's load-bearing fields — not the full response shape, which
churns harmlessly. The top silent risk it guards is a renamed retrieval-tool
author bypassing the hidden-content collapse. ERROR findings exit non-zero;
WARN findings are surfaced but non-fatal so a backup run is never blocked.
Also retires the in-app watch mode and headless StartOS direction from the
roadmap (tool stays a local, manually-run CLI) and updates docs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
First real-data export against v0.4.0 surfaced 66 unknown blocks across
three content types — captured live and added.
Added:
- execution_output (Code Interpreter / container.exec / python tool
output) → tool_result block. output=content.text,
tool_name=author.name, is_error=metadata.aggregate_result.status,
summary=metadata.reasoning_title
- system_error → error tool_result with tool_name=author.name
- tether_browsing_display: spinner placeholders (empty result+summary)
skip silently with DEBUG log; defensive populated-case branch maps
to tool_result (untested in real data)
- tool_result block schema: optional `summary` field rendered as
italic line between header and fence
- tool_result rendering: tool_name appears in header when present
(e.g. `📤 Result: container.exec`); existing tool_name=None calls
unchanged
- _ROLE_LABELS["tool"] = ("🔧 Tool", "tool")
Fixed:
- chatgpt.normalize_conversation reads `conversation_id` as fallback
for `id`. Live API uses conversation_id; fixtures use id.
Pre-fix: empty id in YAML frontmatter and missing context in
WARNING logs.
Tests: 11 new (192 total, 0 failures). Fixture extended with 4
tool-output cases (execution_output success, empty execution_output
that should skip, system_error, tether_browsing_display spinner).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Extracts per-message content into a typed `blocks` list (text, code,
thinking, tool_use, tool_result, image_placeholder, file_placeholder,
unknown) and renders them at exporter write time. Voice transcripts,
Custom Instructions, and image references now appear in exports
instead of being silently dropped.
Foundation:
- src/blocks.py: pure block constructors, _safe_fence (fence-corruption
defense, verified live in Joplin), _blockquote_prefix, render
- src/loss_report.py: per-run tally surfaced as INFO summary at end of
export so silently-dropped data becomes visible
Providers:
- ChatGPT: dispatch on content_type produces typed blocks; voice shapes
(audio_transcription, audio_asset_pointer, real_time_user_audio_video_
asset_pointer) locked from live DevTools capture; Custom Instructions
bug fix (parts-vs-direct-fields); role filter lifted; hidden-context
marker driven by is_visually_hidden_from_conversation flag
- Claude: defensive dispatch for text/thinking/tool_use/tool_result/image
with recursive nested-block flattening; untested against real rich-
content data — fix-forward in v0.4.1
Exporter:
- Markdown renders from blocks at write time via render_blocks_to_markdown;
backward-compat fallback to content for any pre-v0.4.0 cached data
Tests:
- 27 new tests across providers, exporters, CLI; fixtures rebuilt with
real-shape ChatGPT voice + Custom Instructions cases
- 181/181 pass
Behavior changes (intentional):
- JSON output omits content; consumers should read blocks
- Per-conversation message counts increase (Custom Instructions, image-
only, tool-only messages now appear)
- Existing exports not auto-re-rendered; users wanting fresh output run
cache --clear then export
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Core features:
- Add `joplin` command: syncs exported Markdown to Joplin via local REST API
- Notebooks auto-created per provider+project (e.g. "ChatGPT - My Project")
- Idempotent: notes updated (not duplicated) on re-run; note ID tracked in manifest
- Add `--project` filter to `export` and `list` commands (substring or 'none')
- Add ChatGPT Projects support via CHATGPT_PROJECT_IDS env var
Config:
- Add JOPLIN_API_TOKEN, JOPLIN_API_URL, JOPLIN_REQUEST_TIMEOUT
- Version now read from importlib.metadata (single source of truth: pyproject.toml)
- Bump version to 0.2.0
Quality:
- Explicit Timeout handling in JoplinClient with actionable error messages
- token validation (validate_token) separate from connectivity (ping)
- Remove debug_auth.py, debug_claude.py, and untracked .har file
- Add *.har to .gitignore (may contain auth cookies/session tokens)
- Update README, CHANGELOG, FUTURE.md to reflect v0.2.0
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>