38 Commits
Author SHA1 Message Date
JesseMarkowitzandClaude Opus 5.5 23c6e1f512 fix: push a FAILED notification when a scheduled run dies without reporting
sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
2026-10-05 07:14:16 -04:00
JesseMarkowitzandClaude Opus 5.5 64068bb19b fix: fork subagents recursed forever in the Claude Code export
A fork subagent's transcript opens with a copy of the parent turn that spawned it, its own Agent call included; folding that call re-entered the same fork until RecursionError, failing every daily sync since 2026-09-24. Skip a spawn call while its own subagent is being expanded, and strip the <fork-boilerplate> preamble.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
2026-10-05 07:14:16 -04:00
JesseMarkowitzandClaude Opus 5.5 b6ce636891 fix: detect claude.ai's 403 session-invalid as an expired key; ChatGPT .1 cookie is optional
claude.ai answers an invalid or expired sessionKey with 403 account_session_invalid, never 401, so the refresh-your-cookie message could not fire for Claude. Auth detection is now a provider decision (_is_auth_failure); Claude matches the 403 on its error code so a real permission error still reports as itself. README and .env.example no longer claim both ChatGPT cookie chunks are required.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
2026-10-05 07:14:01 -04:00
JesseMarkowitz d3745e1de4 fix: test suite sent real push notifications via the developer's .env 2026-08-18 14:26:05 -04:00
JesseMarkowitz 2d5fcb26f5 release: v0.9.0, and rewrite FUTURE.md around what is actually planned
FUTURE.md had become a 648-line archaeological record: six shipped roadmap
items, two dropped, two implemented, an investigation trail for each, and a
backlog closed as not needed — with the two genuinely planned items buried
at the bottom. It is now 139 lines, planned work first.

- Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon
  behind decisions now recorded in one line each — why Brave cookie
  extraction is not viable, why the ChatGPT token's expiry cannot be read
  client-side, how the drift canary was designed — which would be expensive
  to rediscover.

- Roadmap is now two items: the StartOS service (including the 2026-08-18
  decision that clients upload to StartOS storage and the service owns the
  Joplin connection), and the README split.

- Everything shipped or dropped is off the roadmap. Decisions not to build
  survive as a one-line table so they are not re-proposed.

- Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0.

Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased]
section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the
launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy
notifications, gizmo_id project attribution and the `projects` command, and
the splitlines data-loss fix in both local providers.

Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose —
exported conversations are Markdown and may contain private content — with
each doc re-included by name, so a new doc is silently untracked rather than
rejected. The archive would not have been committed at all. Recorded in the
README-split roadmap item, since `docs/*.md` will hit exactly this.

Also refreshes the package description, which still named only ChatGPT and
Claude after two local providers were added.

Carries the documentation and table-of-contents work from earlier today.
2026-08-18 13:51:34 -04:00
JesseMarkowitz 55b9ce12f6 added ntfy support for scheduled runs 2026-08-18 11:21:26 -04:00
JesseMarkowitz b0aa05a2b1 fix sync bug for scheduled work 2026-08-18 10:31:28 -04:00
JesseMarkowitz 999429f61e add alias to avoid explicit python commands, scheduled runs and bug fixes 2026-08-18 09:59:15 -04:00
JesseMarkowitz fe5ed341ad added support for codex provider. fixed bug in claude code export (lines split when shouldn't be) 2026-08-18 08:01:34 -04:00
JesseMarkowitz d083aea135 feat: projects command — find the project IDs the config is missing
Reading gizmo_id during normalization fixed attribution, but it cannot
answer "which projects am I missing?": normalization only runs on
conversations being exported, and a normal run skips everything already
cached. Discovering the gaps would have meant --force re-exporting the
whole archive.

`ai-chat-exporter projects` does it directly. It lists conversations,
collects the g-p- ids they belong to, resolves display names, and prints a
table marking which are absent from .env plus a paste-ready
CHATGPT_PROJECT_IDS line. --write applies it; --deep falls back to one
detail request per conversation when the listing does not carry gizmo_id
(unverified which shape this account returns, so the command reports which
path it took rather than assuming).

This matters beyond tidiness: attribution is now self-correcting, but the
listing pass still needs the ids. Conversations that live only inside a
project never appear in the default listing, so an unlisted project's chats
are not merely misfiled — they are never fetched.

6 CLI tests: reporting, the already-configured case, the g-p- guard, --deep,
the hint when --deep is needed, and --write. 324 pass.
2026-08-17 13:19:34 -04:00
JesseMarkowitz 710889b65f fix: read project attribution from the conversation's gizmo_id
A conversation in a project absent from CHATGPT_PROJECT_IDS exported into
no-project/ even though its own payload names the project. Found while
investigating the media 403s: bpi-f3-case-options sits in
g-p-6a4edf5160848191927b05da20f49151, which is not among the 13 configured
ids, so it filed under no-project.2026.

Both existing sources are bounded by what the user configured. The detail
response is not: it carries gizmo_id. Use it as a third fallback, after the
listing annotation and the project map, and cache the result into the map.
Attribution now stays correct with no list to maintain, and moving a chat
into a new project stops silently misfiling it.

Only g-p- ids are treated as projects — a custom GPT is not a project and
must not become a folder.

CHATGPT_PROJECT_IDS still matters for the listing pass: conversations that
live only inside a project never appear in the default listing, so an
unconfigured project's chats can be missed entirely. Each one is now
reported once per run with the id to add, which turns "some chats are
missing" into a line to paste.

8 tests cover precedence, the g-p- guard, map caching and the once-per-run
report. 318 pass.
2026-08-17 12:49:04 -04:00
JesseMarkowitz dfa0645fba close the media 403 investigation: the bytes are gone
The render check settled it. Scrolling the whole conversation found sets of
images that do NOT display in ChatGPT — a set of 6 and a set of 4 — and they
are exactly our failures: the 6 are the records that 404 outright, the 4 are
the refused attachment batch we dumped (image.png, image(1).png,
image(2).png + 1). Every set that displays downloaded fine.

So the 403 was never OpenAI withholding something. It is the same failure
their own UI hits. All 19 are unrecoverable.

What the investigation established, now recorded in media.py and the
changelog so it is not repeated:
  - uploads do not expire (36/36 sampled, 2025-09 → 2026-08, still live), so
    export cadence was never the variable
  - 7 records 404; 12 report state=ready with a library_file_id
  - /files/{id}/download mints the signed estuary/content URL the UI fetches,
    and refuses the survivors regardless of headers, Authorization, gizmo_id,
    conversation_id, Referer or namespace; the Library id is file_not_found
  - all survivors were created 2026-07-14 within minutes of each other yet
    appear in conversations predating that date — a Library migration that
    kept the metadata and lost the bytes

Dropping the nine one-off probes; their findings live in the code comments
and changelog, and git history has the scripts if they are ever wanted.
2026-08-17 12:47:05 -04:00
JesseMarkowitz d30a9510bb tools: find what the browser sends that we don't
The decisive fact arrived: the images DO render in ChatGPT. So a valid
signature exists for files that hand us 403, and there is a route to find.

The rest is now pinned down. /files/{id}/download is the minting endpoint —
a working file's download_url is the same estuary/content URL the browser
uses. Signatures are per-file: swapping an id into a working URL returns
"Invalid signature or expired URL", and a ts without a sig gets the same.
So the difference is in the request, not the route.

Our call sends Authorization: Bearer plus cookies. A browser leans on
cookies and adds client headers a gated endpoint may check. This walks
those variants — no Authorization, conversation Referer, oai-device-id,
oai-language, sec-fetch-*, an image Accept, a conversation_id param —
against a file that is currently refused, and reports which mints a URL.

If none do, it prints the DevTools recipe for finding the minting call by
searching the Network panel for the file id.
2026-08-17 11:07:18 -04:00
JesseMarkowitz b1da1df986 tools: probe a refused file, not a deleted one
Confirmed by the last run: a working file's download_url is
chatgpt.com/backend-api/estuary/content?cid&id&p&sig&ts&v — the same route
the browser uses. So /files/{id}/download is the minting endpoint, and a
gizmo file's 403 is a refusal to mint the signature. That is why no amount
of scoping helped; we were turned away at the only door that issues them.
Nothing else in the estuary namespace serves files: 404 across the board.

But section B tested file_000000001b3071f5…, one of the seven DELETED
files, because it took the first failure without checking its class. Those
results say nothing about the eleven recoverable ones.

- Pick the target by probing /files/{id} and taking one that answers 200,
  skipping (and naming) the deleted ones.
- Add two experiments: estuary/content with a ts but no sig, since
  validation asked only for id/p/ts and never sig; and a working file's
  freshly minted URL with the refused id swapped in, which shows whether
  the signature is bound to the file.
2026-08-17 10:57:52 -04:00
JesseMarkowitz 726f57bdf9 tools: probe the estuary namespace
A URL copied from the browser gave us the route we never tried:

    /backend-api/estuary/content?id=…&ts=496382&p=fs&cid=1&sig=…&v=0

Fetched with no cookies it returns 403 {"detail":"File stream access
denied."}, so the signature rides on the session rather than replacing it.
Every earlier probe lived under /backend-api/files/*; estuary/* is new
ground.

estuary_probe.py checks two things:

A. What /files/{id}/download hands back for a file that works. If its
   download_url is an estuary URL, that endpoint is the minting step, and a
   gizmo file's 403 is a refusal to mint — which is why no amount of
   scoping helped.
B. Whether the estuary namespace exposes a route that serves a refused
   file directly.

It also replays a pasted URL through the exporter's session, and says
plainly whether the id in that URL is one of the refused files — the two
captured so far were working images, so they showed the shape without
telling us whether the broken ones have a URL at all.
2026-08-17 10:50:23 -04:00
JesseMarkowitz 024bfde030 tools: test whether a browser image URL works from the exporter
The endpoint hunt is done guessing. Results:

  /files/{libfile}/download    200 {"error_code":"file_not_found",
                                    "error_type":"GetDownloadLinkError"}
                               — twice, on two different Library ids, so the
                                 Library id is simply not valid there
  /library/*                   404 across every shape
  POST on the download route   405
  /content?asset_pointer=      422 requiring query params id, ts, p

That last one is the find: /backend-api/content is a signed route, and the
web UI has the values we cannot compute. So the question shifts from "which
endpoint" to "is the UI's URL reusable outside the browser".

try_pasted_url.py takes a URL copied from DevTools and fetches it three
ways — bare, through the exporter's authenticated session, and with the
Authorization header removed in case bearer and signature conflict. The
pattern says whether the exporter can fetch these at all, and therefore
whether it is worth hunting for what mints the signature.

Before any of that: check whether the images still render in ChatGPT's own
UI. If they show broken there, file_not_found is literally true, these 11
are lost like the other 7, and no route exists to find.
2026-08-17 10:27:57 -04:00
JesseMarkowitz ffd01ebcf3 tools: print the error body the Library endpoint returns
The pairing came back perfectly clean, and it is the whole diagnosis:

    has library_file_id  → refused (11/11)
    no  library_file_id  → deleted (7/7)

And /files/{libfile_id}/download answered 200 — not 403, not 404 — with
{status, error_code, error_type, error_message}. That endpoint accepts the
Library id; it is returning an application-level error inside an HTTP 200.
The probe printed only the keys and dropped the message, which is the one
thing that says what the call is missing.

- show() now prints the full body for every response, and detects a
  non-JSON 200 as possible raw bytes.
- Added routes worth ruling in or out: the same call as POST, with a
  conversation_id, /files/{lib}/content, /content?asset_pointer=, and four
  Library listing endpoints — a listing usually reveals both the id form
  the UI uses and the route that actually serves bytes.
- Probes a second Library id too, so one odd file cannot mislead us.
2026-08-17 10:14:18 -04:00
JesseMarkowitz 3f82b35a55 tools: try fetching refused images by their Library ID
Dumping the raw message metadata found the identity the asset pointer never
carried:

    "id":              "file_000000003454722f9481506b96aed510"   ← refused
    "library_file_id": "libfile_4eb82f478fe081919127e2eba9886e86"
    "source":          "local"

The exporter only knows the sediment id from the asset_pointer and asks
/files/{sediment}/download, which 403s for these. The Library is a separate
store with its own ids, so we have been asking for the conversation-scoped
copy of a file that now lives in the Library. That also fits the 2026-07-14
creation-time cluster: a Library migration would mint exactly these new
records.

Neither gizmo_id, conversation_id, a /gizmos path nor a project Referer
helped (403/404), so scope was never the missing piece — identity was.

library_download_probe.py pairs each refused sediment id with its
library_file_id from message.metadata.attachments and tries the endpoints
that could serve it.
2026-08-17 09:33:05 -04:00
JesseMarkowitz 03646009b9 tools: probe how to download a gizmo-scoped file
metadata_diff found the discriminator. Every refused image carries
use_case="gizmo" — ChatGPT's name for Projects and custom GPTs — while
everything that downloads is image_gen or multimodal. All are state=ready,
so nothing is damaged: the plain /files/{id}/download endpoint just will
not serve a project-scoped file.

Second clue: all 11 were created 2026-07-14 within minutes of each other,
yet appear in conversations dated 2025-11 through 2026-08, several
predating their own creation_time. Something on July 14 re-created them as
gizmo-scoped copies and repointed the conversations — which is also why
the losses start in July. Not a policy change, an event.

gizmo_download_probe.py dumps the raw conversation part carrying one of
these assets (it may simply name the scope the download wants) and then
tries the plausible calls — gizmo_id/conversation_id query params, a
/gizmos/{id}/files path, a project Referer. A 200 is the fix.
2026-08-17 09:22:48 -04:00
JesseMarkowitz a3ac279e39 tools: find what separates a refused image from a served one
branch_check killed the abandoned-branch hypothesis — every lost image is
on the live branch, and the 3 images that do sit on abandoned branches are
alive. But it turned up something the earlier probing missed by sampling
four IDs from one family and generalising: of the 19 failures, only 7 are
actually gone. The other 12 answer /files/{id} with 200 and full metadata
and refuse only /download. They exist, and may be recoverable.

Three states, then: gone (404), refused (200 meta + 403 download), working
(200 both). Since metadata comes back for the refused ones, the
discriminator can be read straight off — fetch it for every file in each
state and compare fields, flagging any field whose values never overlap
between states.

Also re-checks /download now: if a file refused during the export serves
today, those 403s were transient and a retry pass recovers them, which is
a completely different fix from anything permanent.
2026-08-17 09:18:32 -04:00
JesseMarkowitz 021a76c628 tools: test whether lost images sit on abandoned branches
Expiry is now ruled out: 36/36 sampled images from 2025-09 through
2026-08 are still live server-side, so nothing dies of age and export
cadence is not the variable. The loss is per-asset — 2026-07-09 kept 38
images and lost 12 in one conversation.

Next hypothesis: those images belong to messages that were edited or
regenerated. ChatGPT stores a conversation as a tree and editing forks it,
stranding the superseded messages on an abandoned branch. The exporter
walks every node (chatgpt.py:963), so it exports those branches too — and
an attachment unreachable from live history is a natural GC target.

branch_check.py fetches the raw conversation, derives the live branch by
walking parent links up from current_node, and cross-tabulates every image
against (on the live branch?) x (still downloadable?). If the lost images
are all off-branch and nothing on the live branch is missing, this is not
data loss at all — it is attachments to messages that were replaced.
2026-08-17 08:58:20 -04:00
JesseMarkowitz b9e8896b33 tools: distinguish "captured in time" from "still alive"
analyze_media_age.py claimed age was ruled out because an image from
2025-11 was saved while one from 2026-08 was not. That conclusion does not
follow. "Saved" means some earlier run downloaded it, not that ChatGPT
still holds it — an image captured in June looks saved forever after, even
if it died in July. A month with no losses shows the exports were timely,
not that the assets survived.

- probe_survival_by_month.py: sample images already on disk, grouped by
  their conversation's month, and ask /files/{id} whether each still
  exists today. Old months still 200 → no expiry, and cadence did not save
  them. Old months now 404 → uploads do expire and cadence is the whole
  ballgame.
- analyze_media_age.py: stop asserting the unsupported verdict; say what
  the number does and does not show, and point at the probe.

One signal there is immune to the confound and survives: 2026-07-09 kept
38 images and lost 12. Same conversation, same day, opposite outcomes —
no retention policy does that, so at least part of this is per-asset.
2026-08-17 08:53:54 -04:00
JesseMarkowitz 04191eed8c tools: analyze whether media loss is age-based
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
(![source](media/…) when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.

Offline, no token, no API calls.
2026-08-17 08:45:21 -04:00
JesseMarkowitz f40b25001a tools: add one-off probe for the media 403s
The improved error reporting landed, and the answer it produced is
{"detail":"Forbidden"} — generic, no reason. The cause has to be narrowed
by experiment instead, so collect the experiments in one script:

- classify the failed assets offline from the exported placeholders
  (user_upload vs model_generated) — no API call needed
- probe a known-good asset in the same session as a control, to rule the
  session in or out
- retry the 403 with ChatGPT-Account-Id, which the exporter never sends
  and which workspace-scoped resources can require
- compare /files/{id} against /files/{id}/download

Temporary: delete once the cause is known, or fold into `doctor` if the
check earns a permanent home.
2026-08-17 08:08:01 -04:00
JesseMarkowitz 395ea19ca8 fix: surface the response body on 4xx so 403s are diagnosable
Media downloads logged "HTTP Error 403:" with no reason. That string is
curl_cffi's raise_for_status() format, "HTTP Error {code}: {reason}", and
HTTP/2 carries no reason phrase — so the message said nothing, and
_make_request threw the response body away. The provider's JSON `detail`
is the only explanation available for a refused asset.

- base._make_request: end non-retryable statuses with a ProviderError
  carrying the body's detail/error/message (redacted, truncated to 300
  chars) instead of a bare raise_for_status().
- media: bucket 403 as `forbidden` in the run summary, separately from
  `download-error` — "the asset is gone" and "we were refused" are
  different problems.
- utils.redact_secrets: match secret key names per word. Exact matching
  let access_token, api_key, and session-token through into logged
  bodies; "keywords"/"monkey"/"tokenizer" stay intact.
- tests/test_config.py: test_defaults depended on the absence of a local
  .env — load_config() calls load_dotenv(override=False), which restored
  the variable the test had just deleted. Stub dotenv discovery.

305 tests pass.
2026-08-17 07:58:02 -04:00
JesseMarkowitzandClaude Opus 4.8 1f5a445ada feat: v0.8.0 — Claude Code subagent capture, own Joplin notebook, git-root repo tags, multi-root scanning
- Subagents: fold Task-tool transcripts (subagents/*.jsonl) inline as
  collapsible <details> blocks at the spawn point; their own tool traffic
  collapses under the same policy; nesting handled via toolUseId matching
- Joplin: Claude Code gets its own top-level AI-ClaudeCode notebook;
  update_note sets parent_id so notes self-heal/relocate on re-sync
- Repo [tags] in titles via git-root detection (nearest .git ancestor),
  home-wide and cross-workspace; CLAUDE_CODE_REPO_TAG_IGNORE escape hatch
- Multi-root scanning: CLAUDE_CODE_DIR as os.pathsep list + CLAUDE_CONFIG_DIR;
  merged by launch-folder, newer-mtime wins on UUID collision
- 298 tests passing

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 05:04:41 -04:00
JesseMarkowitzandClaude Opus 4.8 bbcb29c856 docs: close backlog — project feature-complete for now
Mark FUTURE.md as feature-complete as of v0.7.0: active roadmap is empty and
the deprioritized backlog (Joplin --force, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reclassification, Obsidian output,
token-expiry notifications, search) is closed as not needed, kept for
reference only. No further work planned.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:51:05 -04:00
JesseMarkowitzandClaude Opus 4.8 1e016ea652 canary: detect provider API schema drift; release v0.7.0
The export reads ChatGPT's and Claude's undocumented internal web APIs,
which can change shape without notice; the worst failure for a backup tool
is a silent one (skipped/mis-parsed content with no error). Add a `canary`
command + BaseProvider.check_drift() (overridden by ChatGPT/Claude) that
fetches one listing page + one conversation per provider and asserts only
the normalizer's load-bearing fields — not the full response shape, which
churns harmlessly. The top silent risk it guards is a renamed retrieval-tool
author bypassing the hidden-content collapse. ERROR findings exit non-zero;
WARN findings are surfaced but non-fatal so a backup run is never blocked.

Also retires the in-app watch mode and headless StartOS direction from the
roadmap (tool stays a local, manually-run CLI) and updates docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:37:24 -04:00
JesseMarkowitzandClaude Opus 4.8 ef603cf659 doctor: report ChatGPT token health via /api/auth/session error field
ChatGPT session tokens are JWEs whose `exp` is encrypted and unreadable
client-side, so the old JWT-decode path could never yield an expiry, and the
`/api/auth/session` `expires` is a misleading rolling window (it advances on
every call even for a dead token). The honest signal is the `error` field:
`RefreshAccessTokenError` means the session token is dead while `expires` and
a stale `accessToken` are still echoed (verified live).

- Add ChatGPTProvider.session_health() and a "ChatGPT token active" doctor
  check based on it; drop the dead JWE/exp decode path in doctor and auth.
- Fix _fetch_access_token to fail fast on a set `error` instead of returning
  the stale accessToken (which produced a confusing downstream 401).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:37:01 -04:00
JesseMarkowitzandClaude Opus 4.8 4a7ee5773f Remove auth --from-browser browser cookie extraction
Browser cookie auto-extraction is not viable: modern Chromium App-Bound
Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level
layer that cannot be decrypted off disk without admin rights and AV-flagged
SYSTEM impersonation, and fails on Brave specifically. Drop the
`--from-browser` flag, `_auth_from_browser`, `src/browser_tokens.py`, its
test, and the browser-cookie3 dependency. Auth is manual (DevTools wizard)
only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:35:42 -04:00
JesseMarkowitz 3cf0b1eaa8 fix: force re-render campaign tracking, cache dir loading, Joplin link preservation. Co-Authored-By: Fable 5 2026-06-12 22:39:12 -04:00
JesseMarkowitz 456975ad50 fix: --force re-render progression, cache dir loading, and Joplin link preservation. Co-Authored-By: Fable 5 <noreply@anthropic.com> 2026-06-12 21:25:47 -04:00
JesseMarkowitz 9e1a8ab7cb feat: v0.6.0 — collapse policy, session limiter, Claude Code provider, prune, browser auth, media downloads 2026-06-12 18:26:14 -04:00
JesseMarkowitzandClaude Sonnet 4.6 557994f7d9 fix: persist created_at in cache so Joplin note titles get date prefix
mark_exported() was discarding created_at from the metadata dict because
it wasn't in the hardcoded stored-key list, so the joplin sync always
saw an empty date and omitted the prefix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 11:36:21 -04:00
JesseMarkowitzandClaude Sonnet 4.6 e9b2e42893 feat: v0.5.0 — nested Joplin notebooks, date-prefixed note titles, flat year folders
Joplin notebooks now use a two-level hierarchy: AI-ChatGPT / <project> and
AI-Claude / <project> instead of a single flat title. Note titles are prefixed
with the conversation created_at date (YYYY-MM-DD). Export folders collapse
provider/project/year into a single provider/project.year directory.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 11:05:39 -04:00
JesseMarkowitzandClaude Opus 4.7 68e8d532be feat: v0.4.1 — ChatGPT tool-output content types and conv_id fix
First real-data export against v0.4.0 surfaced 66 unknown blocks across
three content types — captured live and added.

Added:
- execution_output (Code Interpreter / container.exec / python tool
  output) → tool_result block. output=content.text,
  tool_name=author.name, is_error=metadata.aggregate_result.status,
  summary=metadata.reasoning_title
- system_error → error tool_result with tool_name=author.name
- tether_browsing_display: spinner placeholders (empty result+summary)
  skip silently with DEBUG log; defensive populated-case branch maps
  to tool_result (untested in real data)
- tool_result block schema: optional `summary` field rendered as
  italic line between header and fence
- tool_result rendering: tool_name appears in header when present
  (e.g. `📤 Result: container.exec`); existing tool_name=None calls
  unchanged
- _ROLE_LABELS["tool"] = ("🔧 Tool", "tool")

Fixed:
- chatgpt.normalize_conversation reads `conversation_id` as fallback
  for `id`. Live API uses conversation_id; fixtures use id.
  Pre-fix: empty id in YAML frontmatter and missing context in
  WARNING logs.

Tests: 11 new (192 total, 0 failures). Fixture extended with 4
tool-output cases (execution_output success, empty execution_output
that should skip, system_error, tether_browsing_display spinner).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-05 09:25:55 -04:00
JesseMarkowitzandClaude Opus 4.7 473d02f71a feat: v0.4.0 — rich content support with typed blocks and loss visibility
Extracts per-message content into a typed `blocks` list (text, code,
thinking, tool_use, tool_result, image_placeholder, file_placeholder,
unknown) and renders them at exporter write time. Voice transcripts,
Custom Instructions, and image references now appear in exports
instead of being silently dropped.

Foundation:
- src/blocks.py: pure block constructors, _safe_fence (fence-corruption
  defense, verified live in Joplin), _blockquote_prefix, render
- src/loss_report.py: per-run tally surfaced as INFO summary at end of
  export so silently-dropped data becomes visible

Providers:
- ChatGPT: dispatch on content_type produces typed blocks; voice shapes
  (audio_transcription, audio_asset_pointer, real_time_user_audio_video_
  asset_pointer) locked from live DevTools capture; Custom Instructions
  bug fix (parts-vs-direct-fields); role filter lifted; hidden-context
  marker driven by is_visually_hidden_from_conversation flag
- Claude: defensive dispatch for text/thinking/tool_use/tool_result/image
  with recursive nested-block flattening; untested against real rich-
  content data — fix-forward in v0.4.1

Exporter:
- Markdown renders from blocks at write time via render_blocks_to_markdown;
  backward-compat fallback to content for any pre-v0.4.0 cached data

Tests:
- 27 new tests across providers, exporters, CLI; fixtures rebuilt with
  real-shape ChatGPT voice + Custom Instructions cases
- 181/181 pass

Behavior changes (intentional):
- JSON output omits content; consumers should read blocks
- Per-conversation message counts increase (Custom Instructions, image-
  only, tool-only messages now appear)
- Existing exports not auto-re-rendered; users wanting fresh output run
  cache --clear then export

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-04 23:17:18 -04:00
Jesse.MarkowitzandClaude Sonnet 4.6 4798edcea7 docs: update README for chunked ChatGPT session cookies
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-19 22:32:01 -04:00
40 changed files with 11157 additions and 576 deletions
+82 -2
View File
@@ -6,10 +6,18 @@
# --- ChatGPT ---
# How to get: open chatgpt.com in Chrome → F12 → Application tab
# → Cookies → https://chatgpt.com → find the two cookie chunks:
# → Cookies → https://chatgpt.com → find the session token cookie. Chrome splits a
# cookie only above ~4KB, so you will see ONE of these two layouts:
#
# __Secure-next-auth.session-token (the whole value) → CHATGPT_SESSION_TOKEN
# (leave _1 empty)
# or, when the token was large enough to be split:
# __Secure-next-auth.session-token.0 (starts with "eyJ") → CHATGPT_SESSION_TOKEN
# __Secure-next-auth.session-token.1 (the remainder) → CHATGPT_SESSION_TOKEN_1
# Token type: JWE. Typically valid for ~7 days.
#
# CHATGPT_SESSION_TOKEN_1 is OPTIONAL — leave it empty when there is no .1 cookie.
# But if a .1 cookie does exist, you must copy it: a partial .0 fails silently
# (HTTP 200 with no accessToken). Token type: JWE. Typically valid for ~7 days.
CHATGPT_SESSION_TOKEN=
CHATGPT_SESSION_TOKEN_1=
@@ -26,6 +34,53 @@ CHATGPT_PROJECT_IDS=
# Token type: opaque string. Typically valid for ~30 days.
CLAUDE_SESSION_KEY=
# --- Claude Code (local agent sessions) ---
# The claude-code provider reads local Claude Code transcripts. By default it
# scans ~/.claude/projects/ (plus $CLAUDE_CONFIG_DIR/projects when that is set).
# To scan additional roots — e.g. other machines' sessions copied onto this box —
# set a ':'-separated list of projects roots. Sessions are merged by folder.
#CLAUDE_CODE_DIR=~/.claude/projects:/mnt/backup/laptop/.claude/projects
#
# Session titles are tagged with the git repos they touched, e.g.
# "Resume StartWRT work [start-technologies]". To never tag specific repos,
# list their names here (comma-separated).
#CLAUDE_CODE_REPO_TAG_IGNORE=some-repo,another-repo
# --- Codex (local agent sessions) ---
# The codex provider reads local Codex CLI rollout files. By default it scans
# ~/.codex/sessions/ (plus $CODEX_HOME/sessions when CODEX_HOME is set).
# To scan additional roots, set a ':'-separated list of sessions roots.
#CODEX_DIR=~/.codex/sessions:/mnt/backup/laptop/.codex/sessions
#
# As with Claude Code, session titles are tagged with the git repos they
# touched. To never tag specific repos, list their names here (comma-separated).
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
# --- Launcher ---
# Read by the ai-chat-exporter wrapper scripts, not by the Python code. The
# wrapper warns when run from outside the repo, because cache/ and exports/
# resolve against the current directory and the wrong one silently starts a
# separate archive. Set to 1 to silence that warning.
#AI_CHAT_EXPORTER_QUIET_CWD=1
# --- Notifications (ntfy) ---
# Push the result of a run to ntfy so an unattended archive reports back — the
# log file, the systemd journal and Task Scheduler's exit code are all pull-only.
# Unset NTFY_TOPIC disables notifications entirely.
#NTFY_TOPIC=my-archive-topic
#
# Self-hosting? Point at your own server.
#NTFY_SERVER=https://ntfy.sh
#
# Bearer token, for access-controlled topics. A topic on public ntfy.sh is
# readable by anyone who knows its name — notifications therefore carry counts
# and a machine name only, never conversation titles.
#NTFY_TOKEN=
#
# always (default) — notify on every run; failure — only when something failed;
# off — never.
#NTFY_NOTIFY=always
# --- Output ---
# Where exported Markdown files are written (default: ./exports)
EXPORT_DIR=./exports
@@ -36,6 +91,31 @@ EXPORT_DIR=./exports
# provider/year → exports/claude/2024/file.md (ignores projects)
OUTPUT_STRUCTURE=provider/project/year
# What to do with content that was invisible in the provider's web UI
# (file-retrieval tool dumps, hidden context like Custom Instructions).
# These dumps can be 90% of a conversation's bytes. Options:
# placeholder (default) → one-line placeholder with tool name and size
# full → keep everything (pre-v0.6.0 behavior)
# omit → drop entirely (still counted in the run summary)
EXPORTER_HIDDEN_CONTENT=placeholder
# Download conversation assets (images, audio) next to the Markdown, into a
# media/ folder, and inline them. Options:
# images (default) → images only
# all → also audio/voice clips and other files
# off → keep text placeholders, download nothing
# Downloaded media is uploaded to Joplin as resources on the next `joplin` run.
EXPORTER_DOWNLOAD_MEDIA=images
# Cap how many conversations are downloaded per export run (per provider).
# Runs are resumable — a capped run continues where it stopped next time.
# Keeps big backfills from looking like scraper traffic. Unset = unlimited.
#MAX_CONVERSATIONS_PER_RUN=25
# Seconds between consecutive API requests (small random jitter is added).
# Default 1.0; set 0 to disable pacing.
#REQUEST_DELAY=1.0
# --- Joplin ---
# Automate importing exported conversations into Joplin as notes.
# Requires Joplin desktop running with the Web Clipper service enabled.
+1
View File
@@ -22,6 +22,7 @@ exports/
!tests/fixtures/*.json
!README.md
!FUTURE.md
!FUTURE-ARCHIVE.md
!CHANGELOG.md
# Cache and logs
+164 -14
View File
@@ -3,19 +3,169 @@
All notable changes to this project will be documented here.
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
## [0.2.0] - Unreleased
### Added
- Joplin import automation: `joplin` command syncs exported Markdown files to Joplin as notes
- Notebooks created automatically per provider+project (`ChatGPT - My Project`, etc.)
- Re-running is safe: notes are updated, not duplicated (Joplin note ID stored in manifest)
- `JOPLIN_API_TOKEN`, `JOPLIN_API_URL`, `JOPLIN_REQUEST_TIMEOUT` config variables
- Configurable request timeout with clear error messages and actionable hints on timeout
- `--project` filter on `export` and `list` commands (case-insensitive substring or `none`)
- ChatGPT Projects support via `CHATGPT_PROJECT_IDS` env var
## [Unreleased]
### Fixed
- **Every scheduled Claude Code sync since 2026-09-24 died with `RecursionError`.** Claude Code's `fork` subagents write a transcript that opens with a copy of the parent turn that spawned them — the fork's own `Agent` call included. `_extract_messages` folds a subagent inline whenever it meets its spawn call, so it met that copy inside the fork, folded the same fork again, and recursed until Python's limit. One fork anywhere in `~/.claude/projects` was enough to fail the whole provider; seven existed across three sessions, and the Codex half of the run, unaffected, hid the cause behind a generic exit 1.
The recursive pass now carries the spawn ids being expanded around it and skips a tool_use whose id is among them — the enclosing subagent block already stands for that call. Because the set accumulates, a longer cycle (A spawns B, whose transcript re-spawns A) stops too. The fork's `<fork-boilerplate>` preamble — generic worker rules the harness prepends — is now stripped with the other harness tags, so a fork block opens with its actual directive.
`TestSubagentFold.test_fork_containing_its_own_spawn_call` reproduces the on-disk shape (a `fork-context-ref` record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original `RecursionError` against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block.
- **A scheduled run that crashed sent no notification, and the providers that survived said "OK".** `sync` pushes its ntfy result from the end of a run it finished, so the `RecursionError` above — and any crash, any exit before the sync starts (the ToS gate, a cache error), a launcher that cannot build its venv — sent nothing at all. Worse, the unit runs each provider as its own `sync`, so codex kept pushing a low-priority "AI archive OK" every morning for the eleven days claude-code was dead: a broken provider was indistinguishable from a quiet day.
The systemd unit's `ExecStart` is now `scheduling/run-sync.sh`, which carries the old per-provider loop and pushes a high-priority **FAILED** notification for any run that exits non-zero without the app's "Sync completed with failures" banner (printed right after the app's own push, so an app-reported failure is not reported twice). The push names the provider and, for a crash, the exception class only — `claude-code: crashed (RecursionError)` — never its message, holding to `src/notify.py`'s counts-only rule for a topic anyone can read. It honours `NTFY_NOTIFY=off` and reads `NTFY_*` the way the app does: environment first, then `.env`. Re-run `install-systemd-timer.sh` to pick it up. The Windows task has no equivalent yet.
Verified against a fake launcher and a local capture server: a crash pushes with the class and without the message, an exit before the sync pushes, an app-reported failure and a success push nothing extra, `off` pushes nothing, and the run still exits 1 if any provider failed.
- **An expired Claude session key reported a raw JSON dump instead of how to fix it.** `_make_request` routed only **401** to the auth handler (`src/providers/base.py`), and claude.ai does not use 401 — an invalid or expired `sessionKey` comes back as `403 permission_error` with `details.error_code = account_session_invalid`. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: `HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}`, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it."
Measured live 2026-09-20 against `GET /api/organizations`: a valid key returns 200, while an expired key, a deliberately malformed key and **no cookie at all** return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated."
Auth detection is now a provider decision rather than a hardcoded status. `BaseProvider._is_auth_failure(response)` defaults to 401 and `ClaudeProvider` overrides it to add 403 **matched on `account_session_invalid`**, not on the bare status — so a genuine permission error, which carries a different code, is still reported as itself rather than being mislabelled an expired key. `_handle_401` is renamed `_handle_auth_failure` and takes the response, because a handler named for one status that must handle two is how this stayed hidden; its messages now state the status actually observed instead of asserting "401 Unauthorized". ChatGPT is unaffected: it does not override the default, so the deleted-asset 403 path is untouched.
Seven regression tests cover the split (`TestAuthFailureDetection`), including the two that matter most: a Claude 403 with a *different* error code must **not** be treated as an auth failure, and a ChatGPT 403 must not either. The docs that repeated the wrong premise — `README.md`'s expiry table and "When Tokens Expire" section, the `auth` wizard's on-screen note, and the `ClaudeProvider` docstring — are corrected in the same change.
- **The docs claimed both ChatGPT cookie chunks were required; they are not.** `README.md` stated flatly that "ChatGPT splits large session tokens across two cookies to stay under the browser's 4KB cookie limit. Both are required," and `.env.example` documented only the chunked layout — so a machine whose session token happens to fit in a single `__Secure-next-auth.session-token` cookie looked broken, with the user hunting for a `.1` that does not exist. Chrome splits a cookie only above ~4KB, so the layout varies by session size and the *same account* can be chunked on one machine and not on another.
The code was already correct: `CHATGPT_SESSION_TOKEN_1` is optional (`src/providers/chatgpt.py:162`) and the `auth` wizard already told you to paste a lone cookie into `.0` and leave `.1` blank. Only the reference docs were wrong, and they are the ones read when setting up a new machine.
Measured 2026-09-20 against `/api/auth/session`, reassembling a real 4089-byte token to test each naming: chunked `.0`+`.1` → 200 with an `accessToken`; the whole value under the unchunked name → 200 with an `accessToken`; the whole value under `.0` alone → 200 with an `accessToken`. The server reassembles a complete value sent under `.0`, so both layouts authenticate as the code already assumed. A *partial* `.0` with its `.1` omitted is the one combination that fails, and it fails **silently** — HTTP 200 with no `accessToken` rather than an error — which is now documented in both files alongside the correction.
- **The test suite sent real push notifications to the developer's phone.** `TestSyncCommand` invokes the actual `sync` command, which calls `load_config()`, which calls `load_dotenv()` — so the real `.env` was loaded and its live `NTFY_TOPIC` used for the POST. Every `pytest` run fired three or four pushes, including a fabricated "codex: 3 conversation(s) failed to export" straight out of a fixture, which is worse than noise: it reports a failure that never happened. Nothing appeared in `cache/logs/exporter.log` to explain it, because every test invocation passes `--no-log-file`.
A `tests/conftest.py` autouse fixture now neutralises the environment for every test: `NTFY_TOPIC`/`NTFY_TOKEN` are emptied, and `NTFY_SERVER` and `JOPLIN_API_URL` are pointed at a closed local port, so a stray topic cannot reach the internet and a test cannot write notes into a real Joplin instance. The values are **emptied rather than deleted** — `load_dotenv(override=False)` skips only keys already present, so deleting one lets `.env` put it back. Verified by instrumenting `requests` across a full run: zero outbound requests, where the same instrumentation without the fixture records POSTs to the live ntfy topic.
## [0.9.0] - 2026-08-18
### Fixed
- **An em dash in a notification title silently dropped the notification.** HTTP header values are latin-1 at best and `requests` raises on anything outside it, so the first real send failed with `'latin-1' codec can't encode character '\u2014'`. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
- **The terms-of-service gate exited 0 without a terminal.** `click.prompt` raises `Abort` on a closed stdin, which the handler treated as a user Ctrl-C and exited 0 — so a scheduled run on a machine that had never acknowledged the notice would report success having archived nothing. Non-interactive invocations now exit 1 with an explanation of how to clear the gate once by hand. Found by running the new systemd unit rather than by reading the code.
- **A single U+0085 in a transcript silently dropped a whole record.** Both local providers split session files with `str.splitlines()`, which breaks not just on `\n` but on U+0085 (NEL), U+2028 and U+2029 — all of which are legal *inside* a JSON string and are written literally by Codex (Rust does not escape non-ASCII). One NEL in captured command output shredded one record into unparseable fragments; the parser logged "skipped 3 unparseable line(s)" and lost the record. Found while exporting a real rollout. Both providers now split on `\n` only, and both have regression tests that write their fixtures with `ensure_ascii=False` — with `json.dumps`' default the hazardous characters are escaped and the bug cannot reproduce.
- **Deleted uploads are no longer reported as permission errors.** ChatGPT's `/backend-api/files/{id}/download` answers a *missing* asset with `403 {"detail":"Forbidden"}`, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned `404 {"detail":"File not found"}` on `/files/{id}`, while assets that downloaded fine returned 200 on both in the same session, and `ChatGPT-Account-Id` made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as `expired-or-missing`. A 403 on an asset that *does* still exist is left alone as `forbidden` — that one would be a real problem.
- **4xx errors now report why.** `_make_request` ended non-retryable statuses with `raise_for_status()`, whose curl_cffi message is `HTTP Error {code}: {reason}` — and HTTP/2 carries no reason phrase, so a refused request logged as bare `HTTP Error 403:` and the response body (the only explanation the provider gives) was discarded. The body's `detail`/`error`/`message` is now carried into the `ProviderError`, redacted and truncated. This is what made the media 403s on `GET /backend-api/files/{id}/download` undiagnosable.
- **`redact_secrets` missed compound key names.** It matched keys exactly, so `access_token`, `api_key`, and `session-token` passed through un-redacted into debug-logged response bodies; matching now applies per word ("keywords", "monkey", "tokenizer" stay intact).
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
## [0.1.0] - Unreleased
### Added
- Initial implementation: ChatGPT and Claude export via internal web APIs
- Markdown and JSON exporters
- Local cache/manifest for incremental sync
- CLI with export, list, cache, doctor, and auth commands
- **ntfy push notifications for unattended runs (`NTFY_TOPIC`).** A scheduled archive reports only to places you have to go and look at — the log file, the systemd journal, Task Scheduler's exit code — so a run that quietly failed every morning would stay quiet. `sync` now pushes its result: success carries per-provider counts at low/default priority, failure carries the reason at high priority with an alert tag, so the two are distinguishable at a glance on a phone. `ai-chat-exporter notify` shows the settings and `--test` sends a test push. `NTFY_NOTIFY=failure` limits it to failures, `off` disables, `--notify`/`--no-notify` override per run. Never fatal: an unreachable ntfy logs a warning and the run still reports its real exit code.
The payload is **counts and a machine name only, never conversation titles** — a topic on public ntfy.sh is readable by anyone who knows its name, and a title is the first line of what you asked. The machine name is included because several machines archive into one topic, where "2 exported" means nothing on its own. `NTFY_SERVER` points at a self-hosted instance and `NTFY_TOKEN` authenticates against an access-controlled topic.
- **`ai-chat-exporter` / `ai-chat-exporter.cmd` launchers — no virtualenv ceremony.** `cd` into the repo and run; the wrapper creates `.venv`, installs dependencies on first run, and reinstalls when `pyproject.toml` changes. A fresh clone goes from nothing to a working command in one step (measured: ~10s), on Linux/macOS and Windows alike, which matters for a tool meant to run on several machines. `cmd.exe` searches the current directory before `PATH`, so Windows needs no `.\` prefix. The working directory is deliberately not changed — `.env`, `cache/` and `exports/` still resolve against it, which is what lets one checkout archive different machines into different places — but the wrapper now warns when you run it from elsewhere, because a different `cache/manifest.json` silently starts a *second* archive rather than failing.
- **`sync` command — `export` then `joplin` in one invocation, with a real exit code.** Intended for schedulers (and the "trivial add-on" FUTURE.md §7 anticipated): it exits non-zero if any conversation failed to export or any note failed to sync, so a scheduled run that achieved nothing is distinguishable from one that had nothing to do. `--skip-joplin` exports only; `--joplin-optional` downgrades an unreachable Joplin to a warning, since the export has already captured the local transcripts and the notes rebuild from the cache on the next run that finds Joplin up.
- **Daily scheduling for both platforms.** `scheduling/install-systemd-timer.sh` (systemd user timer, `Persistent=true` so a machine that was off catches up at boot) and `scheduling/Register-AiChatSyncTask.ps1` (per-user Task Scheduler entry, `-StartWhenAvailable`). `--provider` is repeatable in both, because the right set differs per machine: the local providers need no credentials and always work unattended, while a web provider whose session token has expired would fail the job every single day and train you to ignore it.
- **Codex CLI provider (`--provider codex`).** Archives local Codex agent transcripts from `~/.codex/sessions/**/rollout-*.jsonl` — local-only, like `claude-code`: no tokens, no rate limits, no ToS exposure. Sessions land in their own top-level `AI-Codex` Joplin notebook, with the same prose-only default, repo tags (`CODEX_REPO_TAG_IGNORE`) and multi-root scanning (`CODEX_DIR`, plus `$CODEX_HOME/sessions`).
Codex writes each session twice in one file and the choice between the two layers is the whole design. `response_item` records are the model-facing wire format, where a tool call arrives as *JavaScript* (`tools.exec_command({...})`) because Codex's `exec` tool is code-mode; `event_msg`/`item_completed` records are Codex's own typed items, already decoded into `CommandExecution`/`FileChange`/`Extension` with argv, cwd, exit code and output as fields. Measured over 7 sessions on 0.147.0 (2026-08-18), the typed layer is 1:1 with the raw layer for prose (91 `AgentMessage` ↔ 91 assistant messages, sharing ids) and additionally omits every piece of harness plumbing — all 51 `developer`-role messages plus the 7 `# AGENTS.md instructions…` and 1 `<environment_context>` injections — which the Claude Code provider has to strip by regex. So the typed layer is parsed for content.
Its one gap is that it records what *ran*, not what was *attempted*: 26 of 180 `exec_command` calls produced no item (14 sandbox launch failures, 6 user aborts, ~5 still running at turn end, 1 failure). The raw layer is therefore read for a call count only, and placeholders report the shortfall — `3 calls: exec_command ×3 (+2 did not complete)` — instead of silently under-reporting. `wait` calls are process polls, not attempts, and are excluded.
**Reasoning is not exportable from Codex.** All 345 reasoning records carry `encrypted_content`, with `summary` empty in the raw layer and `summary_text`/`raw_content` empty in the typed layer, in every session. It is always dropped and counted; unlike Claude Code, `EXPORTER_HIDDEN_CONTENT=full` cannot surface it. The sidecar SQLite databases (`state_5.sqlite`, `thread_history_1.sqlite`) are deliberately not read: `thread_history_projection_state` tracks a byte offset into the rollout file, so the JSONL is canonical and the DB derived, and its `title` column is just the first user message truncated.
**Codex Cloud is out of scope, verified rather than assumed.** Cloud tasks are reachable at `chatgpt.com/backend-api/api/codex/tasks{,/list}` — the same host and `/backend-api` root the ChatGPT provider already uses — but local CLI sessions are never uploaded there, so it is not an alternative source for these transcripts. `codex cloud list` confirmed the account holds no cloud tasks. The provider makes no network calls.
- **`projects` command — discover the project IDs your config is missing.** `CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is invisible to the listing pass, so its conversations are never fetched. The command reports every project your conversations belong to, marks which are absent from `.env`, and prints a paste-ready line (`--write` applies it). It reads project ids from the conversation listing when they are there and falls back to `--deep`, one detail request per conversation, when they are not.
- **Project attribution now reads the conversation's own `gizmo_id`.** Previously the project name came only from `CHATGPT_PROJECT_IDS`, so a conversation in a project you had not listed exported into `no-project/` even though its payload names its project. The detail response carries `gizmo_id`, so it is used as a fallback after the listing annotation and the project map — attribution stays correct without maintaining a list, and moving a chat into a new project no longer silently misfiles it. Only `g-p-` ids count: a custom GPT is not a project and must not become a folder. Each unconfigured project is reported once per run, naming the id to add, because the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that live only inside a project never appear in the default listing.
### Changed
- Media download failures are bucketed as `forbidden` (403 — the file record survives) separately from `download-error`, so the run summary distinguishes it from `expired-or-missing` (404).
### Notes
- **ChatGPT media 403s: investigated and closed (2026-08-17).** 19 images across 7 conversations would not download. They are unrecoverable, and not because of anything the exporter or the export schedule did. Findings, recorded so this is not re-litigated: uploads do **not** expire (36/36 sampled images from 2025-09 through 2026-08 are still live, so export cadence is not a factor); the failures split into 7 records that 404 outright and 12 that report `state: "ready"` with a `library_file_id`; `/files/{id}/download` is the endpoint that mints the signed `estuary/content` URL the web UI fetches, and it refuses the survivors with a bare 403 regardless of headers, `Authorization`, `gizmo_id`, `conversation_id`, Referer, or namespace, while the Library id is rejected as `file_not_found`. Decisively, those same images render blank in ChatGPT's own UI — nothing is being withheld from the exporter. All the survivors were created 2026-07-14 within minutes of each other yet appear in conversations predating that date, pointing at a Library migration that kept the metadata and lost the bytes.
## [0.8.0] - 2026-07-06
Focused on the Claude Code provider after the tool's on-disk layout changed and
Claude Code sessions proved hard to find in Joplin.
### Added
- **Subagent capture.** Claude Code now stores subagent (Task-tool) transcripts as separate `<session>/subagents/agent-*.jsonl` files (with an `agent-*.meta.json` sidecar). These were previously invisible to the exporter — all delegated work (reviews, research, plans) was lost. Each subagent is now folded into its parent session inline at the `Task`/`Agent` call that spawned it, rendered as a collapsible `<details>` block labeled with its `agentType`/`description`. The subagent's own tool traffic is collapsed under the same `EXPORTER_HIDDEN_CONTENT` policy as the main dialogue; nesting is handled at any depth via `toolUseId` matching.
- **Repo tags in Claude Code titles.** Sessions launched from a workspace root all land in one folder-named notebook, so titles now carry the repos each session touched, e.g. `Resume StartWRT project work [start-technologies]` — visible in the note list and matched by Joplin search. A file's repo is the git repository it lives in (nearest ancestor with a `.git`), resolved from tool_use paths anywhere in the filesystem — so cross-workspace work is captured and non-repo noise (config dirs, one-off files, reference dirs) is excluded because it isn't a git repo. Frequency-ordered, capped at 3. Escape hatch: `CLAUDE_CODE_REPO_TAG_IGNORE` (comma-separated repo names). Tags reflect the current git layout, so a since-deleted/moved repo drops from the tag on re-export.
- **Multiple projects roots.** `CLAUDE_CODE_DIR` now accepts an `os.pathsep`-separated list of roots (a single path stays backward compatible), and `CLAUDE_CONFIG_DIR`'s `projects/` tree is scanned automatically when set. Sessions from all roots are merged by launch-folder; a session UUID present in two roots keeps the newer-mtime copy.
### Changed
- **Claude Code gets its own top-level Joplin notebook, `AI-ClaudeCode`** (was nested under `AI-Claude` alongside Claude web chats, which made dev sessions hard to find). Existing Claude Code notes self-heal into the new notebook on the next `joplin` run — `update_note` now sets `parent_id`, so a changed provider→notebook mapping relocates notes in place instead of duplicating them. After migrating, the emptied `AI-Claude/{Myworkspace,Services,…}` notebooks can be deleted by hand (Joplin does not auto-remove empty folders).
### Notes
- The sibling `<session>/tool-results/*.txt` sidecars (externalized large tool outputs) are intentionally not captured — tool_result content is collapsed under the default policy anyway.
## [0.7.0] - 2026-06-28
### Added
- `canary` command — probes the ChatGPT/Claude web APIs for schema drift against the fields the parser actually depends on (one listing page + one conversation per provider), not the full response shape. Flags the silent-failure risks for a backup tool: a renamed retrieval-tool author bypassing the hidden-content collapse, a new `content_type`, drifted message fields, or non-empty Claude `attachments`/`files` the normalizer ignores. ERROR findings exit non-zero; WARN findings are surfaced but non-fatal. See `FUTURE.md` §10.
- `doctor` now reports a real ChatGPT token-health check ("ChatGPT token active") based on the `/api/auth/session` `error` field. ChatGPT session tokens are JWEs whose `exp` is encrypted and unreadable client-side; the previous decode path could never yield an expiry. The honest signal is `error == "RefreshAccessTokenError"` (verified live), which means the session token is dead even though `expires`/`accessToken` still echo stale values. See `FUTURE.md` §9.
### Fixed
- `ChatGPTProvider._fetch_access_token` now checks the `/api/auth/session` `error` field and fails fast with a clear "refresh your token" message. Previously it returned the stale `accessToken` present on a dead session, producing a confusing downstream 401 instead of an actionable auth error.
### Removed
- `auth --from-browser` and the `browser-cookie3` dependency (`src/browser_tokens.py`). Browser cookie auto-extraction is not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level layer that can't be decrypted off disk without Administrator rights and AV-flagged SYSTEM impersonation, and fails on Brave specifically (symptom: `Unable to get key for cookie decryption`). Auth is manual (DevTools wizard) only. See `FUTURE.md` §2 for the full rationale.
## [0.6.0] - 2026-06-12
First tagged release. The project was developed through several internal
milestones (0.1.0–0.5.0) that were never published; their changes are
consolidated here.
### Added
**Core export & sync**
- ChatGPT and Claude export via their internal web APIs, with a local cache/manifest for incremental, resumable sync (only new or updated conversations are fetched; each is written to the manifest immediately).
- Markdown and JSON exporters. Markdown rendering happens at exporter-write time (providers produce typed blocks; exporters render them), which keeps the door open for future Obsidian/HTML output.
- CLI: `export`, `list`, `cache`, `doctor`, `auth`, `joplin`, `prune`.
- `--project` filter on `export`/`list`/`joplin` (case-insensitive substring, or `none`).
- ChatGPT Projects support via `CHATGPT_PROJECT_IDS` (project conversations are fetched separately from the default listing).
**Joplin integration**
- `joplin` command syncs exported Markdown to Joplin as notes; notebooks are created automatically and nested per provider+project. Re-running is safe — the Joplin note ID is stored in the manifest, so notes are updated, not duplicated.
- `JOPLIN_API_TOKEN`, `JOPLIN_API_URL`, `JOPLIN_REQUEST_TIMEOUT` config, with actionable error messages on timeout/connection failure.
**Rich content (typed blocks)**
- Messages carry an ordered `blocks` list: text, code, thinking, tool_use, tool_result, citation, image_placeholder, file_placeholder, unknown.
- ChatGPT voice mode: `audio_transcription` parts render as text; audio asset pointers render as `📎 File attached` placeholders with size/duration. Custom Instructions (`user_editable_context`/`model_editable_context`) now appear (were silently dropped) with a `> ℹ️ Hidden context` marker.
- ChatGPT `execution_output`, `system_error`, and `tether_browsing_display` render as `tool_result` blocks (with `tool_name`, `is_error`, optional `summary`); transient browse spinners skip silently.
- Defensive Claude block extraction (text/thinking/tool_use/tool_result/image, including nested-block flattening).
- `_safe_fence` picks a backtick fence longer than any run in the content, so embedded triple-backticks can't corrupt rendering (verified live in Joplin).
- Data-loss visibility: a `LossReport` summary at the end of every `export` run breaks down `unknown blocks` and `extraction failures` by raw type; `unknown` blocks render as visible `> ⚠️ Unsupported content` with the raw type and observed keys.
**Invisible-content collapse (`EXPORTER_HIDDEN_CONTENT=full|placeholder|omit`, default `placeholder`; `export --hidden-content`)**
- Messages that were invisible in the ChatGPT web UI are collapsed to one-line placeholders instead of exported in full. Two triggers: (a) tool-role retrieval dumps identified by `author.name` (`file_search`, `myfiles_browser`) — ChatGPT re-injects the full text of attached/project files on every tool run, measured at 86% of all content bytes across the three largest conversations and **not** hidden-flagged; (b) messages flagged `is_visually_hidden_from_conversation` (Custom Instructions). Code-execution results, web-search output, and all dialogue are untouched.
- New `collapsed` block renders as `> 🔧 **Tool output** — file_search (15.0 KB) — omitted (EXPORTER_HIDDEN_CONTENT=full to keep)` (ℹ️ Hidden context variant for hidden-flagged messages). The `LossReport` gains a `collapsed by policy` section (count per origin + total KB) so the omission stays visible.
**Claude Code session provider (`--provider claude-code`)**
- Archives local agent transcripts from `~/.claude/projects/*/*.jsonl` (override with `CLAUDE_CODE_DIR`) — no tokens, no rate limits, no ToS exposure. Prose-only by default per the hidden-content policy: dialogue and deliverables kept; tool traffic grouped into one placeholder per activity run (`> 🔧 Tool output — 14 calls: Read ×9, Bash ×2 (86KB) — omitted`); thinking dropped (counted in the summary); subagent (`isSidechain`) and harness (`isMeta`, snapshots, command tags) records stripped. Titles from the last `ai-title` record; project from the working-directory basename. Joplin notebooks nest under the `AI-Claude` parent. Measured: 19MB of JSONL → 804KB of Markdown across 25 sessions.
**Binary/image downloads (`EXPORTER_DOWNLOAD_MEDIA=images|all|off`, default `images`; `export --download-media`)**
- Downloads ChatGPT conversation assets into a `media/` folder beside each export and inlines them — images become real `![](media/…)` embeds, files become local links. Two-hop fetch via `/backend-api/files/{id}/download` → signed URL (verified live). `all` also pulls voice-mode audio (≈162MB of clips in this archive, transcripts already in text — hence the images-only default). Idempotent (skips assets already on disk), atomic writes, `600` perms. Expired/missing assets (old generated images 404) keep their placeholder and are counted as `media failed` — never fatal.
- Joplin resource upload: the `joplin` command uploads downloaded media as resources and rewrites `media/…` links to `:/resourceId`, so images render inside notes (verified live: resources created with correct mime/size and linked to the note). Resource IDs are tracked per-conversation in the manifest, so re-syncs reuse them instead of duplicating.
**Token setup & throughput**
- `auth --from-browser [brave|chrome|chromium|edge|firefox]`: extracts ChatGPT session-token cookies and the Claude `sessionKey` straight from a local browser's cookie store (via `browser-cookie3`; defaults to Brave), validates each against the live API, and writes `.env`. Tokens that fail validation never overwrite a working entry; clear fallback when the browser is absent or the cookie DB is locked.
- Session limiter: `--max-conversations N` / `MAX_CONVERSATIONS_PER_RUN` cap downloads per run (per provider); capped-out conversations are reported as "deferred" and resumed next run.
- Polite pacing: `REQUEST_DELAY` (default 1.0s, ±25% jitter, `0` disables) between consecutive API requests, with the existing 429 backoff as the reactive net.
**Re-rendering**
- `export --force` re-exports every conversation even if cached and unchanged, so the whole archive can be re-rendered after a formatting/feature change without `cache --clear`. Runs as a tracked campaign (a stamped start time in the manifest): each run re-renders the least-recently-exported conversations, the "still to go" count shrinks toward zero, finished providers do no further work, and the campaign auto-completes (reporting "Force re-render complete"). Combines with `--max-conversations` to spread the work across runs.
- `mark_exported` now preserves `joplin_note_id` / `joplin_synced_at` / `joplin_resources` across re-exports (refreshing `exported_at`), so a re-export followed by `joplin` updates the existing notes instead of creating duplicates.
- Fix: `.env` is now loaded at the very start of every command, so `CACHE_DIR` and `LOG_FILE` are honored (previously the cache silently used `./cache` regardless of `CACHE_DIR`).
**Archive hygiene**
- `prune` deletes export files not referenced by the manifest (old-layout trees, no-ID orphans), with listing, confirmation, `--dry-run`, `-y`, and empty-directory sweep. Refuses to run when the manifest references no files (so `cache --clear` + `prune` can't wipe the archive). First live run removed 420 stale files (9.4 MB).
- `doctor` verifies manifest ↔ disk integrity: every recorded `file_path` must exist on disk.
### Changed
- ChatGPT role filter that dropped `tool`/`system` messages is **lifted**; all roles route through normal extraction (truly empty messages skip via the empty-content guard).
- `BaseProvider.normalize_conversation` accepts an optional `LossReport` parameter.
- `ChatGPTProvider.normalize_conversation` reads `conversation_id` as a fallback for `id` (live detail responses use `conversation_id`; fixtures use `id`).
- `Config` validates `EXPORTER_HIDDEN_CONTENT`, `EXPORTER_DOWNLOAD_MEDIA`, `MAX_CONVERSATIONS_PER_RUN`, and `REQUEST_DELAY`, and logs the active values at startup.
- New dependency: `browser-cookie3==0.20.1`.
### Fixed
- Custom Instructions (`user_editable_context`/`model_editable_context`) were silently dropped from every conversation (parts-vs-direct-fields mismatch).
### Migration
- JSON exports: messages contain typed `blocks` and may omit the legacy `content` field — external consumers should prefer `blocks`.
- To re-render existing exports with the current rendering and collapse policy: `python -m src.main cache --clear` then `python -m src.main export`, followed by `joplin` to update notes (and upload media resources). Per-conversation message counts may increase as previously-dropped Custom Instructions, image-only turns, and tool-only turns now appear.
### Test suite
- 264 tests, all passing.
+660
View File
@@ -0,0 +1,660 @@
# FUTURE.md — archived 2026-08-18 (pre-v0.9.0 cleanup)
This is the full `FUTURE.md` as it stood before the v0.9.0 cleanup, kept
because it carries the investigation trails behind decisions that are now
recorded in one line each: why Brave cookie extraction is not viable, why the
ChatGPT token's expiry cannot be read client-side, how the drift canary was
designed, and the reasoning behind each closed backlog item.
Nothing here is planned work. The live roadmap is in `FUTURE.md`.
---
# Planned Future Work
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
>
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
> active roadmap is empty and the remaining backlog below has been **closed as
> not needed** — the tool does what it's needed to do as a local, manually-run
> backup CLI. Items are kept for reference only; revisit on demand if a real
> need shows up. Nothing here is planned work.
Items completed in each release are moved to the changelog. Items below the
roadmap were designed for but intentionally not implemented. The codebase is
structured to make each of these additions straightforward if ever revived.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
---
# Roadmap (decided 2026-06-12)
Priorities reflect the tool's primary purpose — a trustworthy backup so that
conversation data is not lost if a provider account is ever closed — plus the
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
local, manually-run CLI; the headless/StartOS direction was dropped
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
(manual refresh is sufficient).
**Now (in order):**
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
(full-archive `export --force` re-export completed 2026-06-13)
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
needs admin + SYSTEM impersonation that AV flags as credential theft, and
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
**Soon, but later:**
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
manually-run CLI, so no in-app polling loop is needed
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
sufficient (source convos live in the cloud and can be re-downloaded;
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
also retires the headless token-freshness problem — manual weekly refresh
is fine.
**Active:**
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
via the `/api/auth/session` `error` field (not `expires`). See §9.
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
`canary` command. See §10.
**Deprioritized** (entries kept at the bottom of this file; revisit on
demand): `--force` flags, per-conversation cache reset, official export-ZIP
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
scope — no significant usage to archive.
---
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, with one scoping correction
from live recon: the retrieval dumps are NOT flagged
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
the discriminator (the hidden flag only marks Custom Instructions and small
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
completed 2026-06-13.
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
the user's own attached documents, the same files dumped dozens of times
per conversation. These messages were invisible in the ChatGPT web UI;
they appear in exports because v0.4.0 lifted the role filter to fix silent
data loss. Custom Instructions hidden-context blocks are a minor secondary
case (~2KB, once per conversation) — the originally planned
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
file contains zero hidden-context blocks.
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
`placeholder`, plus a `--hidden-content` CLI override on `export`.
- `placeholder` renders affected messages as one line with type and size:
`> 🔧 Tool output (file_search, 24KB) — omitted
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
roughly halves; worst files shrink ~90%.
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
decision must happen in the provider, not the renderer); (b) messages
flagged `is_visually_hidden_from_conversation`, including
`user_editable_context` / `model_editable_context` (Custom
Instructions) — subsumes the old suppress-hidden-context idea.
- **Scope — keep at full size:** code-execution `tool_result` blocks and
web-search results; those are usually content the user wants.
- Count collapsed messages in the post-export summary so the omission
stays visible (mirror the LossReport presentation, but as intentional
policy, not loss).
Re-export workflow after shipping: `cache --clear` + `export` (same as
the v0.4.0 migration).
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
different mechanism (see below).
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
symptom hit in the field:
```
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
```
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
rights** and behavior that AV/EDR flags as infostealer activity. The one
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
this, so no off-disk reader is a reliable, non-invasive fit.
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
(launch the browser with `--remote-debugging-port`, read cookies via
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
hand. With the headless/StartOS direction dropped (#8), manual DevTools
refresh is the accepted approach — no automated extraction is needed.
## 3. Claude Code Session Provider — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
during implementation: `isSidechain` records are subagent transcripts (skipped),
`isMeta` marks harness-generated user records (skipped), and listing/normalized
`updated_at` must both use file mtime or the cache would re-export every
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
Archive local Claude Code session transcripts. No tokens, no rate limits,
no ToS risk — the data is already on disk but lives in a single JSONL per
session that Claude Code may clean up, and it contains deliverables
(reviews, plans, analyses) that exist nowhere else.
Decisions (2026-06-12):
- **Rendering: prose-only.** Keep user prompts and assistant text
(including full deliverable write-ups); collapse tool activity to
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
86KB — omitted`); **exclude thinking blocks**.
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
in v0.5.0).
Data facts (measured 2026-06-12):
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
Currently 29 sessions, 18.8MB total, largest 3.8MB.
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
activity is what makes these exports readable.
- Record types: `user` / `assistant` (Anthropic-style `message.content`
block arrays) plus harness records: `ai-title` (use for note title and
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
`permission-mode`, `system` (skip). Strip harness noise from user
messages (`<local-command-caveat>`, `<command-name>` blocks).
Implementation shape: new `src/providers/claude_code.py` implementing the
`BaseProvider` interface — `list_conversations` scans project dirs,
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
existing block schema (content is already block-shaped: text / tool_use /
tool_result / thinking). Incremental sync via file mtime/size recorded in
the existing manifest. Project name derives from the munged cwd dirname.
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
(refuses to prune right after `cache --clear`). First live run removed 420
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
manifest↔disk integrity (293/293 after the run).
A backup is only trustworthy if the on-disk tree matches the manifest.
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
v0.4.1) sit alongside current exports and would double-sync into Joplin.
- `prune` command: delete export files not referenced by the manifest.
`--dry-run` (default off, but always print the list before deleting)
shows what would be removed and why (old layout / orphan / unknown).
- `doctor` extension: verify every manifest entry's `file_path` exists on
disk; report missing files (re-export candidates) and unreferenced files
(prune candidates).
## 5. Binary Content Downloads — SHIPPED v0.6.0
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
`download_url`; a second GET yields the bytes (works for user uploads;
older AI-generated images 404 — expired server-side, handled gracefully).
Asset refs come in three shapes — `sediment://file_…`,
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
images-only is the default and audio is opt-in via `all`.
Original notes below.
**Priority note (2026-06-12): "later, but soon" — under the
backup-if-account-closes goal, embedded images are part of the data that
would be lost; placeholders alone don't preserve them.**
v0.4.0 ships placeholders for images and audio assets but does not download
the binary content. The `_safe_fence`-wrapped placeholders include the asset
reference (`sediment://...` or `file-service://...`), MIME type, size, and
duration where available; the actual bytes are not preserved.
Next steps:
- Download attached images alongside the Markdown export, save under a
`media/` sibling directory with a stable filename derived from the asset
reference.
- Replace `image_placeholder` rendering with an inline `![](relative/path)`
reference once the file is on disk.
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
rewrite the rendered Markdown to use `:/resourceId` references, and track
the resource ID in the cache manifest so re-syncs stay idempotent.
- DALL-E images on the assistant side: not observed in this user's data; the
code path exists (`source = "model_generated"`) but is untested.
The block-level schema is already in place — only the file-fetch + rewrite
layer needs to be added. See the `image_placeholder` and `file_placeholder`
block definitions in `src/blocks.py`.
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
the re-run picked it up.
Cap how many conversations are downloaded in a single `export` run so the
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
most importantly on the very first export, which otherwise fetches the
entire conversation history in one session. Because every run is resumable
(the manifest records each conversation immediately), a capped run simply
exports the first N pending conversations and the next run picks up where
it left off. This keeps traffic looking like a human-paced session rather
than a scraper, reducing the risk of rate limiting or account flags.
Two complementary pieces:
1. **Session cap** — `--max-conversations N` flag (and
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
`export` command, slice the pending list after the cache filter:
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
counts (reuse the message format from the 429 early-exit path) and
remind the user to re-run to continue.
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
random jitter) slept between per-conversation detail fetches in
`BaseProvider`, so even a capped run doesn't fire requests
back-to-back. The existing 429 backoff in `_request` stays as the
reactive safety net.
Note: the conversation *listing* (paginated, 100/page) still runs in full
each time so the cache comparison works — the cap applies to the heavy
per-conversation detail fetches, which dominate request volume.
This is a stepping stone to the StartOS service: a capped, politely-paced
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
traffic profile a headless deployment needs.
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
whatever hosts the tool: a user cron line locally, and on the long-term
StartOS target the platform's own scheduling. Either way the tool only needs
to do one capped, politely-paced `export` + `joplin` run and exit — which it
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
that chains `export` then `joplin` for a single cron line is a trivial
add-on, but the polling loop itself is off the roadmap.
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
Not pursuing a headless StartOS service. The local, manually-run CLI is
sufficient: the source conversations live in the providers' clouds and can be
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
so durability is covered without a server in the loop.
Dropping this also retires the one genuinely hard sub-problem it carried —
session-token freshness without a browser. There is no headless context to
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
extraction remains a dead end regardless — see #2.)
### REOPENED as a TODO (2026-08-18) — centralization, not durability
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
decision rested on "the source conversations live in the providers' clouds and
can be re-downloaded". **That is no longer true of half the providers.**
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
that exist *only* on the machine that produced them — Codex prunes its rollout
files, and neither is recoverable from any cloud. The re-download premise now
covers the web providers only.
The new motivation is consolidation rather than durability: work is split across
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
conversations from everywhere, instead of per-machine islands that only meet
inside Joplin.
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what it is uniquely able to do —
read that machine's local transcripts, and hold the browser session for the web
providers. What changes is where the output goes: instead of syncing to Joplin
itself, a local run **uploads its conversations to the StartOS storage area**,
and the StartOS service owns the Joplin connection for the whole corpus.
That inverts today's arrangement, where every machine talks to its own Joplin
desktop, and it removes two problems we already have:
- **The Joplin-availability race disappears from the clients.** A scheduled run
currently has to find Joplin desktop open on that same machine — the
2026-08-18 09:02 timer run exported fine and then skipped the sync because
Joplin did not start until 09:07. Uploading to a server that is always up has
no such window, and `--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note-update logic run once, server-side, against one manifest — rather than
each machine independently deciding what a notebook is called and racing to
update the same note.
What this would need, and what it would *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers can keep running interactively on the
machine that has the browser, pushing their output to the server. Only the
local providers need to run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter currently does not
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing would make that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should probably keep them rather than move them.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
Shipped: `doctor` now adds a "ChatGPT token active" check via
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
decode path was removed, and `_fetch_access_token` now fails fast on a set
`error` instead of returning a stale token. Tests in
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
below for the record.
Goal: if it's cheap to tell how much longer a token will work, show it on
`doctor`. Findings from live recon:
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
never yields an expiry for the real tokens (falls to the "not decodable"
branch).
- **`/api/auth/session` exposes an `expires`** (the provider already calls
this endpoint in `_fetch_access_token`; the response includes `expires`
alongside `accessToken`). Live value observed 2026-06-28:
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
assumption and the lived weekly-refresh cadence, so it is almost certainly
the NextAuth **rolling session window** (re-extended on every call), not
the point at which the pasted token actually 401s. Displaying it verbatim
would give false confidence.
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
was actually dead (conversations API → 401), `/api/auth/session` still
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
So `expires` is a **rolling session window that rolls forward on every call
even for a dead token**; displaying it would actively lie. The same response
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
token `error` is absent/null; when the session token is dead NextAuth can't
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
health check.
- **Design (ready to build):**
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
falsy and `accessToken` present; on `RefreshAccessTokenError` report
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
work) and do NOT surface `expires`.
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
without checking `error`, so it proceeds with a stale token and yields a
confusing downstream 401 instead of a clear "refresh your token" message.
Check `error` there and fail fast.
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
ChatGPT and Claude). It fetches one listing page + one conversation per
provider and asserts only the normalizer's load-bearing fields, emitting
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
trail below for the record.
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
change shape without notice. The worst failure for a backup tool is *silent*:
a response-schema change that makes the exporter skip or mis-parse content
without erroring. `doctor` currently checks token validity, reachability, and
manifest↔disk integrity — but not "does the provider's response still look
like what the parser expects."
To investigate: a lightweight schema/shape assertion on a known-good sample
of each provider's listing + conversation-detail responses (presence and type
of the fields the normalizers rely on), surfaced as a `doctor` check or a
dedicated canary.
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
- **Dependency surface (assert ONLY these — see below for why):**
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
`project.name`, `chat_messages[]`.
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
(list of typed blocks), `created_at`.
- **Key finding — full-shape diffing is the wrong design.** Claude's
`settings` object is full of volatile internal codenames that churn
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
canary would fire on every UI experiment. The canary MUST target the
normalizer's load-bearing fields only, not the whole response. (Aligns with
the drop-noise-don't-retain-it principle.)
- **Real Claude messages are flat `text`/`sender`** — in this archive every
message had a string `text` and NO `content` block list (0 rich blocks
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
thinking, image, …) is an **unexercised theoretical path**; drift there
can't be "caught" by a canary because it never runs on real data — it's a
safety net for if Claude ever switches to block content. The canary should
assert the flat shape and *warn if `content` ever appears as a list* (that
itself is the drift event that would activate the dormant code).
- **Possible silent-loss spot (separate from drift):** Claude messages carry
`attachments` and `files` arrays (empty in this sample) that the normalizer
ignores entirely. If a user ever attaches files in Claude, they'd be
dropped without a LossReport entry. Worth a follow-up check.
**Live recon — ChatGPT captured 2026-06-28:**
- **Dependency surface (assert ONLY these):**
- listing item: `id`, `title`, `update_time`/`create_time`.
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
`update_time`, `mapping` (non-empty).
- mapping node: `message`, `children` (the tree walk depends on both);
message: `author.role`, `author.name`, `content.content_type`,
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
- **content_type vocabulary observed (all currently handled):** `text`,
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
`execution_output`, `reasoning_recap`, `user_editable_context`,
`tether_browsing_display`. A *new* content_type already degrades gracefully
(visible `unknown` block + WARNING + LossReport tally) — so content_type
drift is **already non-silent**. The canary just needs to confirm the known
set still parses to non-empty blocks.
- **The genuinely silent drift risk — `author.name` collapse keys.**
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
confirms `file_search` is live (and `web.run`/`python` are correctly left
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
stops** and the archive re-bloats with no error or LossReport entry. This is
the top canary target: assert that retrieval-dump tool authors are still
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
- **Second silent risk — empty `content.parts`.** A `text` message whose
`parts` field is renamed/emptied yields zero blocks and is skipped with only
a debug log = silent loss. Canary should assert a sampled `text` message
produces a non-empty block.
**Recon complete for both providers. Canary design (ready to build):** a
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
listing page + one conversation and asserts the dependency-surface fields
above by presence+type — NOT full shape (Claude `settings` codenames prove
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
surface as a warning, never a hard error (a backup tool must still run).
---
# Deprioritized — CLOSED as not needed (2026-06-28)
These were considered and intentionally **not** built. Closed, not planned —
the tool is feature-complete for its purpose. Kept for reference in case a
real need ever revives one: Joplin `--force`, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
output, token-expiry notifications (also moot — see §9), and a search
command. (Also closed, outside this list: handling Claude `attachments`/
`files`, which the canary will flag if they ever appear in real data.)
## Export `--force` Flag — SHIPPED v0.6.0
Implemented 2026-06-12: `export --force` passes `force=True` to
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
preserves Joplin links across re-exports, so a forced re-render + `joplin`
updates existing notes instead of duplicating them.
## Joplin `--force` Flag
Similarly, add `--force` to the `joplin` command to re-sync all cached
conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
Implementation: in `get_joplin_pending()`, return all entries that have a
`file_path` when `force=True`, ignoring `joplin_synced_at`.
## Per-Conversation Cache Reset
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
single conversation without clearing the entire provider cache.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
## Official API Fallback
If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface
(`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON.
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
this user needs (project data), so it isn't a real fallback here.
## Reclassify o1/o3 Reasoning Subparts
v0.4.0 leaves dict parts inside `text` content_type messages with shape
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
shape was inferred from a code comment, not captured live). Once a real
reasoning conversation is captured, reclassify these as `thinking` blocks.
## Obsidian Vault Output
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
## Token Expiry Notifications
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
"proactively notify before expiry" idea is blocked by the same finding: a
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
window rather than the real refresh cadence). Any heads-up — a `doctor`
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
resolving the §9 open question of what signal is actually trustworthy.
## Search Command
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
## Split the README into Separate Documents
**TODO (2026-08-18).** The README is 830 lines / 5,562 words / 37 KB — about a
25-minute read, with 61 headings. An H2-only table of contents was added the
same day and helps navigation, but it treats the symptom: the file is doing at
least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
Not done yet, and not urgent, because it has a real cost the TOC does not: any
existing link into a README section (a bookmark, a note, another repo, a commit
message) breaks when that section moves to another file. Worth doing when the
README next needs substantial editing anyway, rather than as a change of its own.
Two things to decide when it happens:
- Whether `docs/` renders acceptably on the Gitea instance that hosts this repo
(relative links between Markdown files do work there, but worth confirming
before splitting rather than after).
- Whether the anchors in the split files stay stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors are derived from heading text, so they break silently on a reword —
the same fragility the TOC already carries.
+119 -138
View File
@@ -1,163 +1,144 @@
# Planned Future Work
Items completed in each release are moved to the changelog. Items here are
designed for but not yet implemented. The codebase is structured to make each
of these additions straightforward.
> **Status 2026-08-18 (v0.9.0).** The tool archives four providers — two web
> (`chatgpt`, `claude`) and two local agent-transcript (`claude-code`, `codex`)
> — on a schedule, on Linux and Windows, reporting results by push notification.
> Two items below are genuinely planned. Everything else has shipped or been
> decided against.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
Completed work moves to the changelog; this file holds only what is *not* built
yet. Decisions not to build something are recorded at the bottom in one line
each, so they are not re-proposed — the full investigation trails behind them
are in `FUTURE-ARCHIVE.md`.
---
## Export `--force` Flag (v0.2.x)
# Roadmap
Add `--force` to the `export` command to re-export already-cached conversations
without permanently clearing the entire manifest. Useful for re-generating files
after changing the Markdown template or output structure.
## 1. StartOS Service — one corpus, not per-machine islands
Implementation: pass a `force=True` flag to `cache.get_new_or_updated()`, which
returns all conversations regardless of cache state when force is True.
Each machine currently archives to its own `exports/` and its own Joplin. Work
is split across boxes — coding sessions (`claude-code`, `codex`) on the Linux
machine, web chats (`chatgpt`, `claude`) on the Windows one — so there is no
single place where all conversations exist together.
Current workaround: `python -m src.main cache --clear` then re-run export.
This was dropped on 2026-06-28 and **reopened 2026-08-18**, because the
reasoning behind the drop has gone stale. It rested on "the source conversations
live in the providers' clouds and can be re-downloaded". That is no longer true
of half the providers: `claude-code` and `codex` transcripts exist *only* on the
machine that produced them, Codex prunes its rollout files, and neither is
recoverable from any cloud. The motivation is also different from the one
weighed then — consolidation, not durability.
## Joplin `--force` Flag (v0.2.x)
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what only it can do: read that
machine's local transcripts, and hold the browser session for the web providers.
What changes is where the output goes. Instead of syncing to Joplin itself, a
local run **uploads its conversations to the StartOS storage area**, and the
StartOS service owns the Joplin connection for the whole corpus.
Similarly, add `--force` to the `joplin` command to re-sync all cached
conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
That inverts today's arrangement and removes two problems we already have:
Implementation: in `get_joplin_pending()`, return all entries that have a
`file_path` when `force=True`, ignoring `joplin_synced_at`.
- **The Joplin-availability race leaves the clients.** A scheduled run currently
has to find Joplin desktop open on that same machine — the 09:02 timer run on
2026-08-18 exported fine and then skipped the sync because Joplin did not
start until 09:07. A server that is always up has no such window, and
`--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note updates run once, server-side, against one manifest — rather than each
machine independently deciding what a notebook is called and racing to update
the same note.
## Per-Conversation Cache Reset (v0.2.x)
What this needs, and what it does *not*:
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
single conversation without clearing the entire provider cache.
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers keep running interactively on the
machine that has the browser, and push their output to the server. Only the
local providers would run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter does not track today:
`claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing makes that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should keep them rather than hand them off.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 2. Split the README into separate documents
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61
headings. An H2-only table of contents was added 2026-08-18 and helps, but it
treats the symptom: the file is doing at least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
Not urgent, because it has a real cost the TOC did not: any existing link into a
README section — a bookmark, a note, another repo, a commit message — breaks
when that section moves to another file. Worth folding into the next substantial
README edit rather than doing as a change of its own.
Two things to settle when it happens:
- **`.gitignore` ignores `*.md` on purpose** — exported conversations are
Markdown and may contain private content — and re-includes each doc by name
(`!README.md`, `!FUTURE.md`, …). New files under `docs/` will be silently
ignored, with no error, until `!docs/*.md` is added. This already bit the
creation of `FUTURE-ARCHIVE.md` on 2026-08-18.
- Whether `docs/` renders acceptably on the Gitea instance hosting this repo.
Relative links between Markdown files do work there, but confirm before
splitting rather than after.
- Whether anchors in the split files are stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors derive from heading text and break silently on a reword — the same
fragility the TOC already carries.
---
## Official API Fallback (v0.3.0)
# Decided against
If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
Recorded so they are not re-proposed. Full reasoning and the recon behind each
is in `FUTURE-ARCHIVE.md`.
The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface
(`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON.
| Item | Verdict |
|------|---------|
| Brave/Chromium cookie auto-extraction | **Not viable** (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
| In-app watch/polling loop | **Dropped** (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as `sync` plus a systemd timer and a Windows scheduled task. |
| Proactive token-expiry notification | **Blocked, and moot** (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and `/api/auth/session`'s `expires` is a rolling window that advances even for a dead token. `doctor`'s `error`-based health check covers the real need. |
| Official export-ZIP fallback | **Closed** (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. `BaseProvider` still admits a `FileProvider` if that changes. |
| Joplin `--force` flag | Closed as not needed (2026-06-28). |
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
| Search command | Closed as not needed (2026-06-28). `grep`/`ripgrep` over `EXPORT_DIR` covers it. |
| Claude `attachments` / `files` handling | Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
---
## Rich Content Support (v0.4.0)
# Completed
Currently only text content is exported. Future versions should handle:
Detail for each release is in `CHANGELOG.md`.
### Claude
- Artifacts (code, documents, HTML) — export as separate files, link from Markdown
- Uploaded images — download and embed or link
- Extended thinking/reasoning blocks — include as collapsible sections
- Tool call results and web search citations — include as footnotes or appendices
### ChatGPT
- DALL-E generated images — download and embed or link
- Code Interpreter outputs — export code and results
- File attachments — download and reference
- Voice transcripts — include as text
Implementation note: the normalized message schema already includes a
`content_type` field placeholder. When this work begins, extend the schema
rather than replacing it. Non-text content already logs a WARNING when
encountered so users can see what was skipped.
---
## Scheduled / Watch Mode (v0.5.0)
Add a `watch` command (or cron integration helper) to run exports automatically
on a schedule:
```bash
python -m src.main watch --interval 6h # poll every 6 hours
```
This would run `export` + `joplin` in sequence, then sleep. Alternatively,
provide a `cron` command that prints the correct crontab line for the user's
setup.
Implementation: simple loop with `time.sleep()`, or emit a crontab entry
string that calls the export and joplin commands in sequence. A `--once`
flag would do a single run then exit (useful for cron itself).
---
## Obsidian Vault Output (v0.5.0)
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
---
## Joplin Nested Notebooks (future)
Currently notebooks are flat: `ChatGPT - My Project`. Joplin supports nested
notebooks via `parent_id`. A future option (`JOPLIN_NESTED_NOTEBOOKS=true`)
could create a two-level hierarchy:
```
ChatGPT/
My Project/
No Project/
Claude/
Budget Tracker/
```
Implementation: `get_or_create_notebook` would first find/create the provider
notebook, then find/create the project notebook as a child.
---
## Token Expiry Notifications (future)
Proactively warn when a token is close to expiry (within 48h for ChatGPT),
rather than only surfacing the warning at startup. Options:
- Add an `expiry` subcommand that prints token status and exits non-zero if
any token is expired or expiring soon (useful in scripts/cron)
- Send a desktop notification via `notify-send` (Linux) or `osascript` (macOS)
when a token is within 24h of expiry
---
## Search Command (future)
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
- **v0.1.0** — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- **v0.2.0** — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- **v0.4.0** — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via `LossReport` and visible `unknown` blocks
- **v0.5.0** — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- **v0.6.0** — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT`); Claude Code session provider; `prune` + manifest integrity; binary content downloads; session limiter and request pacing; `export --force`
- **v0.7.0** — `canary` drift detection; real ChatGPT token-health check on `doctor`; removal of the non-viable browser-cookie path
- **v0.8.0** — Claude Code coverage reopened: subagent capture as folded `<details>`, repo `[tags]` in titles, its own `AI-ClaudeCode` notebook with self-healing note moves, multi-root scanning (`CLAUDE_CODE_DIR`, `CLAUDE_CONFIG_DIR`)
- **v0.9.0** — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms; `sync` with a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution via `gizmo_id` and the `projects` command; a silent data-loss fix in both local providers (`splitlines` breaking on U+0085/U+2028/U+2029)
+456 -27
View File
@@ -6,7 +6,28 @@ Supports incremental sync — only new or updated conversations are exported on
---
## ⚠️ Terms of Service Warning
## Contents
- [Terms of Service Warning](#terms-of-service-warning)
- [Installation](#installation)
- [First Run: Run Doctor](#first-run-run-doctor)
- [Getting Your Session Tokens](#getting-your-session-tokens)
- [The `auth` Command](#the-auth-command)
- [`.env` Setup](#env-setup)
- [ChatGPT Projects](#chatgpt-projects)
- [Claude Code Sessions](#claude-code-sessions)
- [Codex Sessions](#codex-sessions)
- [Scheduling a Daily Run](#scheduling-a-daily-run)
- [Output Structure](#output-structure)
- [CLI Reference](#cli-reference)
- [How the Cache Works](#how-the-cache-works)
- [Troubleshooting](#troubleshooting)
- [Future Work](#future-work)
- [Security Notes](#security-notes)
---
## Terms of Service Warning
**Read this before using this tool.**
@@ -32,10 +53,31 @@ This tool is designed for a single user backing up their own conversations. Do n
```bash
git clone <repo-url>
cd ai-chat-exporter
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cd AIChatExporter
./ai-chat-exporter doctor
```
That's the whole install. The `ai-chat-exporter` wrapper creates `.venv` and
installs dependencies on first run, and reinstalls whenever `pyproject.toml`
changes — so there is no `python3 -m venv` / `source .venv/bin/activate` to
remember, on this machine or the next one you clone onto. Every command in this
README works the same way:
```bash
./ai-chat-exporter export --provider all
./ai-chat-exporter sync
```
The wrapper does **not** change directory: `.env`, `cache/` and `exports/` all
resolve against your current directory, which is what lets one checkout archive
different machines into different places. Run it from the repo. (It warns if you
don't, because a different working directory means a different `cache/manifest.json`
— which would re-export everything and orphan your existing Joplin notes.)
Prefer the traditional route, or want the test dependencies? That still works:
```bash
python3 -m venv .venv && source .venv/bin/activate && pip install -e ".[dev]"
```
### Windows
@@ -44,12 +86,28 @@ No admin access required. Run these in **Command Prompt** (`cmd.exe`) — it's t
```bat
git clone <repo-url>
cd ai-chat-exporter
python -m venv .venv
.venv\Scripts\activate
pip install -e ".[dev]"
cd AIChatExporter
ai-chat-exporter doctor
```
`ai-chat-exporter.cmd` does the same bootstrap as the POSIX wrapper — it creates
`.venv` and installs dependencies on first run. No `python -m venv`, no
`.venv\Scripts\activate`.
**How you invoke it differs between the two Windows shells:**
| Shell | Command |
|---|---|
| Command Prompt (`cmd.exe`) | `ai-chat-exporter export --provider all` |
| PowerShell | `.\ai-chat-exporter.cmd export --provider all` |
`cmd.exe` searches the current directory before `PATH` and resolves the bare name
through `PATHEXT`, so it finds `ai-chat-exporter.cmd` with no prefix and no
extension. PowerShell deliberately does *not* search the current directory, and
`.\ai-chat-exporter` there would resolve to the extensionless POSIX script, which
PowerShell cannot execute — so name the `.cmd` explicitly. Command Prompt is the
simpler of the two here.
All `ai-chat-exporter` commands work identically in Command Prompt.
**Using PowerShell instead?** If you prefer PowerShell, you may need to allow script execution first (one-time, current user only):
@@ -79,7 +137,17 @@ Before anything else, validate your setup:
ai-chat-exporter doctor
```
This checks token presence, format, expiry, directory permissions, disk space, and live API connectivity. Fix any failures before proceeding.
This checks token presence, format, token health (via the `/api/auth/session` `error` field), directory permissions, disk space, and live API connectivity. Fix any failures before proceeding.
### Checking for API drift
The exporter reads ChatGPT's and Claude's undocumented internal web APIs, which can change shape without notice. The worst failure for a backup tool is *silent* — a response change that makes the exporter skip or mis-parse content without erroring. Run the canary to catch that early:
```bash
ai-chat-exporter canary
```
It probes one conversation per provider and checks only the fields the parser depends on (a renamed retrieval-tool author that would bypass the content-collapse, a new content type, drifted message fields, ignored attachments). `OK` means the live schema still matches; `WARN`/`ERROR` tells you exactly what changed. Worth running before a large export or on a schedule.
---
@@ -87,12 +155,22 @@ This checks token presence, format, expiry, directory permissions, disk space, a
Session tokens are how your browser stays logged in. This tool uses them to access your chat history on your behalf.
Tokens are entered manually — copy them from your browser's DevTools and run the wizard:
```bash
ai-chat-exporter auth
```
The wizard detects your OS, shows the correct DevTools shortcut, and writes the values to `.env` without echoing them to the terminal.
> **Why no auto-extraction from the browser?** Modern Chromium browsers (Chrome 127+, current Brave) encrypt cookies on Windows with **App-Bound Encryption**: the key is bound to the browser through a SYSTEM-level service, so reading cookies off disk requires Administrator rights and a PsExec-style SYSTEM impersonation that antivirus flags as credential theft — and it still fails on Brave specifically. There is no reliable, non-invasive way to pull these cookies automatically, so the manual DevTools flow below is the supported path.
### Token Lifetimes
| Provider | Cookie Name | Lifetime | Expiry Detection |
|----------|-------------|----------|-----------------|
| ChatGPT | `__Secure-next-auth.session-token` | ~7 days | JWT `exp` claim (decoded automatically) |
| Claude | `sessionKey` | ~30 days | Only detectable via 401 response |
| ChatGPT | `__Secure-next-auth.session-token` (split into `.0` + `.1` when over ~4KB) | refresh ~weekly | `error` field of `/api/auth/session` — `doctor` reports "ChatGPT token active". The token is an encrypted JWE, so its `exp` is **not** readable client-side, and the `expires` field is a misleading rolling window; the `error` (`RefreshAccessTokenError` when dead) is the honest signal. |
| Claude | `sessionKey` | ~30 days | Opaque token — only detectable from an API rejection. claude.ai answers an invalid session with **403** `permission_error` / `account_session_invalid`, **not** 401; `doctor` reports it on the "Claude API reachable" row. |
### Finding Tokens in Chrome DevTools
@@ -102,16 +180,38 @@ Session tokens are how your browser stays logged in. This tool uses them to acce
4. In the left panel, expand **Cookies** and click the site URL
5. Find the cookie by name and copy its **Value**
**ChatGPT:** go to `https://chatgpt.com` → find `__Secure-next-auth.session-token` → copy Value (starts with `eyJ`)
**ChatGPT:** go to `https://chatgpt.com` → find the session token cookie. You will
see **one of two layouts**, depending on how large your session token is:
- **One cookie**, `__Secure-next-auth.session-token` — copy Value → `CHATGPT_SESSION_TOKEN`, and leave `CHATGPT_SESSION_TOKEN_1` empty.
- **Two cookies**, `__Secure-next-auth.session-token.0` and `.1` — copy `.0` (starts with `eyJ`) → `CHATGPT_SESSION_TOKEN`, and `.1` → `CHATGPT_SESSION_TOKEN_1`.
Chrome splits a cookie only when it exceeds ~4KB, so a larger session is chunked
and a smaller one is not — the same account can differ from machine to machine.
`CHATGPT_SESSION_TOKEN_1` is optional; both layouts authenticate, because the
server reassembles a complete value sent under the `.0` name.
What does *not* work is sending a **partial** chunk — `.0` on its own when a `.1`
exists. That fails silently: `/api/auth/session` answers HTTP 200 with no
`accessToken` rather than an error. If you see two cookies, copy both.
**Claude:** go to `https://claude.ai` → find `sessionKey` → copy Value
### When Tokens Expire
When a token expires you'll see a `401 Unauthorized` error. To refresh:
An expired token shows up as an authentication error naming the cookie to
refresh and how. The status differs by provider — ChatGPT reports 401, while
claude.ai reports **403 "Invalid authorization"** (`account_session_invalid`) —
so don't read a 403 from Claude as a permissions problem with your account.
To refresh:
- Re-run the `auth` wizard: `ai-chat-exporter auth`
- Or manually update the value in your `.env` file
`ai-chat-exporter doctor` is the quickest check: the "token set" rows only test
that a value is present, so an expired credential passes those and fails on the
"API reachable" row.
---
## The `auth` Command
@@ -138,7 +238,8 @@ cp .env.example .env
| Variable | Description |
|----------|-------------|
| `CHATGPT_SESSION_TOKEN` | Your ChatGPT JWT session token (`eyJ…`) |
| `CHATGPT_SESSION_TOKEN` | ChatGPT session token chunk `.0` (starts with `eyJ…`) |
| `CHATGPT_SESSION_TOKEN_1` | ChatGPT session token chunk `.1` (the remainder) |
| `CHATGPT_PROJECT_IDS` | Comma-separated ChatGPT project IDs (see below) |
| `CLAUDE_SESSION_KEY` | Your Claude session key |
@@ -148,6 +249,10 @@ cp .env.example .env
|----------|---------|-------------|
| `EXPORT_DIR` | `./exports` | Where to write exported Markdown files |
| `OUTPUT_STRUCTURE` | `provider/project/year` | Folder structure (see below) |
| `EXPORTER_HIDDEN_CONTENT` | `placeholder` | What to do with content invisible in the provider's web UI (file-retrieval tool dumps, Custom Instructions): `placeholder` collapses to a one-line note with size, `full` keeps everything, `omit` drops it. Dumps can be 90% of a conversation's bytes. |
| `EXPORTER_DOWNLOAD_MEDIA` | `images` | Download conversation assets into a `media/` folder beside each export and inline them: `images` (images only), `all` (also audio/voice clips), `off` (text placeholders only). Downloaded media is uploaded to Joplin as resources on the next `joplin` run. |
| `MAX_CONVERSATIONS_PER_RUN` | unlimited | Cap downloads per export run (per provider). Runs are resumable, so re-running continues where the cap stopped — useful to spread a large first export across several sessions. |
| `REQUEST_DELAY` | `1.0` | Seconds between consecutive API requests, with small random jitter, so traffic stays human-paced. Set `0` to disable. |
### Joplin
@@ -157,12 +262,26 @@ cp .env.example .env
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
### Notifications
| Variable | Default | Description |
|----------|---------|-------------|
| `NTFY_TOPIC` | — | [ntfy](https://ntfy.sh) topic to push run results to. Unset disables notifications entirely. |
| `NTFY_SERVER` | `https://ntfy.sh` | Point at your own host if self-hosting. |
| `NTFY_TOKEN` | — | Bearer token, for access-controlled topics. |
| `NTFY_NOTIFY` | `always` | `always` notifies on every run, `failure` only when something failed, `off` never. |
A topic on public ntfy.sh is readable by anyone who knows its name, so
notifications carry per-provider counts and a machine name only — never
conversation titles. See [Getting notified](#getting-notified).
### Cache & logging
| Variable | Default | Description |
|----------|---------|-------------|
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
| `AI_CHAT_EXPORTER_QUIET_CWD` | — | Set to `1` to silence the launcher's warning when run from outside the repo. Read by the `ai-chat-exporter` wrapper scripts, not by Python; the scheduler installers set it, since they always set the correct working directory. |
---
@@ -172,6 +291,28 @@ ChatGPT project conversations are stored separately from your main conversation
### Finding your project IDs
The quickest way is to let the exporter find them:
```
ai-chat-exporter projects
```
It lists every project your conversations actually belong to, marks which are
missing from `.env`, and prints a paste-ready `CHATGPT_PROJECT_IDS=` line
(`--write` updates `.env` for you). If your ChatGPT account does not include
the project on conversation summaries, add `--deep` and it reads each
conversation's detail instead — slower, one request per conversation, but
complete.
Why it matters: project attribution is resolved from each conversation's own
`gizmo_id`, so exports file correctly whether or not a project is configured.
But the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that
live only inside a project never appear in the default conversation list, so an
unlisted project's chats are never fetched at all. Every export run also names
any unconfigured project it encounters.
To find them by hand instead:
1. Open ChatGPT and click a Project in the left sidebar
2. Look at the browser URL — it will look like:
`https://chatgpt.com/g/g-p-68c2b2b3037c8191890036fb4ae3ed9f-my-project/project`
@@ -187,6 +328,161 @@ The `auth` wizard can also guide you through this step interactively.
---
## Claude Code Sessions
The `claude-code` provider archives your local [Claude Code](https://claude.com/claude-code) agent transcripts — no tokens, no API, no ToS exposure. Sessions are read from `~/.claude/projects/`.
```bash
ai-chat-exporter export --provider claude-code
ai-chat-exporter joplin --provider claude-code
```
Exports are prose-only by default: your prompts and Claude's write-ups are kept, tool activity is grouped into one-line placeholders (`> 🔧 Tool output — 14 calls: Read ×9, Bash ×2 (86KB) — omitted`), and internal reasoning is dropped (counted in the run summary). Set `EXPORTER_HIDDEN_CONTENT=full` to keep everything. Sessions become notes under their own top-level **`AI-ClaudeCode`** Joplin notebook, in a sub-notebook per launch folder. The provider is included in `--provider all` whenever a sessions directory exists.
**Subagents.** Claude Code stores subagent (Task-tool) transcripts as separate files under `<session>/subagents/`; each is folded into its parent session inline, as a collapsible `<details>` block labeled with the subagent's type and description (its own tool traffic is collapsed like the main dialogue).
**Repo tags.** Sessions launched from a workspace root all share one folder-named notebook, so titles carry the repos each session touched — `Resume StartWRT project work [start-technologies]` — for at-a-glance scanning and search. A file's repo is the git repository it lives in (nearest ancestor with a `.git`), resolved from the tool paths in the transcript, so work is tagged wherever it happened — even across workspaces — and config/one-off files are ignored (they aren't repos). Repos are frequency-ordered and capped at 3. To never tag specific repos, set `CLAUDE_CODE_REPO_TAG_IGNORE` (comma-separated). Note this reads your current git layout, so a repo you later delete or move drops from the tag on re-export.
**Multiple locations.** By default the provider scans `~/.claude/projects/` plus `$CLAUDE_CONFIG_DIR/projects` when `CLAUDE_CONFIG_DIR` is set. To scan additional roots (e.g. other machines' sessions copied onto this box), set `CLAUDE_CODE_DIR` to a `:`-separated list:
```bash
CLAUDE_CODE_DIR="$HOME/.claude/projects:/mnt/backup/laptop/.claude/projects"
```
Sessions from all roots are merged by folder (no per-machine label); if the same session UUID appears in two roots, the newer copy wins. Note the exporter only sees this machine's disk and cannot recover sessions Claude Code has already pruned — run it regularly.
---
## Codex Sessions
The `codex` provider archives your local [Codex CLI](https://chatgpt.com/codex) agent transcripts — same deal as Claude Code: no tokens, no API, no ToS exposure. Rollout files are read from `~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl`.
```bash
ai-chat-exporter export --provider codex
ai-chat-exporter joplin --provider codex
```
Exports are prose-only by default, with the same placeholder format as Claude Code, into their own top-level **`AI-Codex`** Joplin notebook. Repo tags work the same way (`CODEX_REPO_TAG_IGNORE` to suppress), and additional roots can be scanned with `CODEX_DIR` (`:`-separated); `CODEX_HOME`'s `sessions/` is picked up automatically when that variable is set.
Three things differ from Claude Code, all forced by how Codex stores its data:
**Reasoning cannot be exported.** Codex encrypts it at rest — every reasoning record carries `encrypted_content` with no plaintext summary in any layer of the file. It is always dropped and counted; `EXPORTER_HIDDEN_CONTENT=full` cannot bring it back.
**Incomplete tool calls are reported.** Codex writes each session twice in one file: its own typed items (what actually ran) and the raw model-facing wire format (everything attempted). The provider reads the typed layer — it is already decoded, and it omits harness plumbing that would otherwise need stripping — but cross-checks the raw layer for calls that never produced a result, so placeholders read `3 calls: exec_command ×3 (+2 did not complete)`. Those are commands that failed to launch, that you aborted, or that were still running when the turn ended.
**Cloud tasks are out of scope.** `codex cloud` tasks run server-side and are reachable at `chatgpt.com/backend-api/api/codex/tasks`, but local CLI sessions are never uploaded there, so the cloud API is not an alternative source for these transcripts and this provider stays entirely offline. If you start using `codex cloud exec`, those transcripts would be cloud-only and would need separate work.
---
## Scheduling a Daily Run
`sync` chains `export` then `joplin` in one invocation, which is what a scheduler
wants — one command, and a meaningful exit code so a failed run is visible
instead of silent.
```bash
./ai-chat-exporter sync # every configured provider
./ai-chat-exporter sync --provider codex # just one
```
### Linux (systemd user timer)
```bash
./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
./scheduling/install-systemd-timer.sh --uninstall
```
Defaults to 09:00 daily (`--time 21:30` to change). `Persistent=true` means a
machine that was off at the scheduled time runs the archive at next boot rather
than skipping the day. To archive while logged out, `loginctl enable-linger $USER`.
Check on it with `systemctl --user list-timers aichat-sync.timer` and
`journalctl --user -u aichat-sync.service -n 50`.
### Windows (Task Scheduler)
```powershell
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
```
Per-user task, no admin rights needed. `-StartWhenAvailable` is the counterpart
of systemd's `Persistent=true`.
### What to know before you rely on it
**List only the providers that work unattended on that machine.** The local
providers (`claude-code`, `codex`) need no credentials and always work. The web
providers depend on a session token that expires and can only be refreshed by
hand via DevTools — so on a machine where that token is stale, scheduling them
means a failed run every single day, which is a good way to learn to ignore
failures you actually want to see. That is why `--provider` is repeatable in both
installers: schedule the coding machine for `claude-code` + `codex`, the browser
machine for `chatgpt` + `claude`.
**Both installers pass `--joplin-optional`.** If Joplin desktop isn't running,
the sync warns instead of failing: the export has already captured the local
transcripts (the part that can disappear), and the notes are rebuilt from the
cache on the next run that finds Joplin up.
**One provider failing does not skip the rest.** Both installers loop over the
providers in a single action rather than one action each, because systemd
`oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only
the last action's result. Every provider is attempted; the run still exits
non-zero if any failed. On Linux the loop is `scheduling/run-sync.sh`, which the
unit's `ExecStart` calls.
### Getting notified
A scheduled run is silent by default. Output goes to three pull-only places: the
exporter's own log (`cache/logs/exporter.log`), the systemd journal on Linux
(`journalctl --user -u aichat-sync.service`), and Task Scheduler's
`LastTaskResult` on Windows.
To have runs report back, set an [ntfy](https://ntfy.sh) topic in `.env`:
```bash
NTFY_TOPIC=my-archive-topic
```
Then check it works before waiting on a scheduled run:
```bash
ai-chat-exporter notify # show settings
ai-chat-exporter notify --test # send a test push
```
`sync` pushes a result whenever a topic is configured — a success carries the
per-provider counts at low/default priority, a failure carries the reason at
high priority with an alert tag, so a failed archive is distinguishable from a
quiet one on your phone. `NTFY_NOTIFY=failure` notifies only on failure; `off`
disables it; `--notify` / `--no-notify` override per run.
`sync` can only push from the end of a run it finished. A crash, an exit before
the sync starts (the ToS gate, a cache error) or a launcher that can't build its
venv sends nothing — and since each provider pushes separately, the ones that
succeeded still say "OK", so a dead provider looks like a quiet day. On Linux,
`scheduling/run-sync.sh` closes that gap: any run that exits non-zero without the
app having reported it gets a high-priority **FAILED** push naming the provider
and, for a crash, the exception's class (`claude-code: crashed (RecursionError)`)
— the class only, never its message, which can carry a conversation title. The
traceback is in the journal. The Windows task has no such backstop yet.
The message includes the **machine name**, which matters because both machines
archive into one topic. It contains counts only — never conversation titles. A
topic on public ntfy.sh is readable by anyone who knows its name, so if you want
it private, self-host (`NTFY_SERVER`) or use an access-controlled topic with
`NTFY_TOKEN`.
A notification is never fatal: if ntfy is unreachable, the run logs a warning and
still reports its real exit code.
**Acknowledge the ToS notice once, interactively.** It's stored in the cache
manifest per machine. Until then a scheduled run exits 1 with an explanation
rather than hanging on a prompt no one can answer.
---
## Output Structure
All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks.
@@ -251,10 +547,10 @@ Each provider+project combination maps to a flat Joplin notebook created automat
### `auth` — Interactive token setup
```bash
ai-chat-exporter auth
ai-chat-exporter auth # manual wizard (DevTools flow)
```
Guided wizard to find and save session tokens and ChatGPT project IDs. Detects OS and shows the correct DevTools shortcut.
Guided wizard to find and save session tokens and ChatGPT project IDs. Detects OS and shows the correct DevTools shortcut, then writes the values to `.env`.
### `doctor` — Health check
@@ -295,7 +591,22 @@ ai-chat-exporter export --output /path/to/my/notes
ai-chat-exporter export --dry-run
```
Options: `--provider [chatgpt|claude|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--dry-run`
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--hidden-content [full|placeholder|omit]`, `--download-media [images|all|off]`, `--max-conversations N`, `--force`, `--dry-run`
**Re-rendering the whole archive after an upgrade.** New formatting or features (collapse policy, media downloads) only change conversations as they're re-exported. To re-render everything you already have, use `--force` — it re-exports every conversation even if unchanged, **without** `cache --clear`, so your Joplin note links are preserved (a later `joplin` run updates the existing notes instead of duplicating them).
A force re-render runs as a tracked "campaign": each run re-renders the least-recently-exported conversations, the "still to go" count shrinks each run, and finished providers do no further work. The simplest approach is one uncapped pass:
```bash
ai-chat-exporter export --force # re-renders everything in one paced pass
ai-chat-exporter joplin # update notes + upload media
```
Or spread the load with `--max-conversations`, re-running until it reports "Force re-render complete":
```bash
ai-chat-exporter export --force --max-conversations 50 # repeat until complete
```
### `list` — List conversations
@@ -342,7 +653,119 @@ Reads the local export cache and pushes each exported Markdown file to Joplin as
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
4. Joplin desktop must be open when you run this command
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--project NAME`, `--dry-run`
### `sync` — Export and sync in one run
```bash
# The whole archive run: export, then push to Joplin
ai-chat-exporter sync
# One provider
ai-chat-exporter sync --provider codex
# Export only; don't touch Joplin
ai-chat-exporter sync --skip-joplin
# Joplin being closed is a warning, not a failure (used by the schedulers)
ai-chat-exporter sync --joplin-optional
```
Equivalent to `export` followed by `joplin` with the same `--provider`. Intended
for scheduled runs — see [Scheduling a Daily Run](#scheduling-a-daily-run).
Unlike the individual commands, `sync` sets a **meaningful exit code**: non-zero
if any conversation failed to export or any note failed to sync. A provider whose
listing call fails outright (an expired web session token being the usual cause)
counts its whole batch as failed. A provider that is simply unconfigured, or that
had nothing new, is ordinary success. `export` on its own always exits 0, which
is fine when you're reading the summary table and useless to a scheduler.
`--joplin-optional` downgrades an unreachable Joplin to a warning: the export has
already captured the local transcripts, and the notes are rebuilt from the cache
by the next run that finds Joplin open.
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--since YYYY-MM-DD`, `--hidden-content [full|placeholder|omit]`, `--max-conversations N`, `--skip-joplin`, `--joplin-optional`, `--notify/--no-notify`, `--dry-run`
Note this is a deliberate subset of `export`'s options — `--format`, `--output`,
`--project`, `--download-media` and `--force` are not passed through. Use
`export` directly for those. (`--download-media` still applies from `.env`; the
flag is only a per-run override.)
### `notify` — Push-notification settings and test
```bash
# Show the current settings
ai-chat-exporter notify
# Send a test push to confirm the topic works
ai-chat-exporter notify --test
```
Shows the resolved ntfy configuration and which machine name will appear in the
title. See [Getting notified](#getting-notified) for what a scheduled run sends.
Options: `--test`
### `prune` — Delete stale export files
```bash
# Preview what would be deleted
ai-chat-exporter prune --dry-run
# Delete (asks for confirmation; -y skips the prompt)
ai-chat-exporter prune
```
Deletes export files no longer referenced by the cache manifest — leftovers
from old folder layouts or fixed bugs that would otherwise double-sync into
Joplin. Refuses to run when the manifest is empty (e.g. right after
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
command separately verifies that every manifest entry's file exists on disk.
### `projects` — Discover ChatGPT project IDs
```bash
# List the projects your conversations belong to
ai-chat-exporter projects
# Also inspect conversations whose listing entry doesn't name a project
ai-chat-exporter projects --deep
# Write the discovered IDs straight into .env
ai-chat-exporter projects --write
```
`CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is
invisible to the listing pass — conversations that live *only* inside that
project are never fetched at all. This reports every project your conversations
belong to, marks the ones absent from `.env`, and prints a paste-ready line.
`--deep` fetches each conversation's detail when the listing doesn't name its
project: complete, but one request per conversation, so it's slow.
Options: `--deep`, `--write`
### `canary` — Check for provider API drift
```bash
ai-chat-exporter canary
ai-chat-exporter canary --provider chatgpt
```
The web providers are undocumented internal APIs that can change shape without
notice, and the failure mode is silent — a renamed field means content is
quietly dropped rather than an error being raised. The canary fetches one
listing page and one conversation per provider and asserts only the fields the
normalizer actually depends on.
Findings are `ERROR` (a load-bearing field is missing or mistyped — the parser
will break or silently lose data) or `WARN` (something unfamiliar appeared;
worth investigating, not necessarily broken). **Exits non-zero on any ERROR**,
so it can be scheduled or run in CI. Local providers have no remote schema and
are not probed.
Options: `--provider [chatgpt|claude|all]`
### `cache` — Manage the sync manifest
@@ -382,12 +805,12 @@ To force a full re-export: `ai-chat-exporter cache --clear` then re-run export.
## Troubleshooting
### `401 Unauthorized`
### `Authentication failed` (401, or 403 from Claude)
Your session token has expired.
- Run `ai-chat-exporter auth` to get a new token interactively
- Or manually copy a fresh cookie value into your `.env` file
Note: Claude's `sessionKey` is an opaque string — the only way to know it's expired is the 401 error. ChatGPT JWTs have an `exp` claim that the `doctor` command can decode and display.
Note: neither token's expiry can be read client-side. Claude's `sessionKey` is an opaque string, and claude.ai reports an invalid one as **403** "Invalid authorization" (`account_session_invalid`), not 401. ChatGPT's token is an encrypted JWE; `doctor` reads the `error` field of `/api/auth/session` instead. See [When Tokens Expire](#when-tokens-expire).
### `429 Rate Limited`
The tool automatically pauses, saves progress, and exits with a clear message showing how many conversations were exported vs remaining. Just re-run the same export command to resume — the cache picks up exactly where it left off.
@@ -421,7 +844,13 @@ Make sure you've added the project IDs to `CHATGPT_PROJECT_IDS` in your `.env`.
The provider's internal API may have changed. Run with `--debug`, sanitize the output (remove any personal content), and check the project's GitHub Issues for known fixes.
### Non-text content warnings
Images, code interpreter outputs, DALL-E generations, and Claude artifacts are not exported in v0.2.0. A WARNING is logged for each skipped item. See `FUTURE.md` for the roadmap.
Since v0.4.0, rich content is preserved as typed blocks in the export. ChatGPT voice transcripts render as text and audio assets as `📎 File attached` placeholders with size and duration metadata. Anything the extractor doesn't recognise renders as a visible `> ⚠️ Unsupported content` block naming the type and observed keys, *and* increments a counter in the post-export summary so you can tell whether real content is being silently skipped.
### Embedded images and media
Since v0.6.0, image attachments are downloaded by default into a `media/` folder beside each export and inlined as `![](media/…)`. Audio and other files download only with `EXPORTER_DOWNLOAD_MEDIA=all`. On the next `joplin` run these are uploaded as Joplin resources so they render inside the note. Assets that have expired on the provider's side (older AI-generated images often do) keep their text placeholder and are tallied as `media failed` in the run summary — they're never fatal to the export. Set `EXPORTER_DOWNLOAD_MEDIA=off` to skip downloads entirely.
### Tool output / hidden context collapsed
Since v0.6.0, content that was invisible in the ChatGPT web UI is collapsed by default to one-line placeholders like `> 🔧 Tool output — file_search (15.0 KB) — omitted`. This is ChatGPT's file-retrieval tool re-injecting your attached files on every run — it can be 90% of a conversation's bytes while containing none of the dialogue. Custom Instructions collapse the same way (`> ℹ️ Hidden context`). The post-export summary lists every collapsed message by origin with the total KB omitted. To keep everything, set `EXPORTER_HIDDEN_CONTENT=full` in `.env` or pass `--hidden-content full`.
### Empty export / all conversations skipped
No new or updated conversations since your last run. To verify: `ai-chat-exporter cache --show`. To force a full re-export: `ai-chat-exporter cache --clear`.
@@ -435,12 +864,12 @@ No new or updated conversations since your last run. To verify: `ai-chat-exporte
## Future Work
See `FUTURE.md` for planned features:
See `FUTURE.md` for the full roadmap. Current priorities:
- **v0.2.x** — `export --force` flag; `joplin --force` flag; per-conversation cache reset
- **v0.3.0** — Official API fallback: parse export ZIP files from ChatGPT/Claude settings
- **v0.4.0** — Rich content: images, artifacts, code interpreter output, extended thinking
- **v0.5.0** — Watch/scheduled mode; Obsidian vault output
- **A StartOS service** that centralises every machine's conversations into one
corpus and owns the Joplin connection, so each machine only has to upload
(`FUTURE.md` §8)
- **Splitting this README** into a short overview plus separate documents
---
+67
View File
@@ -0,0 +1,67 @@
#!/usr/bin/env bash
# ai-chat-exporter — run the exporter without activating a virtualenv.
#
# cd /path/to/AIChatExporter
# ./ai-chat-exporter export --provider all
#
# Creates .venv and installs dependencies on first run, so a fresh clone on a
# new machine needs no `python3 -m venv` / `source .venv/bin/activate` ceremony.
# Reinstalls automatically when pyproject.toml changes.
#
# Working directory is deliberately NOT changed: EXPORT_DIR, CACHE_DIR and .env
# discovery are all relative to your current directory, which is what lets the
# same checkout archive different machines into different places. Run it from
# the repo (see the warning below) unless you mean otherwise.
#
# Windows equivalent: ai-chat-exporter.cmd (same directory).
set -euo pipefail
DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
VENV="$DIR/.venv"
PY="$VENV/bin/python"
STAMP="$VENV/.deps-stamp"
# The cache manifest is read from ./cache by default. Running from somewhere
# else silently starts a *new, empty* archive rather than failing, which would
# re-export everything and orphan the existing Joplin notes. Warn, don't block —
# a deliberate second archive is a legitimate thing to want.
if [ "$PWD" != "$DIR" ] && [ -z "${AI_CHAT_EXPORTER_QUIET_CWD:-}" ]; then
echo "warning: running from $PWD, not $DIR" >&2
echo " cache/ and exports/ resolve against the current directory," >&2
echo " so this may start a separate archive. Set" >&2
echo " AI_CHAT_EXPORTER_QUIET_CWD=1 to silence this." >&2
fi
find_python() {
for candidate in python3 python; do
if command -v "$candidate" >/dev/null 2>&1; then
# Needs >=3.11 (pyproject requires-python).
if "$candidate" -c 'import sys; sys.exit(0 if sys.version_info >= (3, 11) else 1)' 2>/dev/null; then
command -v "$candidate"
return 0
fi
fi
done
return 1
}
if [ ! -x "$PY" ]; then
BOOTSTRAP_PY="$(find_python)" || {
echo "error: no python3 >= 3.11 found on PATH — install it and re-run." >&2
exit 1
}
echo "Creating virtualenv in $VENV …" >&2
"$BOOTSTRAP_PY" -m venv "$VENV"
rm -f "$STAMP"
fi
# Install (or refresh) dependencies when the venv is new or pyproject changed.
if [ ! -f "$STAMP" ] || [ "$DIR/pyproject.toml" -nt "$STAMP" ]; then
echo "Installing dependencies …" >&2
"$PY" -m pip install --quiet --upgrade pip
"$PY" -m pip install --quiet -e "$DIR"
touch "$STAMP"
fi
exec "$PY" -m src.main "$@"
+83
View File
@@ -0,0 +1,83 @@
@echo off
rem ai-chat-exporter.cmd - run the exporter without activating a virtualenv.
rem
rem Command Prompt (cmd.exe) - the documented way:
rem cd C:\path\to\AIChatExporter
rem ai-chat-exporter export --provider all
rem
rem cmd.exe searches the current directory before PATH and resolves the bare
rem name through PATHEXT (which includes .CMD), so no ".\" and no extension are
rem needed. The extensionless POSIX sibling is ignored: it is not in PATHEXT.
rem
rem PowerShell does NOT search the current directory, and ".\ai-chat-exporter"
rem there would resolve to the extensionless POSIX script, which PowerShell
rem cannot run. In PowerShell, name this file explicitly:
rem .\ai-chat-exporter.cmd export --provider all
rem
rem Creates .venv and installs dependencies on first run, so a fresh clone needs
rem no "python -m venv" / ".venv\Scripts\activate" ceremony. Reinstalls
rem automatically when pyproject.toml changes.
rem
rem The working directory is deliberately NOT changed - EXPORT_DIR, CACHE_DIR
rem and .env discovery are all relative to it. POSIX equivalent: ai-chat-exporter
setlocal enabledelayedexpansion
set "DIR=%~dp0"
if "%DIR:~-1%"=="\" set "DIR=%DIR:~0,-1%"
set "VENV=%DIR%\.venv"
set "PY=%VENV%\Scripts\python.exe"
set "STAMP=%VENV%\.deps-stamp"
rem See the POSIX script for why this warns rather than blocks: cache\ resolves
rem against the current directory, so the wrong one silently starts a second
rem archive instead of failing.
if /i not "%CD%"=="%DIR%" if "%AI_CHAT_EXPORTER_QUIET_CWD%"=="" (
echo warning: running from %CD%, not %DIR% 1>&2
echo cache\ and exports\ resolve against the current directory, 1>&2
echo so this may start a separate archive. Set 1>&2
echo AI_CHAT_EXPORTER_QUIET_CWD=1 to silence this. 1>&2
)
if not exist "%PY%" (
echo Creating virtualenv in %VENV% ... 1>&2
rem The py launcher is the reliable way to get a specific version; fall back
rem to whatever "python" is if it is not installed.
where py >nul 2>&1
if !errorlevel! equ 0 (
py -3 -m venv "%VENV%"
) else (
python -m venv "%VENV%"
)
if not exist "%PY%" (
echo error: could not create a virtualenv - install Python 3.11+ from 1>&2
echo python.org or the Microsoft Store, then re-run. 1>&2
exit /b 1
)
if exist "%STAMP%" del "%STAMP%"
)
rem Staleness check, in pure batch: %%~tF is the file's last-modified stamp, so
rem storing it and comparing strings needs no external process. The obvious
rem alternative - asking PowerShell to compare timestamps - costs ~1s of
rem interpreter startup on *every* command, which is a lot to pay to almost
rem always learn that nothing changed.
set "PYPROJ_TIME="
for %%F in ("%DIR%\pyproject.toml") do set "PYPROJ_TIME=%%~tF"
set "SAVED_TIME="
if exist "%STAMP%" set /p SAVED_TIME=<"%STAMP%"
if not "%SAVED_TIME%"=="%PYPROJ_TIME%" (
echo Installing dependencies ... 1>&2
"%PY%" -m pip install --quiet --upgrade pip
"%PY%" -m pip install --quiet -e "%DIR%"
if !errorlevel! neq 0 (
echo error: dependency installation failed. 1>&2
exit /b !errorlevel!
)
> "%STAMP%" echo !PYPROJ_TIME!
)
"%PY%" -m src.main %*
exit /b %errorlevel%
+2 -2
View File
@@ -4,8 +4,8 @@ build-backend = "setuptools.build_meta"
[project]
name = "ai-chat-exporter"
version = "0.2.1"
description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin"
version = "0.9.0"
description = "Archive ChatGPT, Claude, Claude Code and Codex conversation history to Markdown for personal backup in Joplin"
requires-python = ">=3.11"
dependencies = [
"requests==2.31.0",
+93
View File
@@ -0,0 +1,93 @@
<#
.SYNOPSIS
Register a Windows scheduled task that runs `ai-chat-exporter sync` daily.
.DESCRIPTION
The Windows counterpart to install-systemd-timer.sh. Creates a per-user task
(no admin rights needed) that runs the exporter from this repository, with
the working directory set to the repo so .env, cache\ and exports\ resolve
exactly as they do for an interactive run.
-Provider is repeatable. The CLI takes one provider per run, so each becomes
its own action, executed in order. List only the providers that work
unattended on this machine: a web provider whose session token has expired
fails the task every day, which trains you to ignore the failures you
actually want to notice.
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Provider all -Time 21:30
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
#>
[CmdletBinding()]
param(
[string[]]$Provider = @('all'),
[string]$Time = '09:00',
[string]$TaskName = 'AiChatExporterSync',
[switch]$Unregister
)
$ErrorActionPreference = 'Stop'
$repo = Split-Path -Parent $PSScriptRoot
$launcher = Join-Path $repo 'ai-chat-exporter.cmd'
if ($Unregister) {
Unregister-ScheduledTask -TaskName $TaskName -Confirm:$false -ErrorAction SilentlyContinue
Write-Host "Removed scheduled task '$TaskName'."
return
}
if (-not (Test-Path $launcher)) {
throw "Launcher not found at $launcher"
}
# A single action looping over the providers, rather than one action each.
# Task Scheduler runs multiple actions in order but reports only the last one's
# result, so a failure in an earlier provider would be invisible. The loop keeps
# going after a failure and propagates a non-zero exit code.
#
# /v:on and !RC! are required, not stylistic: cmd expands every %VAR% on a
# command line *before* running any of it, so "exit /b %RC%" would report the
# value RC had before the loop ever ran - i.e. always success. Delayed expansion
# reads it at the point of use.
$loop = ($Provider | ForEach-Object { "`"$launcher`" sync --provider $_ --joplin-optional || set RC=1" }) -join ' & '
$taskArgs = "/v:on /c set RC=0 & $loop & exit /b !RC!"
$actions = New-ScheduledTaskAction -Execute 'cmd.exe' `
-Argument $taskArgs `
-WorkingDirectory $repo
$trigger = New-ScheduledTaskTrigger -Daily -At $Time
# StartWhenAvailable is the counterpart of systemd's Persistent=true: a machine
# that was asleep at the scheduled time runs the archive when it wakes, rather
# than skipping the day entirely.
$settings = New-ScheduledTaskSettingsSet `
-StartWhenAvailable `
-DontStopIfGoingOnBatteries `
-AllowStartIfOnBatteries `
-ExecutionTimeLimit (New-TimeSpan -Hours 2)
Register-ScheduledTask -TaskName $TaskName `
-Action $actions `
-Trigger $trigger `
-Settings $settings `
-Description 'Export AI chat history and sync it to Joplin' `
-Force | Out-Null
Write-Host "Registered '$TaskName' - daily at $Time for: $($Provider -join ', ')"
Write-Host ''
Write-Host 'Command the task will run:'
Write-Host " cmd.exe $taskArgs"
Write-Host " (working directory: $repo)"
Write-Host ''
Write-Host 'Next steps:'
Write-Host " * Run it once now: Start-ScheduledTask -TaskName $TaskName"
Write-Host " * Check the result: Get-ScheduledTaskInfo -TaskName $TaskName"
Write-Host " * Read the log: Get-Content '$repo\cache\logs\exporter.log' -Tail 50"
Write-Host ' * The terms-of-service notice must have been acknowledged'
Write-Host ' interactively once on this machine, or the task exits 1.'
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Install a systemd *user* timer that runs `ai-chat-exporter sync` on a schedule.
#
# ./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
# ./scheduling/install-systemd-timer.sh --provider all --time 21:30
# ./scheduling/install-systemd-timer.sh --uninstall
#
# User units (not system units) are the right scope: the archive is per-user,
# the .env holds that user's session tokens, and Joplin runs in their session.
#
# --provider is repeatable. The CLI takes one provider per run, so each one
# becomes its own ExecStart line, executed in order. Prefer listing only the
# providers that actually work unattended on this machine — a web provider with
# an expired session token will fail the unit every single day, which trains you
# to ignore the failure you actually want to see.
set -euo pipefail
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
UNIT_DIR="$HOME/.config/systemd/user"
NAME="aichat-sync"
TIME="09:00"
PROVIDERS=()
UNINSTALL=0
while [ $# -gt 0 ]; do
case "$1" in
--provider) PROVIDERS+=("$2"); shift 2 ;;
--time) TIME="$2"; shift 2 ;;
--name) NAME="$2"; shift 2 ;;
--uninstall) UNINSTALL=1; shift ;;
-h|--help) sed -n '2,20p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
if [ "$UNINSTALL" -eq 1 ]; then
systemctl --user disable --now "$NAME.timer" 2>/dev/null || true
rm -f "$UNIT_DIR/$NAME.timer" "$UNIT_DIR/$NAME.service"
systemctl --user daemon-reload
echo "Removed $NAME.timer and $NAME.service."
exit 0
fi
[ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all")
for f in "$REPO/ai-chat-exporter" "$REPO/scheduling/run-sync.sh"; do
if [ ! -x "$f" ]; then
echo "error: $f is missing or not executable." >&2
exit 1
fi
done
mkdir -p "$UNIT_DIR"
# WorkingDirectory is the point of the whole unit: .env, cache/ and exports/ all
# resolve against it, so the scheduled run writes to the same archive an
# interactive run from this directory would.
{
echo "[Unit]"
echo "Description=AI chat archive sync"
echo "Documentation=file://$REPO/README.md"
echo "After=network-online.target"
echo "Wants=network-online.target"
echo
echo "[Service]"
echo "Type=oneshot"
echo "WorkingDirectory=$REPO"
echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1"
echo "Environment=AICHAT_SYNC_UNIT=$NAME"
# run-sync.sh attempts every provider even after one fails, and pushes a
# FAILED notification for any run that died without sending its own.
echo "ExecStart=$REPO/scheduling/run-sync.sh ${PROVIDERS[*]}"
} > "$UNIT_DIR/$NAME.service"
# Persistent=true runs a missed schedule at the next boot — the machine being
# off at 09:00 should delay the archive, not skip it. RandomizedDelaySec keeps
# the web providers from being hit at exactly the same second every day.
cat > "$UNIT_DIR/$NAME.timer" <<EOF
[Unit]
Description=Run the AI chat archive sync daily
[Timer]
OnCalendar=*-*-* $TIME:00
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now "$NAME.timer"
echo "Installed $NAME.timer — daily at $TIME for: ${PROVIDERS[*]}"
echo
systemctl --user list-timers "$NAME.timer" --no-pager || true
cat <<EOF
Next steps:
• Timers only run while you have a session. To archive when logged out:
loginctl enable-linger $USER
• Run it once by hand to confirm:
systemctl --user start $NAME.service
• Read the log:
journalctl --user -u $NAME.service -n 50
• The terms-of-service notice must have been acknowledged interactively at
least once on this machine, or the unit exits 1 with an explanation.
EOF
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env bash
# Run `ai-chat-exporter sync` once per provider — the ExecStart of the systemd
# unit that install-systemd-timer.sh writes.
#
# ./scheduling/run-sync.sh claude-code codex
#
# Every provider is attempted even after one fails, and the exit code is
# non-zero if any failed — one ExecStart per provider would stop at the first.
#
# The app pushes its own ntfy result, but only from the end of a run it
# finished. A crash, a non-zero exit before the sync starts (the terms-of-service
# gate, a cache error) or a launcher that can't build its venv sends nothing, and
# because each provider pushes separately, the providers that did succeed still
# send "OK" — so a broken one looks like a quiet day. This script pushes a FAILED
# notification for any run that exited non-zero without the app having reported
# it. (Its "Sync completed with failures" banner prints right after its push.)
#
# The push carries the provider, the exit code and, for a crash, the exception's
# class name — never its message. Same counts-only rule as src/notify.py: on a
# public ntfy topic anyone who guesses the name can read it, and exception text
# can carry conversation titles. The full traceback is in the journal.
set -uo pipefail
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
LAUNCHER="$REPO/ai-chat-exporter"
# NTFY_* as the app resolves them: the environment wins, then .env.
env_value() {
local name=$1 value=${!1:-}
if [ -z "$value" ] && [ -f "$REPO/.env" ]; then
value=$(sed -n "s/^[[:space:]]*$name[[:space:]]*=[[:space:]]*//p" "$REPO/.env" | tail -n 1)
value=${value%%[[:space:]]#*}
value=${value%"${value##*[![:space:]]}"}
value=${value#[\"\']}
value=${value%[\"\']}
fi
printf '%s' "$value"
}
push_failure() {
local body=$1 topic server token policy
topic=$(env_value NTFY_TOPIC)
policy=$(env_value NTFY_NOTIFY | tr '[:upper:]' '[:lower:]')
if [ -z "$topic" ] || [ "$policy" = "off" ]; then
return 0
fi
server=$(env_value NTFY_SERVER)
server=${server:-https://ntfy.sh}
token=$(env_value NTFY_TOKEN)
local args=(-fsS --max-time 15 -o /dev/null
-H "Title: AI archive FAILED - $(hostname -s)"
-H "Tags: rotating_light" -H "Priority: high"
--data-binary "$body")
[ -n "$token" ] && args+=(-H "Authorization: Bearer $token")
curl "${args[@]}" "${server%/}/$topic" \
|| echo "run-sync: could not send the failure notification" >&2
}
[ $# -eq 0 ] && set -- all
rc=0
for provider in "$@"; do
out=$(mktemp)
"$LAUNCHER" sync --provider "$provider" --joplin-optional 2>&1 | tee "$out"
status=${PIPESTATUS[0]}
if [ "$status" -ne 0 ]; then
rc=1
if ! grep -q "Sync completed with failures" "$out"; then
crash=$(grep -oE '^[A-Za-z_][A-Za-z0-9_.]*(Error|Exception)\b' "$out" | tail -n 1)
if [ -n "$crash" ]; then
reason="crashed ($crash)"
else
reason="exited $status before reporting a result"
fi
push_failure "$provider: $reason
journalctl --user -u ${AICHAT_SYNC_UNIT:-aichat-sync} -n 100"
fi
fi
rm -f "$out"
done
exit "$rc"
+451
View File
@@ -0,0 +1,451 @@
"""Typed content blocks for normalized messages.
Providers produce ordered lists of blocks; exporters render them. Living outside
``src/providers/`` deliberately — blocks are a separate concern from extraction
or rendering, shared by both layers.
Block dicts always have ``type`` set to one of the BLOCK_TYPE_* constants.
Construct via the ``make_*`` helpers; never build dicts by hand. The ``unknown``
block constructor REQUIRES a corresponding WARNING log + ``LossReport`` tally
at the call site — see plan §Data-loss visibility.
"""
import json
from pathlib import Path
from typing import Any
BLOCK_TYPE_TEXT = "text"
BLOCK_TYPE_CODE = "code"
BLOCK_TYPE_THINKING = "thinking"
BLOCK_TYPE_TOOL_USE = "tool_use"
BLOCK_TYPE_TOOL_RESULT = "tool_result"
BLOCK_TYPE_CITATION = "citation"
BLOCK_TYPE_IMAGE_PLACEHOLDER = "image_placeholder"
BLOCK_TYPE_FILE_PLACEHOLDER = "file_placeholder"
BLOCK_TYPE_UNKNOWN = "unknown"
BLOCK_TYPE_HIDDEN_CONTEXT_MARKER = "hidden_context_marker"
BLOCK_TYPE_COLLAPSED = "collapsed"
BLOCK_TYPE_SUBAGENT = "subagent"
COLLAPSED_KIND_TOOL_DUMP = "tool_dump"
COLLAPSED_KIND_HIDDEN_CONTEXT = "hidden_context"
UNKNOWN_REASON_UNKNOWN_TYPE = "unknown_type"
UNKNOWN_REASON_EXTRACTION_FAILED = "extraction_failed"
UNKNOWN_REASON_ALL_BLOCKS_FAILED = "all_blocks_failed"
UNKNOWN_REASON_UNKNOWN_FIELD_IN_KNOWN_TYPE = "unknown_field_in_known_type"
_OBSERVED_KEYS_LIMIT = 10
# Role labels for the turns rendered inside a folded subagent <details> block.
_SUBAGENT_ROLE_LABELS = {
"user": "🧑 Human",
"assistant": "🤖 Assistant",
"tool": "🔧 Tool",
}
# ---------------------------------------------------------------------------
# Constructors
# ---------------------------------------------------------------------------
def make_text_block(text: str) -> dict | None:
"""Return a text block, or None if the text is empty/whitespace-only.
Returning None lets callers do ``if block: blocks.append(block)`` and prune
empty blocks at construction time. See plan §Finalizing the message dict.
"""
if not isinstance(text, str) or not text.strip():
return None
return {"type": BLOCK_TYPE_TEXT, "text": text}
def make_code_block(code: str, language: str = "") -> dict | None:
"""Return a code block, or None if code is empty."""
if not isinstance(code, str) or not code.strip():
return None
return {"type": BLOCK_TYPE_CODE, "language": language or "", "code": code}
def make_thinking_block(text: str) -> dict | None:
"""Return a thinking block, or None if empty."""
if not isinstance(text, str) or not text.strip():
return None
return {"type": BLOCK_TYPE_THINKING, "text": text}
def make_tool_use_block(name: str, input_data: Any, tool_id: str | None = None) -> dict:
"""Return a tool_use block.
Always returns a block (no None) — tool calls are meaningful even with
empty inputs.
"""
return {
"type": BLOCK_TYPE_TOOL_USE,
"name": name or "",
"input": input_data if input_data is not None else {},
"tool_id": tool_id,
}
def make_tool_result_block(
output: str,
tool_name: str | None = None,
is_error: bool = False,
summary: str | None = None,
) -> dict:
"""Return a tool_result block.
``summary`` is an optional short human label rendered between header and
fence (e.g. ChatGPT's ``metadata.reasoning_title`` for execution_output).
"""
return {
"type": BLOCK_TYPE_TOOL_RESULT,
"tool_name": tool_name,
"output": output if isinstance(output, str) else str(output),
"is_error": bool(is_error),
"summary": summary,
}
def make_citation_block(
url: str,
title: str | None = None,
snippet: str | None = None,
) -> dict | None:
if not url:
return None
return {
"type": BLOCK_TYPE_CITATION,
"url": url,
"title": title,
"snippet": snippet,
}
def make_image_placeholder(
ref: str,
source: str = "unknown",
mime: str | None = None,
) -> dict:
"""source ∈ {'user_upload', 'model_generated', 'unknown'}."""
return {
"type": BLOCK_TYPE_IMAGE_PLACEHOLDER,
"ref": ref or "",
"source": source,
"mime": mime,
}
def make_file_placeholder(
ref: str,
filename: str | None = None,
mime: str | None = None,
size_bytes: int | None = None,
duration_seconds: float | None = None,
) -> dict:
return {
"type": BLOCK_TYPE_FILE_PLACEHOLDER,
"ref": ref or "",
"filename": filename,
"mime": mime,
"size_bytes": size_bytes,
"duration_seconds": duration_seconds,
}
def make_unknown_block(
raw_type: str,
observed_keys: list[str] | None = None,
reason: str = UNKNOWN_REASON_UNKNOWN_TYPE,
summary: str | None = None,
) -> dict:
"""Construct an unknown block.
Every call site MUST also emit a WARNING log and increment a LossReport
tally — see plan §Data-loss visibility. The block surfaces the loss at
read time; the WARNING surfaces it at run time. Both signals matter.
"""
keys = list(observed_keys or [])[:_OBSERVED_KEYS_LIMIT]
return {
"type": BLOCK_TYPE_UNKNOWN,
"raw_type": raw_type,
"observed_keys": keys,
"reason": reason,
"summary": summary,
}
def make_collapsed_block(
origin: str,
content_type: str,
size_bytes: int,
kind: str,
) -> dict:
"""One-line stand-in for a message omitted by the EXPORTER_HIDDEN_CONTENT policy.
``kind`` ∈ {COLLAPSED_KIND_TOOL_DUMP, COLLAPSED_KIND_HIDDEN_CONTEXT}.
Unlike ``unknown`` blocks (unexpected loss), a collapsed block records an
intentional policy decision — the content was invisible in the provider's
web UI and the user chose not to keep it. Every call site MUST tally via
``LossReport.record_collapsed`` so the omission stays visible in the
post-export summary.
"""
return {
"type": BLOCK_TYPE_COLLAPSED,
"origin": origin or "",
"content_type": content_type or "",
"size_bytes": int(size_bytes or 0),
"kind": kind,
}
def make_subagent_block(
agent_type: str,
description: str,
messages: list[dict],
) -> dict:
"""A folded Claude Code subagent (Task-tool) transcript.
Claude Code stores subagent transcripts as separate ``subagents/*.jsonl``
files; the provider folds each one into its parent session at the point the
``Task``/``Agent`` tool spawned it. ``messages`` is a normal list of
``{role, blocks}`` dicts (the same shape as a conversation's messages),
extracted under the same hidden-content policy as the main dialogue — so the
subagent's own tool traffic is already collapsed inside these blocks.
Rendered as a collapsible ``<details>`` element so the main conversation
stays readable while the deliverable is preserved (and stays searchable).
"""
return {
"type": BLOCK_TYPE_SUBAGENT,
"agent_type": agent_type or "",
"description": description or "",
"messages": messages or [],
}
def make_hidden_context_marker(content_type: str) -> dict:
"""A short prepend block that flags the surrounding message as hidden context.
Driven by the ``metadata.is_visually_hidden_from_conversation`` flag, not by
content_type matching. The marker tells the reader "this message is
hidden in the source UI; we're showing it here for archival fidelity."
"""
return {
"type": BLOCK_TYPE_HIDDEN_CONTEXT_MARKER,
"content_type": content_type or "",
}
# ---------------------------------------------------------------------------
# Rendering
# ---------------------------------------------------------------------------
def render_blocks_to_markdown(blocks: list[dict]) -> str:
"""Render an ordered list of blocks to a single Markdown string.
Blocks are joined with one blank line between them. Pure function; no I/O.
"""
if not blocks:
return ""
rendered: list[str] = []
for block in blocks:
chunk = _render_one(block)
if chunk:
rendered.append(chunk)
return "\n\n".join(rendered)
def _render_one(block: dict) -> str:
btype = block.get("type", "")
if btype == BLOCK_TYPE_TEXT:
return block.get("text", "")
if btype == BLOCK_TYPE_CODE:
lang = block.get("language") or ""
code = block.get("code", "")
fence = _safe_fence(code)
return f"{fence}{lang}\n{code}\n{fence}"
if btype == BLOCK_TYPE_THINKING:
text = block.get("text", "")
quoted = _blockquote_prefix(text)
return f"**💭 Reasoning**\n\n{quoted}"
if btype == BLOCK_TYPE_TOOL_USE:
name = block.get("name", "")
input_data = block.get("input", {})
body_json = json.dumps(input_data, indent=2, sort_keys=False, default=str, ensure_ascii=False)
fence = _safe_fence(body_json)
body = f"{fence}json\n{body_json}\n{fence}"
quoted = _blockquote_prefix(f"🔧 **Tool: {name}**\n{body}")
return quoted
if btype == BLOCK_TYPE_TOOL_RESULT:
output = block.get("output", "")
is_error = bool(block.get("is_error"))
tool_name = block.get("tool_name") or ""
summary = block.get("summary") or ""
icon = "❌" if is_error else "📤"
label = "Result (error)" if is_error else "Result"
if tool_name:
header = f"{icon} **{label}: {tool_name}**"
else:
header = f"{icon} **{label}**"
fence = _safe_fence(output)
body = f"{fence}\n{output}\n{fence}"
if summary:
inner = f"{header}\n*{summary}*\n{body}"
else:
inner = f"{header}\n{body}"
return _blockquote_prefix(inner)
if btype == BLOCK_TYPE_CITATION:
url = block.get("url", "")
title = block.get("title") or url
return f"[{title}]({url})"
if btype == BLOCK_TYPE_IMAGE_PLACEHOLDER:
ref = block.get("ref", "")
source = block.get("source", "unknown")
mime = block.get("mime")
local_path = block.get("local_path")
if local_path:
# Downloaded — render as a real inline image.
return f"![{source or 'image'}]({local_path})"
meta_parts = [source] if source else []
if mime:
meta_parts.append(mime)
meta_parts.append("content not preserved in this export")
meta = ", ".join(meta_parts)
return f"> 🖼️ **Image attached** — `{ref}` ({meta})"
if btype == BLOCK_TYPE_FILE_PLACEHOLDER:
ref = block.get("ref", "")
filename = block.get("filename")
label = filename or ref
mime = block.get("mime")
size_bytes = block.get("size_bytes")
duration = block.get("duration_seconds")
local_path = block.get("local_path")
meta_parts: list[str] = []
if mime:
meta_parts.append(mime)
size_label = _format_size(size_bytes)
if size_label:
meta_parts.append(size_label)
if isinstance(duration, (int, float)) and duration > 0:
meta_parts.append(f"{duration:.2f}s")
if local_path:
# Downloaded — render as a link to the local copy.
meta = ", ".join(meta_parts) if meta_parts else ""
meta_str = f" ({meta})" if meta else ""
return f"> 📎 **File attached** — [{Path(local_path).name}]({local_path}){meta_str}"
meta_parts.append("content not preserved in this export")
meta = ", ".join(meta_parts)
return f"> 📎 **File attached** — `{label}` ({meta})"
if btype == BLOCK_TYPE_UNKNOWN:
raw_type = block.get("raw_type", "?")
reason = block.get("reason", UNKNOWN_REASON_UNKNOWN_TYPE)
keys = block.get("observed_keys") or []
summary = block.get("summary")
first_line = f"⚠️ **Unsupported content** — type `{raw_type}` ({reason})"
lines = [first_line]
if summary:
lines.append(summary)
if keys:
keys_str = ", ".join(f"`{k}`" for k in keys)
lines.append(f"Keys observed: {keys_str}")
return _blockquote_prefix("\n".join(lines))
if btype == BLOCK_TYPE_SUBAGENT:
agent_type = block.get("agent_type") or "subagent"
description = block.get("description") or ""
summary = f"🤖 Subagent: {agent_type}"
if description:
summary += f" — {description}"
parts = ["<details>", f"<summary>{summary}</summary>", ""]
for msg in block.get("messages") or []:
body = render_blocks_to_markdown(msg.get("blocks") or [])
if not body.strip():
continue
role = msg.get("role", "assistant")
label = _SUBAGENT_ROLE_LABELS.get(role, f"💬 {role.capitalize()}")
parts.append(f"**{label}**")
parts.append("")
parts.append(body)
parts.append("")
parts.append("</details>")
return "\n".join(parts)
if btype == BLOCK_TYPE_HIDDEN_CONTEXT_MARKER:
ctype = block.get("content_type", "")
return f"> ℹ️ **Hidden context** — `{ctype}`"
if btype == BLOCK_TYPE_COLLAPSED:
origin = block.get("origin", "")
kind = block.get("kind", "")
if kind == COLLAPSED_KIND_HIDDEN_CONTEXT:
icon, label = "ℹ️", "Hidden context"
else:
icon, label = "🔧", "Tool output"
size_label = _format_size(block.get("size_bytes"))
size_part = f" ({size_label})" if size_label else ""
return (
f"> {icon} **{label}** — `{origin}`{size_part} — omitted "
"(EXPORTER_HIDDEN_CONTENT=full to keep)"
)
# Defensive: a block of unrecognised local type (shouldn't happen if
# constructors are used). Render as visible warning rather than dropping.
return f"> ⚠️ **Internal: unrecognised block type** — `{btype}`"
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _format_size(size_bytes: Any) -> str:
"""Human-readable size ('24.1 KB', '1.05 MB'), or '' when absent/zero."""
if not isinstance(size_bytes, int) or size_bytes <= 0:
return ""
kb = size_bytes / 1024
return f"{kb:.1f} KB" if kb < 1024 else f"{kb / 1024:.2f} MB"
def _safe_fence(text: str) -> str:
"""Return a backtick fence longer than the longest run of backticks in ``text``.
CommonMark requires the closing fence to be at least as long as the opening
fence. Picking N+1 (where N = longest run in content) ensures the content's
own backticks are inert. Minimum is 3.
Verified live against Joplin during planning — see plan
§Backtick-corruption defense.
"""
if not isinstance(text, str):
return "```"
longest_run = 0
current_run = 0
for ch in text:
if ch == "`":
current_run += 1
if current_run > longest_run:
longest_run = current_run
else:
current_run = 0
fence_len = max(3, longest_run + 1)
return "`" * fence_len
def _blockquote_prefix(text: str) -> str:
"""Prefix every line of ``text`` with ``> `` so the whole block renders as a quote.
Empty source lines become ``>`` (no trailing space) so blockquote continuity
is preserved without trailing-whitespace noise.
"""
if not isinstance(text, str):
return ""
out_lines: list[str] = []
for line in text.split("\n"):
if line == "":
out_lines.append(">")
else:
out_lines.append(f"> {line}")
return "\n".join(out_lines)
+100 -3
View File
@@ -84,27 +84,70 @@ class Cache:
if provider not in self._data:
self._data[provider] = {}
self._data[provider][conv_id] = {
# Preserve the Joplin link across re-exports. Without this, re-rendering
# a conversation (e.g. a forced re-export to pick up new formatting)
# would drop joplin_note_id, and the next sync would create a duplicate
# note instead of updating the existing one. exported_at is refreshed,
# so get_joplin_pending() still flags the note for an update.
prev = self._data[provider].get(conv_id) or {}
entry = {
"title": metadata.get("title", ""),
"project": metadata.get("project"),
"created_at": metadata.get("created_at", ""),
"updated_at": metadata.get("updated_at", ""),
"exported_at": datetime.now(tz=timezone.utc).isoformat(),
"file_path": metadata.get("file_path", ""),
}
for carried in ("joplin_note_id", "joplin_synced_at", "joplin_resources"):
if carried in prev:
entry[carried] = prev[carried]
self._data[provider][conv_id] = entry
self._data["last_run"] = datetime.now(tz=timezone.utc).isoformat()
self._save()
def get_new_or_updated(self, provider: str, conversations: list[dict]) -> list[dict]:
def get_new_or_updated(
self,
provider: str,
conversations: list[dict],
force: bool = False,
campaign_at: str | None = None,
) -> list[dict]:
"""Filter a conversation list to only new or updated conversations.
Args:
provider: "chatgpt" or "claude"
conversations: List of raw conversation dicts from the provider.
Each must have an ``id``/``uuid`` and ``updated_at``/``update_time``.
force: When True, return conversations to re-render regardless of
cache freshness — used by ``export --force``.
campaign_at: In force mode, exclude conversations already re-exported
at/after this ISO timestamp (the re-render campaign start), so
each run advances and finished providers do nothing.
Returns:
Subset that needs to be exported.
Subset that needs to be exported, ordered oldest-export-first in
force mode.
"""
if force:
# Re-export for a re-render campaign, ordered least-recently-exported
# first (never-exported sorts first). When ``campaign_at`` is given,
# conversations already re-exported during this campaign
# (exported_at >= campaign_at) are excluded — so finished providers
# do no work and the remaining count shrinks to zero. Each capped run
# re-renders the next-oldest batch. Conversations without an id drop.
entries = self._data.get(provider, {})
def _exported_at(conv: dict) -> str:
conv_id = conv.get("id") or conv.get("uuid", "")
entry = entries.get(conv_id) or {}
return entry.get("exported_at") or ""
candidates = [c for c in conversations if c.get("id") or c.get("uuid")]
if campaign_at is not None:
candidates = [c for c in candidates if _exported_at(c) < campaign_at]
return sorted(candidates, key=_exported_at)
result = []
for conv in conversations:
conv_id = conv.get("id") or conv.get("uuid", "")
@@ -172,6 +215,22 @@ class Cache:
entry["joplin_synced_at"] = datetime.now(tz=timezone.utc).isoformat()
self._save()
def get_joplin_resources(self, provider: str, conv_id: str) -> dict[str, str]:
"""Return the {relative_media_path: joplin_resource_id} map for a conversation."""
entry = self._data.get(provider, {}).get(conv_id)
if not isinstance(entry, dict):
return {}
resources = entry.get("joplin_resources")
return dict(resources) if isinstance(resources, dict) else {}
def set_joplin_resources(self, provider: str, conv_id: str, resources: dict[str, str]) -> None:
"""Persist the {relative_media_path: joplin_resource_id} map for a conversation."""
entry = self._data.get(provider, {}).get(conv_id)
if entry is None:
return
entry["joplin_resources"] = dict(resources)
self._save()
def get_joplin_pending(self, provider: str) -> list[tuple[str, dict]]:
"""Return (conv_id, entry) pairs that need to be synced to Joplin.
@@ -209,10 +268,48 @@ class Cache:
return pending
def all_file_paths(self) -> set[str]:
"""Return every non-empty ``file_path`` in the manifest, across providers.
Used by ``prune`` (files on disk but not here are stale) and by the
doctor integrity check (files here but not on disk are missing).
"""
paths: set[str] = set()
for key, entries in self._data.items():
if not isinstance(entries, dict) or key in (
"version",
"last_run",
"tos_acknowledged_at",
):
continue
for entry in entries.values():
if isinstance(entry, dict) and entry.get("file_path"):
paths.add(entry["file_path"])
return paths
def last_run(self) -> str | None:
"""Return the ISO8601 timestamp of the last export run, or None."""
return self._data.get("last_run")
def exported_at(self, provider: str, conv_id: str) -> str:
"""Return the recorded ``exported_at`` for a conversation, or '' if absent."""
entry = self._data.get(provider, {}).get(conv_id)
return entry.get("exported_at", "") if isinstance(entry, dict) else ""
# Force re-render campaign marker (top-level scalar; skipped by stats/clear,
# which only act on dict-valued provider keys).
def get_force_campaign(self) -> str | None:
"""Return the active force-re-render campaign start timestamp, or None."""
return self._data.get("force_campaign_at")
def set_force_campaign(self, started_at: str) -> None:
self._data["force_campaign_at"] = started_at
self._save()
def clear_force_campaign(self) -> None:
if self._data.pop("force_campaign_at", None) is not None:
self._save()
# ------------------------------------------------------------------
# Private helpers
# ------------------------------------------------------------------
+75 -1
View File
@@ -20,6 +20,12 @@ _CLAUDE_PLACEHOLDER = ""
# Valid OUTPUT_STRUCTURE values
VALID_STRUCTURES = {"provider/project/year", "provider/project", "provider/year"}
# Valid EXPORTER_HIDDEN_CONTENT values (see src/providers/base.py)
VALID_HIDDEN_CONTENT = {"full", "placeholder", "omit"}
# Valid EXPORTER_DOWNLOAD_MEDIA values (see src/media.py)
VALID_DOWNLOAD_MEDIA = {"images", "all", "off"}
class ConfigError(Exception):
"""Raised when required configuration is missing or invalid."""
@@ -43,6 +49,16 @@ class Config:
# Joplin local REST API settings (Web Clipper service)
joplin_api_token: str | None = None
joplin_api_url: str = "http://localhost:41184"
# Policy for content invisible in the provider web UI (retrieval dumps,
# hidden context): full | placeholder | omit
hidden_content: str = "placeholder"
# Session cap: max conversations downloaded per export run (None = unlimited).
# Every run is resumable, so a capped run just continues next time.
max_conversations: int | None = None
# Seconds between consecutive API requests (politeness pacing; 0 disables)
request_delay: float = 1.0
# Asset download policy: images | all | off
download_media: str = "images"
def load_config() -> Config:
@@ -67,6 +83,12 @@ def load_config() -> Config:
joplin_token = os.getenv("JOPLIN_API_TOKEN", "").strip() or None
joplin_url = os.getenv("JOPLIN_API_URL", "http://localhost:41184").strip()
hidden_content = os.getenv("EXPORTER_HIDDEN_CONTENT", "").strip().lower() or "placeholder"
max_conversations_raw = os.getenv("MAX_CONVERSATIONS_PER_RUN", "").strip()
request_delay_raw = os.getenv("REQUEST_DELAY", "").strip()
download_media = os.getenv("EXPORTER_DOWNLOAD_MEDIA", "").strip().lower() or "images"
# Parse CHATGPT_PROJECT_IDS — comma-separated list of gizmo IDs (g-p-xxx)
_project_ids_raw = os.getenv("CHATGPT_PROJECT_IDS", "").strip()
chatgpt_project_ids = [
@@ -90,6 +112,46 @@ def load_config() -> Config:
f"Must be one of: {', '.join(sorted(VALID_STRUCTURES))}"
)
# Validate hidden-content policy
if hidden_content not in VALID_HIDDEN_CONTENT:
errors.append(
f"EXPORTER_HIDDEN_CONTENT '{hidden_content}' is invalid. "
f"Must be one of: {', '.join(sorted(VALID_HIDDEN_CONTENT))}"
)
# Validate media download policy
if download_media not in VALID_DOWNLOAD_MEDIA:
errors.append(
f"EXPORTER_DOWNLOAD_MEDIA '{download_media}' is invalid. "
f"Must be one of: {', '.join(sorted(VALID_DOWNLOAD_MEDIA))}"
)
# Validate session cap
max_conversations: int | None = None
if max_conversations_raw:
try:
max_conversations = int(max_conversations_raw)
except ValueError:
errors.append(
f"MAX_CONVERSATIONS_PER_RUN '{max_conversations_raw}' is not an integer."
)
else:
if max_conversations < 1:
errors.append(
f"MAX_CONVERSATIONS_PER_RUN must be at least 1 (got {max_conversations})."
)
# Validate request pacing
request_delay = 1.0
if request_delay_raw:
try:
request_delay = float(request_delay_raw)
except ValueError:
errors.append(f"REQUEST_DELAY '{request_delay_raw}' is not a number.")
else:
if request_delay < 0:
errors.append(f"REQUEST_DELAY must be >= 0 (got {request_delay}).")
# Validate and decode ChatGPT JWT
chatgpt_expiry: datetime | None = None
if chatgpt_token:
@@ -139,6 +201,10 @@ def load_config() -> Config:
chatgpt_project_ids=chatgpt_project_ids,
joplin_api_token=joplin_token,
joplin_api_url=joplin_url,
hidden_content=hidden_content,
max_conversations=max_conversations,
request_delay=request_delay,
download_media=download_media,
)
_log_startup_summary(config)
@@ -223,7 +289,11 @@ def _log_startup_summary(cfg: Config) -> None:
"Joplin: %s | "
"export_dir=%s | "
"structure=%s | "
"cache_dir=%s",
"cache_dir=%s | "
"hidden_content=%s | "
"max_conversations=%s | "
"request_delay=%.1fs | "
"download_media=%s",
chatgpt_status,
claude_status,
len(cfg.chatgpt_project_ids),
@@ -231,4 +301,8 @@ def _log_startup_summary(cfg: Config) -> None:
cfg.export_dir,
cfg.output_structure,
cfg.cache_dir,
cfg.hidden_content,
cfg.max_conversations if cfg.max_conversations is not None else "unlimited",
cfg.request_delay,
cfg.download_media,
)
+12 -3
View File
@@ -6,6 +6,7 @@ from datetime import datetime, timezone
from pathlib import Path
from typing import Any
from src.blocks import render_blocks_to_markdown
from src.utils import build_export_path, generate_filename
logger = logging.getLogger(__name__)
@@ -15,6 +16,7 @@ _ROLE_LABELS = {
"user": ("🧑 Human", "user"),
"assistant": ("🤖 Assistant", "assistant"),
"system": ("⚙️ System", "system"),
"tool": ("🔧 Tool", "tool"),
}
@@ -125,10 +127,17 @@ class MarkdownExporter:
# Messages
for msg in messages:
role = msg.get("role", "user")
content = msg.get("content", "")
blocks = msg.get("blocks") or []
timestamp = msg.get("timestamp")
if not content or not content.strip():
# Prefer rendering from blocks (v0.4.0+). Backward-compat fallback:
# if blocks is missing/empty AND content exists, render content as-is.
if blocks:
body = render_blocks_to_markdown(blocks)
else:
body = msg.get("content", "") or ""
if not body or not body.strip():
logger.warning(
"[markdown] Skipping empty/whitespace message in conversation %s",
conv_id[:8],
@@ -143,7 +152,7 @@ class MarkdownExporter:
else:
lines.append("")
lines.append(content)
lines.append(body)
lines.append("")
lines.append("---")
lines.append("")
+156 -28
View File
@@ -1,7 +1,10 @@
"""Joplin Data API client for importing notes into Joplin desktop."""
import json
import logging
import os
import re
from pathlib import Path
from typing import Any
import requests
@@ -32,8 +35,8 @@ class JoplinClient:
def __init__(self, base_url: str, token: str) -> None:
self._base_url = base_url.rstrip("/")
self._token = token
# In-memory cache of notebook title → ID to avoid repeated GET /folders
self._notebook_cache: dict[str, str] = {}
# In-memory cache: (parent_id | None, title) → folder ID
self._notebook_cache: dict[tuple[str | None, str], str] = {}
self._notebooks_loaded = False
logger.debug("[joplin] Client initialised with base_url=%s", self._base_url)
@@ -89,13 +92,13 @@ class JoplinClient:
"""Return all Joplin notebooks (folders), handling pagination.
Returns:
List of folder dicts with at least ``id`` and ``title`` keys.
List of folder dicts with at least ``id``, ``title``, and ``parent_id`` keys.
"""
results: list[dict] = []
page = 1
while True:
logger.debug("[joplin] GET /folders page=%d", page)
resp = self._get("/folders", params={"page": page, "fields": "id,title"})
resp = self._get("/folders", params={"page": page, "fields": "id,title,parent_id"})
items = resp.get("items", [])
results.extend(items)
logger.debug("[joplin] /folders page=%d → %d items, has_more=%s", page, len(items), resp.get("has_more"))
@@ -104,11 +107,12 @@ class JoplinClient:
page += 1
return results
def get_or_create_notebook(self, title: str) -> str:
"""Return the Joplin folder ID for ``title``, creating it if needed.
def get_or_create_notebook(self, title: str, parent_id: str | None = None) -> str:
"""Return the Joplin folder ID for ``title`` under ``parent_id``, creating if needed.
Args:
title: Notebook display name (e.g. "ChatGPT - My Project").
title: Notebook display name.
parent_id: ID of the parent folder, or None for a root notebook.
Returns:
Joplin folder ID string.
@@ -116,19 +120,40 @@ class JoplinClient:
if not self._notebooks_loaded:
self._load_notebook_cache()
if title in self._notebook_cache:
folder_id = self._notebook_cache[title]
logger.debug("[joplin] Notebook cache hit: %r → %s", title, folder_id)
key = (parent_id, title)
if key in self._notebook_cache:
folder_id = self._notebook_cache[key]
logger.debug("[joplin] Notebook cache hit: %r (parent=%s) → %s", title, parent_id, folder_id)
return folder_id
# Not found — create it
logger.info("[joplin] Creating notebook: %r", title)
resp = self._post("/folders", {"title": title})
logger.info("[joplin] Creating notebook: %r (parent=%s)", title, parent_id)
data: dict = {"title": title}
if parent_id:
data["parent_id"] = parent_id
resp = self._post("/folders", data)
folder_id = resp["id"]
self._notebook_cache[title] = folder_id
self._notebook_cache[key] = folder_id
logger.debug("[joplin] Notebook created: %r → %s", title, folder_id)
return folder_id
def get_or_create_notebook_path(self, path: list[str]) -> str:
"""Ensure a nested notebook path exists and return the leaf folder ID.
Creates intermediate notebooks as needed.
Args:
path: Ordered list of notebook names, e.g. ["AI-ChatGPT", "No Project"].
Returns:
Joplin folder ID of the deepest (leaf) notebook.
"""
parent_id: str | None = None
for name in path:
parent_id = self.get_or_create_notebook(name, parent_id)
assert parent_id is not None
return parent_id
# ------------------------------------------------------------------
# Notes
# ------------------------------------------------------------------
@@ -153,21 +178,70 @@ class JoplinClient:
logger.info("[joplin] Note created: %r → %s", title, note_id)
return note_id
def update_note(self, note_id: str, title: str, body: str) -> None:
def update_note(
self, note_id: str, title: str, body: str, parent_id: str | None = None
) -> None:
"""Update the title and body of an existing note.
Args:
note_id: Joplin note ID.
title: New note title.
body: New note body (Markdown).
parent_id: If given, also move the note into this notebook. Joplin's
``PUT /notes/:id`` relocates a note when ``parent_id`` is set —
this is how notes self-heal to a new notebook mapping (e.g. the
Claude Code AI-Claude → AI-ClaudeCode migration) on re-sync.
"""
logger.debug(
"[joplin] Updating note %s: %r (%d chars)",
note_id, title, len(body),
)
self._put(f"/notes/{note_id}", {"title": title, "body": body})
data: dict = {"title": title, "body": body}
if parent_id:
data["parent_id"] = parent_id
self._put(f"/notes/{note_id}", data)
logger.info("[joplin] Note updated: %r (%s)", title, note_id)
# ------------------------------------------------------------------
# Resources (embedded media)
# ------------------------------------------------------------------
def create_resource(self, file_path: "Path", title: str | None = None) -> str:
"""Upload a local file as a Joplin resource and return its ID.
Joplin's ``POST /resources`` is multipart: a ``data`` file part plus a
``props`` JSON part. The returned ID is referenced from note bodies as
``:/<id>``.
"""
url = f"{self._base_url}/resources"
props = json.dumps({"title": title or file_path.name})
logger.debug("[joplin] POST /resources (%s)", file_path.name)
try:
with file_path.open("rb") as fh:
resp = requests.post(
url,
params={"token": self._token},
files={
"data": (file_path.name, fh),
"props": (None, props),
},
timeout=_REQUEST_TIMEOUT,
)
resp.raise_for_status()
resource_id = resp.json()["id"]
logger.info("[joplin] Resource created: %s → %s", file_path.name, resource_id)
return resource_id
except requests.exceptions.ConnectionError as e:
raise JoplinError(
"Cannot connect to Joplin. Is Joplin desktop running with Web Clipper enabled?"
) from e
except requests.exceptions.Timeout as e:
raise JoplinError(_timeout_message("POST", "/resources")) from e
except requests.exceptions.HTTPError as e:
raise JoplinError(_http_error_message("POST", "/resources", e)) from e
except (requests.exceptions.RequestException, KeyError, OSError) as e:
raise JoplinError(f"Joplin resource upload failed for {file_path.name}: {e}") from e
# ------------------------------------------------------------------
# HTTP helpers
# ------------------------------------------------------------------
@@ -233,11 +307,14 @@ class JoplinClient:
def _load_notebook_cache(self) -> None:
logger.debug("[joplin] Loading notebook list from Joplin…")
notebooks = self.list_notebooks()
self._notebook_cache = {nb["title"]: nb["id"] for nb in notebooks}
self._notebook_cache = {
(nb.get("parent_id") or None, nb["title"]): nb["id"]
for nb in notebooks
}
self._notebooks_loaded = True
logger.debug("[joplin] Notebook cache loaded: %d notebooks", len(self._notebook_cache))
for title, folder_id in self._notebook_cache.items():
logger.debug("[joplin] %r → %s", title, folder_id)
for (parent_id, title), folder_id in self._notebook_cache.items():
logger.debug("[joplin] (%s) %r → %s", parent_id or "root", title, folder_id)
# ------------------------------------------------------------------
@@ -279,25 +356,76 @@ def _http_error_message(method: str, path: str, e: requests.exceptions.HTTPError
return f"Joplin {method} {path} failed: HTTP {status}{body_snippet}"
# ------------------------------------------------------------------
# Embedded-media rewriting
# ------------------------------------------------------------------
# Markdown image/link targets pointing at the local media/ sibling directory,
# e.g. ![alt](media/file_x.png) or [name](media/clip.wav).
_MEDIA_LINK_RE = re.compile(r"(!?\[[^\]]*\])\((media/[^)\s]+)\)")
def upload_media_and_rewrite(
body: str,
note_dir: Path,
client: "JoplinClient",
resource_map: dict[str, str],
) -> tuple[str, dict[str, str]]:
"""Rewrite local ``media/…`` links to Joplin ``:/resourceId`` references.
Uploads each referenced file as a Joplin resource (once), rewriting both
image embeds and file links. ``resource_map`` (relative-path → resource_id)
is the idempotency record from the cache: known paths are reused, new ones
are added and returned so the caller can persist them. Missing files are
left as-is so the link still points at the on-disk copy.
"""
new_map = dict(resource_map)
def replace(match: re.Match) -> str:
label, rel_path = match.group(1), match.group(2)
resource_id = new_map.get(rel_path)
if resource_id is None:
file_path = (note_dir / rel_path).resolve()
if not file_path.is_file():
logger.warning("[joplin] Media file missing, leaving link: %s", rel_path)
return match.group(0)
resource_id = client.create_resource(file_path)
new_map[rel_path] = resource_id
return f"{label}(:/{resource_id})"
return _MEDIA_LINK_RE.sub(replace, body), new_map
# ------------------------------------------------------------------
# Notebook naming helper
# ------------------------------------------------------------------
_PROVIDER_DISPLAY = {
"chatgpt": "ChatGPT",
"claude": "Claude",
"chatgpt": "AI-ChatGPT",
"claude": "AI-Claude",
# Decision 2026-07-06: Claude Code coding sessions get their own top-level
# notebook (was nested under AI-Claude, which made them hard to find,
# intermixed with Claude web projects and named by dev folder). Existing
# notes self-heal into here on the next sync — update_note moves them.
"claude-code": "AI-ClaudeCode",
# Codex CLI coding sessions get their own top-level notebook for the same
# reason Claude Code does — they are a distinct surface, not a ChatGPT
# project, and burying them under AI-ChatGPT makes both harder to browse.
"codex": "AI-Codex",
}
def notebook_title(provider: str, project: str | None) -> str:
"""Derive a flat Joplin notebook title from provider and project name.
def notebook_path(provider: str, project: str | None) -> tuple[str, str]:
"""Return (parent_notebook, child_notebook) for the given provider and project.
The parent is the top-level provider notebook; the child is the project name.
Examples:
notebook_title("chatgpt", "no-project") → "ChatGPT - No Project"
notebook_title("claude", "budget-tracker") → "Claude - Budget Tracker"
notebook_title("chatgpt", None) → "ChatGPT - No Project"
notebook_path("chatgpt", None) → ("AI-ChatGPT", "No Project")
notebook_path("chatgpt", "no-project") → ("AI-ChatGPT", "No Project")
notebook_path("claude", "budget-tracker") → ("AI-Claude", "Budget Tracker")
"""
prov_display = _PROVIDER_DISPLAY.get(provider, provider.capitalize())
proj = (project or "no-project").replace("-", " ").title()
return f"{prov_display} - {proj}"
parent = _PROVIDER_DISPLAY.get(provider, f"AI-{provider.capitalize()}")
child = (project or "no-project").replace("-", " ").title()
return parent, child
+123
View File
@@ -0,0 +1,123 @@
"""Per-export-run tally for content that was dropped or partially extracted.
Surfaces the loss visibility that the rest of the system promises in its
output (visible ``unknown`` blocks). The summary emitted at the end of
each export is the load-bearing operator-facing signal: if a real content
type starts being silently dropped, this is where it shows up.
Pass a single instance through ``BaseProvider.normalize_conversation`` and
read it back in ``src/main.py`` after the export loop. No global state.
"""
from collections import Counter
from dataclasses import dataclass, field
_TOP_N_BREAKDOWN = 5
@dataclass
class LossReport:
"""Counters for things that didn't render cleanly in an export run."""
# Type-keyed counters. Values are int counts.
unknown_blocks: Counter = field(default_factory=Counter)
extraction_failures: Counter = field(default_factory=Counter)
filtered_roles: Counter = field(default_factory=Counter)
# Messages collapsed/omitted by the EXPORTER_HIDDEN_CONTENT policy —
# intentional, but still surfaced so the omission is never invisible.
collapsed: Counter = field(default_factory=Counter)
collapsed_bytes: int = 0
# Aggregate counters
messages_rendered: int = 0
conversations: int = 0
# Media downloads (EXPORTER_DOWNLOAD_MEDIA): successes / skips / failures
media_downloaded: int = 0
media_downloaded_bytes: int = 0
media_failed: Counter = field(default_factory=Counter)
# Recording -------------------------------------------------------------
def record_unknown(self, raw_type: str) -> None:
self.unknown_blocks[raw_type or "?"] += 1
def record_extraction_failure(self, raw_type: str) -> None:
self.extraction_failures[raw_type or "?"] += 1
def record_filtered_role(self, role: str) -> None:
self.filtered_roles[role or "?"] += 1
def record_collapsed(self, origin: str, size_bytes: int = 0) -> None:
self.collapsed[origin or "?"] += 1
if isinstance(size_bytes, int) and size_bytes > 0:
self.collapsed_bytes += size_bytes
def record_media_downloaded(self, size_bytes: int = 0) -> None:
self.media_downloaded += 1
if isinstance(size_bytes, int) and size_bytes > 0:
self.media_downloaded_bytes += size_bytes
def record_media_failed(self, reason: str) -> None:
self.media_failed[reason or "?"] += 1
def record_message(self) -> None:
self.messages_rendered += 1
def record_conversation(self) -> None:
self.conversations += 1
# Summary ---------------------------------------------------------------
def format_summary(self) -> str:
"""Return a multi-line summary table suitable for INFO logging.
Format pinned by plan §Post-export summary — "(none)" sentinel when a
counter is empty, top-5 breakdown with "+ N more types" overflow.
"""
lines: list[str] = ["[export] Run summary:"]
lines.append(f" conversations: {self.conversations}")
lines.append(f" messages rendered: {self.messages_rendered}")
lines.extend(_format_section("unknown blocks: ", self.unknown_blocks))
lines.extend(_format_section("extraction failures: ", self.extraction_failures))
lines.extend(_format_section("collapsed by policy: ", self.collapsed))
if self.collapsed:
lines.append(
f" ≈{self.collapsed_bytes / 1024:.0f} KB omitted "
"(intentional — EXPORTER_HIDDEN_CONTENT=full to keep)"
)
if self.media_downloaded or self.media_failed:
lines.append(
f" media downloaded: {self.media_downloaded} "
f"(≈{self.media_downloaded_bytes / 1024:.0f} KB)"
)
if self.media_failed:
total_failed = sum(self.media_failed.values())
lines.append(f" media failed: {total_failed}")
for reason, count in self.media_failed.most_common(_TOP_N_BREAKDOWN):
lines.append(f" {reason}={count}")
lines.append(
" filtered roles: "
"(filter lifted in v0.4.0 — counter retained for future use, expected 0)"
)
if self.filtered_roles:
for role, count in self.filtered_roles.most_common(_TOP_N_BREAKDOWN):
lines.append(f" {role}={count}")
return "\n".join(lines)
def _format_section(label: str, counter: Counter) -> list[str]:
"""Render one counter section: header line + indented breakdown lines."""
total = sum(counter.values())
header = f" {label} {total}"
if total == 0:
return [header, " (none)"]
lines = [header]
most_common = counter.most_common()
for raw_type, count in most_common[:_TOP_N_BREAKDOWN]:
lines.append(f" {raw_type}={count}")
if len(most_common) > _TOP_N_BREAKDOWN:
remainder = len(most_common) - _TOP_N_BREAKDOWN
lines.append(f" + {remainder} more types")
return lines
+849 -92
View File
File diff suppressed because it is too large Load Diff
+231
View File
@@ -0,0 +1,231 @@
"""Media resolution — download conversation assets next to the export files.
Walks a normalized conversation's placeholder blocks, downloads each asset via
the provider, saves it under a ``media/`` directory beside the conversation's
Markdown file, and annotates the block with a relative ``local_path`` so the
renderer emits a real inline image / file link instead of a placeholder.
Policy (``EXPORTER_DOWNLOAD_MEDIA``):
images (default) — image_placeholder blocks only
all — also file_placeholder blocks (voice-mode audio etc.;
measured 2026-06-12: 556 clips ≈ 162MB whose transcripts
are already in the exports as text)
off — leave placeholders untouched
Failures (expired assets, network) keep the existing placeholder and are
counted in the run summary — never fatal to the export.
"""
import logging
import os
import tempfile
from pathlib import Path
from src.blocks import BLOCK_TYPE_FILE_PLACEHOLDER, BLOCK_TYPE_IMAGE_PLACEHOLDER
from src.loss_report import LossReport
from src.providers.base import ProviderError
from src.utils import build_export_path, generate_filename
logger = logging.getLogger(__name__)
MEDIA_IMAGES = "images"
MEDIA_ALL = "all"
MEDIA_OFF = "off"
VALID_MEDIA_POLICIES = {MEDIA_IMAGES, MEDIA_ALL, MEDIA_OFF}
_EXT_BY_MIME = {
"image/png": ".png",
"image/jpeg": ".jpg",
"image/gif": ".gif",
"image/webp": ".webp",
"audio/wav": ".wav",
"audio/x-wav": ".wav",
"audio/mpeg": ".mp3",
"audio/mp4": ".m4a",
"audio/aac": ".aac",
"application/pdf": ".pdf",
}
def resolve_media_policy() -> str:
"""Read EXPORTER_DOWNLOAD_MEDIA from the environment, defaulting to images."""
value = os.getenv("EXPORTER_DOWNLOAD_MEDIA", "").strip().lower()
if not value:
return MEDIA_IMAGES
if value not in VALID_MEDIA_POLICIES:
logger.warning(
"EXPORTER_DOWNLOAD_MEDIA=%r is invalid (expected images|all|off) "
"— using 'images'.",
value,
)
return MEDIA_IMAGES
return value
def resolve_media(
normalized: dict,
provider,
export_base: Path,
structure: str,
policy: str,
report: LossReport,
) -> int:
"""Download assets for one normalized conversation. Returns download count.
Mutates placeholder blocks in place (adds ``local_path``). Idempotent:
an asset already on disk is annotated without an API call.
"""
if policy == MEDIA_OFF:
return 0
download = getattr(provider, "download_asset", None)
if download is None:
# Provider has no remote assets (e.g. claude-code local transcripts).
return 0
wanted_types = {BLOCK_TYPE_IMAGE_PLACEHOLDER}
if policy == MEDIA_ALL:
wanted_types.add(BLOCK_TYPE_FILE_PLACEHOLDER)
targets = [
block
for message in normalized.get("messages", [])
for block in message.get("blocks", [])
if block.get("type") in wanted_types and block.get("ref")
]
if not targets:
return 0
# Same path computation the exporter uses, so media/ lands beside the .md.
filename = generate_filename(
normalized.get("title", "Untitled"),
normalized.get("id", ""),
normalized.get("created_at") or "2000-01-01",
)
conv_dir = build_export_path(
export_base,
normalized.get("provider", ""),
normalized.get("project"),
normalized.get("created_at") or "2000-01-01",
filename,
structure,
).parent
media_dir = conv_dir / "media"
downloaded = 0
for block in targets:
ref = block["ref"]
file_id = _safe_asset_name(provider, ref)
if not file_id:
report.record_media_failed("unparseable-ref")
continue
existing = _find_existing(media_dir, file_id)
if existing is not None:
block["local_path"] = f"media/{existing.name}"
continue
try:
content, mime, file_name = download(ref)
except ProviderError as e:
logger.warning(
"[media] Could not download %s: %s", ref[:60], e.original
)
report.record_media_failed(_classify_failure(e))
continue
ext = _pick_extension(mime, file_name)
target = media_dir / f"{file_id}{ext}"
try:
_write_atomic(target, content)
except OSError as e:
logger.error("[media] Could not write %s: %s", target, e)
report.record_media_failed("write-error")
continue
block["local_path"] = f"media/{target.name}"
report.record_media_downloaded(len(content))
downloaded += 1
return downloaded
def _classify_failure(error: ProviderError) -> str:
"""Bucket a download failure for the run summary.
The buckets separate "the asset is gone" from "the asset is there but we
were refused" — different causes, different fixes, so lumping both into
download-error hides which one you have.
Note that a bare 403 from ChatGPT usually means *gone*, not *refused*:
its download endpoint reports a deleted upload as 403 Forbidden. The
provider confirms that against /files/{id} before raising, so a failure
that reaches here still saying "forbidden" is one where the file record
survives.
``forbidden`` does NOT mean recoverable. Investigated exhaustively
2026-08-17 (19 assets): those records report ``state: "ready"`` and carry
a ``library_file_id``, yet /files/{id}/download — the endpoint that mints
the signed estuary/content URL the web UI itself fetches — refuses them,
and the Library id is rejected as ``file_not_found``. Decisively, the same
images render blank in ChatGPT's own UI. Nothing was withheld from us; the
bytes are gone and only the metadata survived, apparently from a Library
migration dated 2026-07-14. Do not spend another afternoon on it: no
header, scope, namespace or id form reaches these. The bucket stays
separate only because the two failures look different on the wire.
"""
detail = str(error.original).lower()
if "404" in detail or "not found" in detail or "no longer exists" in detail:
return "expired-or-missing"
if "403" in detail or "forbidden" in detail:
return "forbidden"
return "download-error"
def _safe_asset_name(provider, ref: str) -> str | None:
"""A stable, filesystem-safe name for the asset (the provider file ID)."""
parser = getattr(provider, "parse_asset_file_id", None)
if parser is None:
from src.providers.chatgpt import parse_asset_file_id as parser
file_id = parser(ref)
if not file_id:
return None
return "".join(c if c.isalnum() or c in "-_." else "_" for c in file_id)
def _find_existing(media_dir: Path, file_id: str) -> Path | None:
"""Return an already-downloaded file for this asset, if present."""
if not media_dir.is_dir():
return None
for candidate in media_dir.glob(f"{file_id}.*"):
if candidate.is_file() and candidate.stat().st_size > 0:
return candidate
return None
def _pick_extension(mime: str | None, file_name: str | None) -> str:
if mime:
base_mime = mime.split(";")[0].strip().lower()
if base_mime in _EXT_BY_MIME:
return _EXT_BY_MIME[base_mime]
if file_name:
suffix = Path(file_name).suffix
if suffix and len(suffix) <= 8:
return suffix.lower()
return ".bin"
def _write_atomic(target: Path, content: bytes) -> None:
target.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_name = tempfile.mkstemp(dir=target.parent, suffix=".tmp")
try:
with os.fdopen(fd, "wb") as fh:
fh.write(content)
os.chmod(tmp_name, 0o600)
os.replace(tmp_name, target)
except OSError:
try:
os.unlink(tmp_name)
except OSError:
pass
raise
+179
View File
@@ -0,0 +1,179 @@
"""Push notifications for unattended runs, via ntfy (https://ntfy.sh).
A scheduled archive is invisible by design: the exporter logs to
``cache/logs/exporter.log``, systemd logs to the journal and Task Scheduler
records an exit code, but all three are *pull* — you only learn a run failed by
going to look. This module pushes a one-line result instead.
Configured entirely from the environment (see ``.env.example``):
NTFY_TOPIC topic name; unset disables notifications entirely
NTFY_SERVER default https://ntfy.sh — set to your own host if self-hosting
NTFY_TOKEN optional bearer token, for access-controlled topics
NTFY_NOTIFY always (default) | failure | off
**Payload is deliberately counts-only.** A public ntfy topic is readable by
anyone who guesses the name — there is no per-topic secret unless you add
NTFY_TOKEN against a server that enforces it. Conversation titles are sensitive
(they are the first line of what you asked), so nothing but provider names and
numbers ever goes into a notification. The machine name is included because two
machines archive into one topic and "2 exported" is meaningless without knowing
which box it came from.
Never fatal: notification is a courtesy, not part of the archive. Every failure
path here logs a warning and returns False, so a down ntfy server can't fail a
run that actually captured your data.
"""
import logging
import os
import socket
import requests
logger = logging.getLogger(__name__)
DEFAULT_SERVER = "https://ntfy.sh"
NOTIFY_ALWAYS = "always"
NOTIFY_FAILURE = "failure"
NOTIFY_OFF = "off"
VALID_NOTIFY_POLICIES = {NOTIFY_ALWAYS, NOTIFY_FAILURE, NOTIFY_OFF}
# ntfy request timeout (connect, read). Short: a notification is never worth
# holding a scheduled run open for.
_TIMEOUT = (5, 10)
def resolve_notify_policy() -> str:
"""Read NTFY_NOTIFY from the environment, defaulting to ``always``."""
value = os.getenv("NTFY_NOTIFY", "").strip().lower()
if not value:
return NOTIFY_ALWAYS
if value not in VALID_NOTIFY_POLICIES:
logger.warning(
"NTFY_NOTIFY=%r is invalid (expected always|failure|off) — using 'always'.",
value,
)
return NOTIFY_ALWAYS
return value
def is_configured() -> bool:
"""True when a topic is set and the policy isn't ``off``."""
return bool(os.getenv("NTFY_TOPIC", "").strip()) and resolve_notify_policy() != NOTIFY_OFF
def machine_name() -> str:
"""Short hostname — two machines share one topic, so this is load-bearing."""
try:
return socket.gethostname().split(".")[0] or "unknown-host"
except Exception:
return "unknown-host"
def _ascii_header(value: str) -> str:
"""Reduce a header value to plain ASCII.
HTTP headers are latin-1 at best, and requests raises on anything outside
it — an em dash in a title is enough to lose the notification entirely
(observed 2026-08-18). The body has no such limit: it is sent as UTF-8
bytes, so only header values need flattening.
"""
replacements = {"\u2014": "-", "\u2013": "-", "\u2018": "'", "\u2019": "'",
"\u201c": '"', "\u201d": '"', "\u2026": "..."}
for bad, good in replacements.items():
value = value.replace(bad, good)
return value.encode("ascii", "replace").decode("ascii")
def send(
title: str,
message: str,
tags: str = "",
priority: str = "",
) -> bool:
"""POST a notification to ntfy. Returns True on success, never raises.
``tags`` is a comma-separated list of ntfy tag names (emoji shortcodes are
rendered as emoji); ``priority`` is one of min/low/default/high/urgent.
"""
topic = os.getenv("NTFY_TOPIC", "").strip()
if not topic:
return False
server = (os.getenv("NTFY_SERVER", "").strip() or DEFAULT_SERVER).rstrip("/")
url = f"{server}/{topic}"
headers = {"Title": _ascii_header(title)}
if tags:
headers["Tags"] = _ascii_header(tags)
if priority:
headers["Priority"] = _ascii_header(priority)
token = os.getenv("NTFY_TOKEN", "").strip()
if token:
headers["Authorization"] = f"Bearer {token}"
try:
resp = requests.post(
url,
data=message.encode("utf-8"),
headers=headers,
timeout=_TIMEOUT,
)
resp.raise_for_status()
logger.debug("[notify] Sent to %s", url)
return True
except Exception as e:
# Deliberately broad: a notification must never fail an archive run.
logger.warning("[notify] Could not send to %s: %s", url, e)
return False
def format_summary(
export_summary: dict | None,
joplin_summary: dict | None,
failures: list[str] | None,
) -> tuple[str, str, str, str]:
"""Build ``(title, message, tags, priority)`` for a completed sync run.
Counts only — see the module docstring on why no titles appear here.
"""
failures = failures or []
host = machine_name()
lines: list[str] = []
total_exported = 0
for prov, counts in (export_summary or {}).items():
exported = counts.get("exported", 0)
skipped = counts.get("skipped", 0)
failed = counts.get("failed", 0)
total_exported += exported
line = f"{prov}: {exported} exported, {skipped} up to date"
if failed:
line += f", {failed} FAILED"
lines.append(line)
synced = sum(
counts.get("created", 0) + counts.get("updated", 0)
for counts in (joplin_summary or {}).values()
)
if joplin_summary:
lines.append(f"joplin: {synced} note(s) created/updated")
if failures:
lines.append("")
lines.extend(failures)
title = f"AI archive FAILED - {host}"
tags = "rotating_light"
priority = "high"
else:
title = f"AI archive OK - {host}"
# A run that captured something is more interesting than a quiet one.
tags = "white_check_mark" if total_exported else "zzz"
priority = "default" if total_exported else "low"
if not lines:
lines = ["nothing to do"]
return title, "\n".join(lines), tags, priority
+188 -13
View File
@@ -1,6 +1,7 @@
"""Abstract base class for AI chat providers."""
import logging
import os
import random
import time
from abc import ABC, abstractmethod
@@ -9,6 +10,7 @@ from typing import Any
import requests
from src.loss_report import LossReport
from src.utils import redact_secrets
# curl_cffi has its own exception hierarchy (rooted at CurlError → OSError),
@@ -31,11 +33,117 @@ logger = logging.getLogger(__name__)
# Request timeouts (connect, read) in seconds
REQUEST_TIMEOUT = (10, 30)
# Longest error-body excerpt to carry into a ProviderError message.
_ERROR_BODY_CHARS = 300
# Retry configuration
MAX_RETRIES = 3
BACKOFF_BASE = 2.0
BACKOFF_MAX = 60.0
# EXPORTER_HIDDEN_CONTENT policy — what to do with content that was invisible
# in the provider's UI (retrieval dumps, hidden-flagged context, agent tool
# traffic). Shared by all providers.
HIDDEN_CONTENT_FULL = "full"
HIDDEN_CONTENT_PLACEHOLDER = "placeholder"
HIDDEN_CONTENT_OMIT = "omit"
VALID_HIDDEN_CONTENT_POLICIES = {
HIDDEN_CONTENT_FULL,
HIDDEN_CONTENT_PLACEHOLDER,
HIDDEN_CONTENT_OMIT,
}
# ---------------------------------------------------------------------------
# API-drift canary (FUTURE.md §10)
# ---------------------------------------------------------------------------
# The export depends on undocumented internal web APIs that can change shape
# without notice. The canary fetches a small live sample and asserts only the
# normalizer's load-bearing fields (NOT full shape — provider settings objects
# churn constantly and would be pure noise). Findings carry a severity:
# ERROR — a load-bearing field is missing/mistyped; the parser will break or
# silently lose data. The backup is no longer trustworthy.
# WARN — something unfamiliar appeared (new content_type, new tool author,
# a dormant field went live). Investigate; not necessarily broken.
# OK — the assertion held.
DRIFT_ERROR = "error"
DRIFT_WARN = "warn"
DRIFT_OK = "ok"
def drift_finding(severity: str, check: str, detail: str = "") -> dict:
"""Build one canary finding."""
return {"severity": severity, "check": check, "detail": detail}
def resolve_hidden_content_policy() -> str:
"""Read EXPORTER_HIDDEN_CONTENT from the environment, defaulting to placeholder."""
value = os.getenv("EXPORTER_HIDDEN_CONTENT", "").strip().lower()
if not value:
return HIDDEN_CONTENT_PLACEHOLDER
if value not in VALID_HIDDEN_CONTENT_POLICIES:
logger.warning(
"EXPORTER_HIDDEN_CONTENT=%r is invalid (expected full|placeholder|omit) "
"— using 'placeholder'.",
value,
)
return HIDDEN_CONTENT_PLACEHOLDER
return value
# Default polite pacing between consecutive API requests (seconds). Keeps
# traffic human-paced instead of bursty. REQUEST_DELAY=0 disables.
DEFAULT_REQUEST_DELAY = 1.0
def resolve_request_delay() -> float:
"""Read REQUEST_DELAY from the environment, defaulting to DEFAULT_REQUEST_DELAY."""
raw = os.getenv("REQUEST_DELAY", "").strip()
if not raw:
return DEFAULT_REQUEST_DELAY
try:
value = float(raw)
except ValueError:
logger.warning(
"REQUEST_DELAY=%r is not a number — using default %.1fs.",
raw,
DEFAULT_REQUEST_DELAY,
)
return DEFAULT_REQUEST_DELAY
if value < 0:
logger.warning(
"REQUEST_DELAY=%r is negative — using default %.1fs.",
raw,
DEFAULT_REQUEST_DELAY,
)
return DEFAULT_REQUEST_DELAY
return value
def _describe_error_body(response: Any) -> str:
"""Summarise an error response body for a log line.
Providers explain 4xx in the body (ChatGPT uses ``detail``), so a bare
status code is not a diagnosis. Prefers ``detail``/``error``/``message``,
falls back to a truncated raw excerpt, and redacts before it is logged.
"""
try:
body = response.json()
except Exception:
try:
text = (response.text or "").strip()
except Exception:
return "no response body"
if not text:
return "empty response body"
return f"body: {text[:_ERROR_BODY_CHARS]}"
if isinstance(body, dict):
for key in ("detail", "error", "message"):
if key in body:
value = redact_secrets(body[key])
return f"{key}: {str(value)[:_ERROR_BODY_CHARS]}"
return f"body: {str(redact_secrets(body))[:_ERROR_BODY_CHARS]}"
# Realistic Chrome User-Agent
USER_AGENT = (
"Mozilla/5.0 (X11; Linux x86_64) "
@@ -75,6 +183,9 @@ class BaseProvider(ABC):
"Accept-Language": "en-US,en;q=0.9",
}
)
# Polite pacing between consecutive requests (REQUEST_DELAY env).
self._request_delay = resolve_request_delay()
self._last_request_at: float | None = None
# ------------------------------------------------------------------
# Abstract interface — subclasses must implement these
@@ -89,13 +200,45 @@ class BaseProvider(ABC):
"""Return the full conversation detail for a single ID."""
@abstractmethod
def normalize_conversation(self, raw: dict) -> dict:
"""Transform provider-specific schema to the common normalized schema."""
def normalize_conversation(self, raw: dict, loss_report: LossReport | None = None) -> dict:
"""Transform provider-specific schema to the common normalized schema.
``loss_report`` accumulates counts of dropped/unhandled content so the
export loop can surface a single summary at the end. When None, providers
construct a throwaway local report (so calling normalize_conversation in
isolation, e.g. from tests, doesn't crash).
"""
def check_drift(self) -> list[dict]:
"""Probe the live API and assert the normalizer's load-bearing fields.
Returns a list of ``drift_finding`` dicts. The default is a no-op
(used by providers with no remote schema risk, e.g. local Claude Code
transcripts); web-API providers override this. See FUTURE.md §10.
"""
return []
# ------------------------------------------------------------------
# Concrete helpers
# ------------------------------------------------------------------
def _pace(self) -> None:
"""Sleep so consecutive requests are at least REQUEST_DELAY apart (±25% jitter).
getattr fallback: tests construct providers via ``__new__``, skipping
``__init__`` — those run unpaced.
"""
delay = getattr(self, "_request_delay", 0)
if delay <= 0:
return
target = delay * random.uniform(0.75, 1.25)
last = getattr(self, "_last_request_at", None)
if last is not None:
wait = target - (time.monotonic() - last)
if wait > 0:
time.sleep(wait)
self._last_request_at = time.monotonic()
def fetch_all_conversations(self, since: datetime | None = None) -> list[dict]:
"""Fetch every conversation, handling pagination automatically.
@@ -192,7 +335,8 @@ class BaseProvider(ABC):
Parsed JSON response body.
Raises:
ProviderError: On 401, exhausted retries, or unrecoverable errors.
ProviderError: On an authentication failure, exhausted retries, or
unrecoverable errors.
"""
kwargs.setdefault("timeout", REQUEST_TIMEOUT)
@@ -201,6 +345,7 @@ class BaseProvider(ABC):
while attempt <= MAX_RETRIES:
attempt += 1
self._pace()
start = time.monotonic()
try:
@@ -217,14 +362,19 @@ class BaseProvider(ABC):
elapsed_ms,
)
# ── 401: token expired / invalid ──────────────────────────
if response.status_code == 401:
self._handle_401()
# _handle_401 raises ProviderError — this line never runs
# ── Auth failure: token expired / invalid ─────────────────
# Not keyed on 401 alone: a provider is free to answer an
# invalid session with some other status, and claude.ai does
# (403). Ask the provider rather than assuming.
if self._is_auth_failure(response):
self._handle_auth_failure(response)
# _handle_auth_failure raises — this line never runs
raise ProviderError(
self.provider_name,
f"{method} {url}",
RuntimeError("401 Unauthorized"),
RuntimeError(
f"HTTP {response.status_code} — not authenticated"
),
)
# ── 429: rate limited ──────────────────────────────────────
@@ -269,7 +419,19 @@ class BaseProvider(ABC):
continue
# ── Other HTTP errors ──────────────────────────────────────
response.raise_for_status()
# Raise with the response body attached. curl_cffi formats
# raise_for_status() as "HTTP Error {code}: {reason}", and
# HTTP/2 carries no reason phrase — so the bare exception
# reads "HTTP Error 403:" and says nothing about the cause.
# The provider's JSON `detail` is the only explanation there is.
if not response.ok:
raise ProviderError(
self.provider_name,
f"{method} {url}",
RuntimeError(
f"HTTP {response.status_code} — {_describe_error_body(response)}"
),
)
# ── Success ────────────────────────────────────────────────
body = response.json()
@@ -320,11 +482,22 @@ class BaseProvider(ABC):
last_exc or RuntimeError("Unknown error"),
)
def _handle_401(self) -> None:
"""Log a clear human-readable message for a 401 and raise ProviderError."""
def _is_auth_failure(self, response: Any) -> bool:
"""Whether this response means the stored credential is not valid.
Defaults to 401. Providers whose API reports an invalid session with a
different status override this — claude.ai returns 403, so keying auth
handling on 401 alone reports an expired key as a generic permission
error and never tells the user to refresh it.
"""
return bool(response.status_code == 401)
def _handle_auth_failure(self, response: Any) -> None:
"""Log a clear human-readable message for an auth failure and raise."""
# Subclasses override to include provider-specific cookie name
msg = (
f"[{self.provider_name}] Authentication failed (401 Unauthorized). "
f"[{self.provider_name}] Authentication failed "
f"(HTTP {response.status_code}). "
"Your session token has likely expired. "
"Run 'ai-chat-exporter auth' to refresh your token."
)
@@ -332,7 +505,9 @@ class BaseProvider(ABC):
raise ProviderError(
self.provider_name,
"authentication",
RuntimeError("401 Unauthorized — token expired"),
RuntimeError(
f"HTTP {response.status_code} — session token expired or invalid"
),
)
@staticmethod
+946 -98
View File
File diff suppressed because it is too large Load Diff
+259 -53
View File
@@ -2,10 +2,29 @@
import logging
import os
from typing import Any
from curl_cffi import requests as curl_requests
from src.providers.base import BaseProvider, ProviderError
from src.blocks import (
UNKNOWN_REASON_EXTRACTION_FAILED,
UNKNOWN_REASON_UNKNOWN_TYPE,
make_image_placeholder,
make_text_block,
make_thinking_block,
make_tool_result_block,
make_tool_use_block,
make_unknown_block,
)
from src.loss_report import LossReport
from src.providers.base import (
BaseProvider,
DRIFT_ERROR,
DRIFT_OK,
DRIFT_WARN,
ProviderError,
drift_finding,
)
logger = logging.getLogger(__name__)
@@ -20,7 +39,9 @@ class ClaudeProvider(BaseProvider):
Cloudflare's bot detection (same issue as chatgpt.com).
Authentication: sessionKey cookie (~30 day lifetime, opaque string).
Expiry cannot be decoded client-side — a 401 is the only signal.
Expiry cannot be decoded client-side, so an API rejection is the only
signal — and claude.ai sends 403 permission_error / account_session_invalid
for an invalid session, not 401. See ``_is_auth_failure``.
"""
provider_name = "claude"
@@ -53,11 +74,41 @@ class ClaudeProvider(BaseProvider):
self._org_id: str | None = None # cached per session
logger.debug("[claude] Session initialised with Chrome TLS impersonation (key: [REDACTED])")
def _handle_401(self) -> None:
# claude.ai answers an invalid or expired sessionKey with 403
# permission_error / account_session_invalid — never 401. Verified live
# 2026-09-20 against GET /api/organizations: a valid key returns 200, while
# an expired key, a garbage key and no cookie at all return byte-identical
# 403s carrying this code. Matching on the code rather than on the bare
# status keeps a genuine permission problem (which would carry a different
# code) reported as itself.
_SESSION_INVALID_CODE = "account_session_invalid"
def _is_auth_failure(self, response: Any) -> bool:
if response.status_code == 401:
return True
if response.status_code != 403:
return False
try:
body = response.json()
except Exception:
return False
if not isinstance(body, dict):
return False
error = body.get("error")
if not isinstance(error, dict):
return False
details = error.get("details")
if not isinstance(details, dict):
return False
return details.get("error_code") == self._SESSION_INVALID_CODE
def _handle_auth_failure(self, response: Any) -> None:
msg = (
"[claude] Authentication failed (401 Unauthorized). "
f"[claude] Authentication failed (HTTP {response.status_code}). "
"Your sessionKey has likely expired (~30 day lifetime). "
"Note: Claude session keys are opaque — a 401 is the only expiry signal. "
"Note: Claude session keys are opaque, and claude.ai reports an "
"invalid one as 403 'Invalid authorization', not 401 — an API "
"rejection is the only expiry signal. "
"To refresh: open claude.ai in Chrome → F12 → Application → Cookies "
"→ find 'sessionKey' → copy the value. "
"Then run 'ai-chat-exporter auth' or update CLAUDE_SESSION_KEY in .env."
@@ -66,7 +117,9 @@ class ClaudeProvider(BaseProvider):
raise ProviderError(
self.provider_name,
"authentication",
RuntimeError("401 Unauthorized — Claude session key expired"),
RuntimeError(
f"HTTP {response.status_code} — Claude session key expired or invalid"
),
)
def _get_org_id(self) -> str:
@@ -161,8 +214,9 @@ class ClaudeProvider(BaseProvider):
return data
def normalize_conversation(self, raw: dict) -> dict:
def normalize_conversation(self, raw: dict, loss_report: LossReport | None = None) -> dict:
"""Transform Claude raw schema to the common normalized schema."""
report = loss_report if loss_report is not None else LossReport()
conv_id = raw.get("uuid") or raw.get("id", "")
title = raw.get("name") or raw.get("title") or "Untitled"
created_at = raw.get("created_at") or raw.get("create_time") or ""
@@ -178,40 +232,37 @@ class ClaudeProvider(BaseProvider):
# Messages
raw_messages = raw.get("chat_messages") or raw.get("messages") or []
messages = []
messages: list[dict] = []
for msg in raw_messages:
role = _map_role(msg.get("sender") or msg.get("role", ""))
if not role:
continue
# Content can be a string or a list of content blocks
content_raw = msg.get("content") or msg.get("text") or ""
content, skipped_types = _extract_claude_text(content_raw, conv_id)
for ctype in skipped_types:
logger.warning(
"[claude] Skipping %s content in conversation %s "
"— rich content not yet supported (see FUTURE.md)",
ctype,
conv_id[:8],
)
content_raw = msg.get("content") if "content" in msg else msg.get("text", "")
blocks = _extract_claude_blocks(content_raw, conv_id, report)
timestamp = msg.get("created_at") or msg.get("timestamp") or None
if content is None:
if not blocks:
logger.debug("[claude] Skipping empty message in conversation %s", conv_id[:8])
continue
content_type = "text"
messages.append(
{
"role": role,
"content": content,
"content_type": "text",
"content_type": content_type,
"timestamp": timestamp,
"blocks": blocks,
}
)
for _ in messages:
report.record_message()
report.record_conversation()
return {
"id": conv_id,
"title": title,
@@ -223,6 +274,70 @@ class ClaudeProvider(BaseProvider):
"messages": messages,
}
def check_drift(self) -> list[dict]:
"""Probe one listing page + one conversation; assert load-bearing fields.
See FUTURE.md §10. Real Claude messages are flat ``text``/``sender``;
the rich ``content`` block path is dormant, so its activation (content
becoming a list) is itself a drift event worth flagging. ``attachments``
/ ``files`` are ignored by the normalizer — non-empty values are a
silent-loss risk.
"""
findings: list[dict] = []
page = self.list_conversations(offset=0, limit=5)
if not page:
return [drift_finding(DRIFT_WARN, "listing",
"empty listing — cannot verify shape")]
item = page[0]
if not (item.get("uuid") or item.get("id")):
findings.append(drift_finding(DRIFT_ERROR, "listing", "summary missing uuid/id"))
if not (item.get("name") or item.get("title")):
findings.append(drift_finding(DRIFT_WARN, "listing", "no name/title on summary"))
if not (item.get("updated_at") or item.get("update_time")):
findings.append(drift_finding(DRIFT_WARN, "listing", "no updated_at on summary"))
conv_id = item.get("uuid") or item.get("id")
if not conv_id:
return findings or [drift_finding(DRIFT_ERROR, "listing", "no id to fetch")]
raw = self.get_conversation(conv_id)
if not (raw.get("uuid") or raw.get("id")):
findings.append(drift_finding(DRIFT_ERROR, "detail", "no uuid/id on detail"))
msgs = raw.get("chat_messages") or raw.get("messages")
if not isinstance(msgs, list) or not msgs:
findings.append(drift_finding(
DRIFT_ERROR, "detail", "chat_messages missing/empty — 0 messages"))
return findings
saw_sender = saw_text = False
flagged_list = flagged_attach = False
for m in msgs:
if m.get("sender") or m.get("role"):
saw_sender = True
content = m.get("content")
if isinstance(content, list) and not flagged_list:
findings.append(drift_finding(
DRIFT_WARN, "content",
"message.content is now a LIST — Claude switched to typed blocks; "
"the dormant _dispatch_claude_block path is now live, verify rendering"))
flagged_list = True
if m.get("text") or content:
saw_text = True
if (m.get("attachments") or m.get("files")) and not flagged_attach:
findings.append(drift_finding(
DRIFT_WARN, "attachments",
"message has non-empty attachments/files — the normalizer ignores "
"these (silent loss); add handling"))
flagged_attach = True
if not saw_sender:
findings.append(drift_finding(DRIFT_ERROR, "messages", "no sender/role on any message"))
if not saw_text:
findings.append(drift_finding(DRIFT_WARN, "messages", "no text/content found"))
if not findings:
findings.append(drift_finding(DRIFT_OK, "schema", "all load-bearing fields present"))
return findings
# ---------------------------------------------------------------------------
# Internal helpers
@@ -242,43 +357,134 @@ def _map_role(sender: str) -> str | None:
return mapping.get(sender.lower()) if sender else None
def _extract_claude_text(
content: str | list | dict, conv_id: str
) -> tuple[str | None, list[str]]:
"""Extract plain text from a Claude content field.
def _extract_claude_blocks(
content: str | list | dict | None, conv_id: str, report: LossReport
) -> list[dict]:
"""Extract typed blocks from a Claude content field.
Returns:
(text_or_None, list_of_skipped_content_types)
Defensive dispatch — zero observed cases of rich Claude content in the
user's archive at planning time, so this is theory-only. Real shapes
will be locked in v0.4.1 once captured. Any unrecognised block type
surfaces as an `unknown` block + WARNING + tally.
"""
skipped: list[str] = []
if content is None:
return []
if isinstance(content, str):
text = content.strip()
return (text if text else None), skipped
block = make_text_block(content)
return [block] if block else []
if isinstance(content, list):
parts: list[str] = []
for block in content:
if isinstance(block, str):
parts.append(block)
elif isinstance(block, dict):
btype = block.get("type", "text")
if btype == "text":
t = block.get("text", "").strip()
if t:
parts.append(t)
else:
skipped.append(btype)
text = "\n".join(parts).strip()
return (text if text else None), skipped
blocks: list[dict] = []
for item in content:
if isinstance(item, str):
block = make_text_block(item)
if block:
blocks.append(block)
elif isinstance(item, dict):
blocks.extend(_dispatch_claude_block(item, conv_id, report))
return blocks
if isinstance(content, dict):
btype = content.get("type", "text")
if btype == "text":
text = content.get("text", "").strip()
return (text if text else None), skipped
else:
skipped.append(btype)
return None, skipped
return _dispatch_claude_block(content, conv_id, report)
return None, skipped
return []
def _dispatch_claude_block(block: dict, conv_id: str, report: LossReport) -> list[dict]:
"""Translate one raw Claude content block into normalized blocks."""
btype = block.get("type", "text")
if btype == "text":
block_obj = make_text_block(block.get("text", "") or "")
return [block_obj] if block_obj else []
if btype == "thinking":
# Claude extended-thinking blocks may use 'thinking' or 'text' field.
text = block.get("thinking") or block.get("text") or ""
block_obj = make_thinking_block(text)
return [block_obj] if block_obj else []
if btype == "tool_use":
return [
make_tool_use_block(
name=block.get("name", "") or "",
input_data=block.get("input"),
tool_id=block.get("id"),
)
]
if btype == "tool_result":
# ``content`` may be a string or a list of nested blocks (recursive).
nested = block.get("content")
output = _flatten_tool_result_content(nested, conv_id, report)
return [
make_tool_result_block(
output=output,
tool_name=None,
is_error=bool(block.get("is_error")),
)
]
if btype == "image":
# Source shape is unverified; try the most likely fields.
source = block.get("source") or {}
ref = ""
if isinstance(source, dict):
ref = (
source.get("file_uuid")
or source.get("media_type")
or source.get("url")
or ""
)
return [make_image_placeholder(ref=ref or "(unknown)", source="user_upload")]
# Unknown block type
keys = list(block.keys())
logger.warning(
"[claude] Unknown block type %r in conversation %s "
"— see plan §Data-loss visibility (rendering as unknown block)",
btype,
conv_id[:8],
)
report.record_unknown(btype or "?")
return [
make_unknown_block(
raw_type=btype or "?",
observed_keys=keys,
reason=UNKNOWN_REASON_UNKNOWN_TYPE,
)
]
def _flatten_tool_result_content(
nested: object, conv_id: str, report: LossReport
) -> str:
"""Flatten Claude tool_result content (string OR list of nested blocks) to text.
Recurses into nested text blocks; any non-text nested block becomes a
visible inline marker so non-text content isn't silently dropped.
"""
if nested is None:
return ""
if isinstance(nested, str):
return nested
if isinstance(nested, list):
chunks: list[str] = []
for item in nested:
if isinstance(item, str):
chunks.append(item)
elif isinstance(item, dict):
btype = item.get("type", "text")
if btype == "text":
chunks.append(item.get("text", "") or "")
else:
keys = list(item.keys())[:10]
report.record_extraction_failure(f"tool_result.{btype}")
chunks.append(
f"[Unsupported nested {btype} block; keys={keys}]"
)
return "\n".join(c for c in chunks if c)
if isinstance(nested, dict):
return _flatten_tool_result_content([nested], conv_id, report)
return str(nested)
+738
View File
@@ -0,0 +1,738 @@
"""Claude Code session provider — archives local agent transcripts.
Reads JSONL session files from ``~/.claude/projects/<munged-cwd>/<uuid>.jsonl``
(override with ``CLAUDE_CODE_DIR``). No tokens, no rate limits, no ToS risk —
but the data lives in single files Claude Code may clean up, and it contains
deliverables (reviews, plans, analyses) that exist nowhere else.
Rendering follows the EXPORTER_HIDDEN_CONTENT policy (decided 2026-06-12):
prose-only by default. User prompts and assistant text are kept; tool_use /
tool_result traffic is collapsed to one grouped placeholder per activity run
(measured: dialogue prose is ~4% of session bytes); thinking blocks are
dropped (counted in the run summary, no placeholder). ``full`` keeps
everything.
Record types in a session file: ``user`` / ``assistant`` carry the dialogue
(Anthropic-style ``message.content`` block arrays); ``ai-title`` carries the
evolving session title (last one wins); ``last-prompt``,
``file-history-snapshot``, ``attachment``, ``permission-mode``, ``system``
are harness records and are skipped. ``isMeta`` records are harness-generated
user records and are skipped.
Subagents (Task tool): Claude Code stores each subagent's transcript as a
separate ``<session>/subagents/agent-*.jsonl`` file with an ``agent-*.meta.json``
sidecar (``agentType``, ``description``, ``toolUseId``). ``_load_subagents``
loads them keyed by ``toolUseId``; at the ``Task``/``Agent`` tool call that
spawned it, the subagent is folded inline as a collapsible ``<details>`` block
(the subagent's own records are flagged ``isSidechain`` and only processed on
this recursive pass — the top-level pass still skips sidechain records).
Scan scope: multiple ``projects/`` roots are supported — ``CLAUDE_CODE_DIR`` may
be an ``os.pathsep``-separated list and ``CLAUDE_CONFIG_DIR``'s ``projects/`` is
included when set. Sessions from all roots are merged by launch-folder; a session
UUID present in two roots keeps the newer-mtime copy. See ``resolve_roots``.
"""
import json
import logging
import os
import re
from collections import Counter
from datetime import datetime, timezone
from pathlib import Path
from src.blocks import (
COLLAPSED_KIND_TOOL_DUMP,
UNKNOWN_REASON_UNKNOWN_TYPE,
make_collapsed_block,
make_image_placeholder,
make_subagent_block,
make_text_block,
make_thinking_block,
make_tool_result_block,
make_tool_use_block,
make_unknown_block,
)
from src.loss_report import LossReport
from src.utils import git_root_name
from src.providers.base import (
BaseProvider,
HIDDEN_CONTENT_FULL,
ProviderError,
VALID_HIDDEN_CONTENT_POLICIES,
resolve_hidden_content_policy,
)
logger = logging.getLogger(__name__)
DEFAULT_PROJECTS_DIR = "~/.claude/projects"
# Tool names whose call spawns a subagent (Task tool) — its separate transcript
# is folded inline. Both spellings have appeared across Claude Code versions.
_SUBAGENT_TOOL_NAMES = {"Task", "Agent"}
def resolve_roots(projects_dir=None) -> list[Path]:
"""Resolve the ordered list of Claude Code ``projects/`` roots to scan.
Precedence:
1. Explicit ``projects_dir`` (a single path or a list) — used by tests.
2. ``CLAUDE_CODE_DIR`` env, split on ``os.pathsep`` (``:``) so multiple
roots can be given; a single path (no separator) stays backward
compatible. Falls back to the default root when unset.
3. Additionally, if ``CLAUDE_CONFIG_DIR`` is set, its ``projects/`` subdir
is appended — this is the common "non-default config dir" case.
Roots are expanded, de-duplicated (order preserved), and returned as-is
(existence is checked by the caller / ``_scan``). Sessions from all roots are
merged by launch-folder; there is no per-source label.
"""
raw: list[str]
if projects_dir is not None:
raw = [str(p) for p in projects_dir] if isinstance(projects_dir, (list, tuple)) \
else [str(projects_dir)]
else:
env = os.getenv("CLAUDE_CODE_DIR")
raw = env.split(os.pathsep) if env else [DEFAULT_PROJECTS_DIR]
config_dir = os.getenv("CLAUDE_CONFIG_DIR")
if config_dir:
raw.append(str(Path(config_dir) / "projects"))
roots: list[Path] = []
for p in raw:
p = p.strip()
if not p:
continue
path = Path(p).expanduser()
if path not in roots:
roots.append(path)
return roots
# Harness-injected tags inside user message text. Stripped so exports contain
# the dialogue, not the CLI plumbing. A record that is nothing but tags
# (e.g. a /model invocation) ends up empty and is skipped.
_HARNESS_TAG_RE = re.compile(
r"<local-command-caveat>.*?</local-command-caveat>"
r"|<command-name>.*?</command-name>"
r"|<command-message>.*?</command-message>"
r"|<command-args>.*?</command-args>"
r"|<command-contents>.*?</command-contents>"
r"|<local-command-stdout>.*?</local-command-stdout>"
r"|<system-reminder>.*?</system-reminder>"
# A fork subagent's first user turn: generic worker rules ahead of the
# fork's actual "Your directive: …", which is kept.
r"|<fork-boilerplate>.*?</fork-boilerplate>",
re.DOTALL,
)
# How many distinct tool names to list in a collapsed-activity placeholder.
_TOOL_NAMES_SHOWN = 4
def _strip_harness_noise(text: str) -> str:
if not isinstance(text, str):
return ""
return _HARNESS_TAG_RE.sub("", text).strip()
class ClaudeCodeProvider(BaseProvider):
"""Local-file provider over Claude Code session transcripts."""
provider_name = "claude-code"
def __init__(
self,
projects_dir: str | Path | None = None,
hidden_content: str | None = None,
) -> None:
super().__init__()
self._projects_dirs = resolve_roots(projects_dir)
self._hidden_content = (
hidden_content
if hidden_content in VALID_HIDDEN_CONTENT_POLICIES
else resolve_hidden_content_policy()
)
# conv_id → session file path, populated by _scan()
self._path_map: dict[str, Path] = {}
# ------------------------------------------------------------------
# BaseProvider interface
# ------------------------------------------------------------------
def list_conversations(self, offset: int = 0, limit: int = 100) -> list[dict]:
full = self._scan()
return full[offset : offset + limit]
def fetch_all_conversations(self, since: datetime | None = None) -> list[dict]:
convs = self._scan()
if since is not None:
since_aware = since if since.tzinfo else since.replace(tzinfo=timezone.utc)
convs = [
c for c in convs
if datetime.fromisoformat(c["updated_at"]) >= since_aware
]
logger.info(
"[claude-code] Found %d session(s) under %s",
len(convs),
", ".join(str(d) for d in self._projects_dirs),
)
return convs
def get_conversation(self, conv_id: str) -> dict:
path = self._path_map.get(conv_id)
if path is None:
# Direct call without a prior listing (e.g. tests) — scan first.
self._scan()
path = self._path_map.get(conv_id)
if path is None or not path.exists():
raise ProviderError(
self.provider_name,
f"get_conversation({conv_id[:8]})",
FileNotFoundError(f"No session file for id {conv_id}"),
)
records = _parse_jsonl(path)
mtime = datetime.fromtimestamp(path.stat().st_mtime, tz=timezone.utc)
return {
"id": conv_id,
"_path": str(path),
"_records": records,
# Subagent (Task-tool) transcripts, keyed by the parent tool_use id
# that spawned them, folded inline during normalization.
"_subagents": _load_subagents(path),
# Listing and normalized updated_at must match, or the cache
# staleness comparison would re-export every session every run.
"_mtime_iso": mtime.isoformat(),
}
def normalize_conversation(self, raw: dict, loss_report: LossReport | None = None) -> dict:
report = loss_report if loss_report is not None else LossReport()
policy = getattr(self, "_hidden_content", None) or resolve_hidden_content_policy()
conv_id = raw.get("id") or ""
records: list[dict] = raw.get("_records") or []
subagents: dict = raw.get("_subagents") or {}
title = _extract_title(records)
# Append the repos this session touched, e.g. "… [repo-a, repo-b]", so
# sessions launched from a workspace root (which all land in one
# folder-named notebook) stay scannable and searchable.
launch_cwd = _extract_launch_cwd(records)
if launch_cwd:
repos = _repos_touched(records, launch_cwd)
if repos:
title = f"{title} [{', '.join(repos)}]"
project = _extract_project(records, raw.get("_path"))
created_at = next(
(r.get("timestamp") for r in records if r.get("timestamp")), ""
)
updated_at = raw.get("_mtime_iso") or next(
(r.get("timestamp") for r in reversed(records) if r.get("timestamp")), ""
)
messages = _extract_messages(records, conv_id, report, policy, subagents)
for _ in messages:
report.record_message()
report.record_conversation()
return {
"id": conv_id,
"title": title,
"provider": self.provider_name,
"project": project,
"created_at": created_at or "",
"updated_at": updated_at or "",
"message_count": len(messages),
"messages": messages,
}
# ------------------------------------------------------------------
# Scanning
# ------------------------------------------------------------------
def _scan(self) -> list[dict]:
existing = [d for d in self._projects_dirs if d.is_dir()]
if not existing:
logger.warning(
"[claude-code] No projects directory exists: %s",
", ".join(str(d) for d in self._projects_dirs),
)
return []
# conv_id → (mtime, conv dict). Roots are merged by launch-folder; if the
# same session UUID appears in two roots (e.g. a live dir and a backup),
# the newer-mtime copy wins so we never emit two entries for one session.
by_id: dict[str, tuple[float, dict]] = {}
for root in existing:
for proj_dir in sorted(p for p in root.iterdir() if p.is_dir()):
for session_file in sorted(proj_dir.glob("*.jsonl")):
try:
stat = session_file.stat()
except OSError:
continue
if stat.st_size == 0:
continue
conv_id = session_file.stem
prev = by_id.get(conv_id)
if prev is not None and prev[0] >= stat.st_mtime:
logger.debug(
"[claude-code] Duplicate session %s in %s; keeping newer copy",
conv_id[:8], root,
)
continue
self._path_map[conv_id] = session_file
title, project, created = _read_session_meta(session_file)
by_id[conv_id] = (
stat.st_mtime,
{
"id": conv_id,
"title": title,
"project": project,
# The --project filter and dry-run table read the
# listing dict, not the normalized conversation.
"_project_name": project,
"created_at": created,
"updated_at": datetime.fromtimestamp(
stat.st_mtime, tz=timezone.utc
).isoformat(),
"_path": str(session_file),
"_project_dir": proj_dir.name,
},
)
# Deterministic order: by bucket dir, then session id.
return sorted(
(conv for _, conv in by_id.values()),
key=lambda c: (c["_project_dir"], c["id"]),
)
# ---------------------------------------------------------------------------
# Internal helpers
# ---------------------------------------------------------------------------
def _read_session_meta(path: Path) -> tuple[str, str | None, str]:
"""Light single-pass scan for listing metadata: (title, project, created_at).
Substring guards keep this cheap — only candidate lines are JSON-parsed.
The full-fidelity extraction happens later in normalize_conversation.
"""
title = ""
project: str | None = None
created = ""
try:
with path.open(encoding="utf-8") as fh:
for line in fh:
if '"ai-title"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
if rec.get("type") == "ai-title" and rec.get("aiTitle"):
title = str(rec["aiTitle"]) # last one wins
continue
if (not created or project is None) and (
'"timestamp"' in line or '"cwd"' in line
):
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
if not created and rec.get("timestamp"):
created = str(rec["timestamp"])
if project is None and rec.get("cwd"):
project = Path(rec["cwd"]).name or None
except OSError as e:
logger.warning("[claude-code] Could not read %s: %s", path, e)
return title, project, created
def _extract_title(records: list[dict]) -> str:
"""Last ai-title record wins; fall back to the first real user prompt."""
title = ""
for rec in records:
if rec.get("type") == "ai-title" and rec.get("aiTitle"):
title = str(rec["aiTitle"])
if title:
return title
for rec in records:
if rec.get("type") != "user" or rec.get("isMeta") or rec.get("isSidechain"):
continue
content = (rec.get("message") or {}).get("content")
if isinstance(content, str):
text = _strip_harness_noise(content)
elif isinstance(content, list):
text = " ".join(
_strip_harness_noise(item.get("text", ""))
for item in content
if isinstance(item, dict) and item.get("type") == "text"
).strip()
else:
text = ""
if text:
return text[:80]
return "Untitled session"
def _extract_project(records: list[dict], path: str | None) -> str | None:
"""Project = basename of the session's working directory."""
for rec in records:
cwd = rec.get("cwd")
if cwd:
name = Path(cwd).name
if name:
return name
# Fallback: the munged directory name (cannot be reliably de-munged
# because '-' is both the path separator and a legal name character).
if path:
return Path(path).parent.name.lstrip("-") or None
return None
def _parse_jsonl(path: Path) -> list[dict]:
"""Read a JSONL file into records, tolerating (and logging) bad lines."""
records: list[dict] = []
bad_lines = 0
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
logger.warning("[claude-code] Could not read %s: %s", path, e)
return records
# split("\n"), not splitlines(): splitlines() also breaks on U+0085,
# U+2028/9 and friends, which are legal *inside* a JSON string. A NEL in
# captured command output shreds one record into unparseable fragments and
# loses it silently (observed 2026-08-18 in a real Codex rollout).
for line in text.split("\n"):
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError:
bad_lines += 1
if bad_lines:
logger.warning(
"[claude-code] %s: skipped %d unparseable line(s)", path.name, bad_lines
)
return records
def _load_subagents(session_path: Path) -> dict:
"""Map spawning tool_use id → ``{"meta": {...}, "records": [...]}``.
Claude Code writes subagent transcripts to
``<session>/subagents/agent-*.jsonl`` with an ``agent-*.meta.json`` sidecar
carrying ``toolUseId`` (which parent ``Task``/``Agent`` call spawned it).
Files without a readable meta (no id to position them) are skipped.
"""
submap: dict[str, dict] = {}
subdir = session_path.parent / session_path.stem / "subagents"
if not subdir.is_dir():
return submap
for jf in sorted(subdir.glob("*.jsonl")):
meta_path = jf.with_suffix(".meta.json")
try:
meta = json.loads(meta_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
logger.warning(
"[claude-code] Subagent %s has no readable meta; skipping", jf.name
)
continue
tool_id = meta.get("toolUseId")
if not tool_id:
continue
submap[tool_id] = {"meta": meta, "records": _parse_jsonl(jf)}
return submap
def _extract_launch_cwd(records: list[dict]) -> str | None:
"""The session's working directory (constant per session; first cwd wins)."""
for rec in records:
cwd = rec.get("cwd")
if cwd:
return cwd
return None
# tool_use input keys that carry a file path.
_TOOL_PATH_KEYS = ("file_path", "path", "notebook_path")
# Optional ignore-list: git repos to never tag (comma-separated names). Rarely
# needed with git-root detection (config/reference dirs are already excluded
# because they aren't git repos), but kept as an escape hatch.
def _ignored_repos() -> set[str]:
env = os.getenv("CLAUDE_CODE_REPO_TAG_IGNORE", "")
return {s.strip() for s in env.split(",") if s.strip()}
# Absolute path-like tokens inside Bash command strings (file_path/path keys are
# matched directly). git-root resolution short-circuits on non-repo paths.
_ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+")
def _repos_touched(
records: list[dict], launch_cwd: str, cap: int = 3, ignore: set[str] | None = None
) -> list[str]:
"""Repos a session touched, for the title's ``[repo-a, repo-b]`` tag.
A file's repo is the **git repository it lives in** (nearest ancestor with a
``.git``), resolved for every absolute path seen in a tool_use input
(``file_path``/``path``/``notebook_path`` and absolute home paths inside Bash
``command`` strings). This works anywhere in the filesystem — not just under
the launch directory — so cross-workspace work is captured, and non-repo
noise (config dirs, one-off files, reference dirs) is excluded because it
isn't a git repo. Ordered by touch frequency, capped with a trailing ``…``.
"""
ignore = _ignored_repos() if ignore is None else ignore
counts: Counter = Counter()
seen: set[str] = set()
def note(p) -> None:
if not isinstance(p, str) or not p:
return
if not p.startswith("/"): # resolve relative paths against the launch cwd
if not launch_cwd:
return
p = str(Path(launch_cwd) / p)
if p in seen:
return
seen.add(p)
name = git_root_name(Path(p))
if name and not name.startswith(".") and name not in ignore:
counts[name] += 1
for rec in records:
msg = rec.get("message") or {}
content = msg.get("content")
if not isinstance(content, list):
continue
for item in content:
if not isinstance(item, dict) or item.get("type") != "tool_use":
continue
inp = item.get("input")
if not isinstance(inp, dict):
continue
for key in _TOOL_PATH_KEYS:
note(inp.get(key))
cmd = inp.get("command")
if isinstance(cmd, str):
for m in _ABS_PATH_RE.finditer(cmd):
note(m.group(0))
ordered = [name for name, _ in counts.most_common()]
if len(ordered) > cap:
return ordered[:cap] + ["…"]
return ordered
def _extract_messages(
records: list[dict],
conv_id: str,
report: LossReport,
policy: str,
subagents: dict | None = None,
include_sidechain: bool = False,
expanding: frozenset[str] = frozenset(),
) -> list[dict]:
"""Normalize Claude Code records into messages.
``subagents`` maps a spawning ``Task``/``Agent`` tool_use id to its separate
transcript ``{"meta": ..., "records": ...}``; when a matching tool_use is
seen its subagent is folded inline as a subagent block (extracted
recursively under the same policy). ``include_sidechain`` is set True for
those recursive subagent passes (subagent records are flagged
``isSidechain``); the top-level pass keeps skipping sidechain records so a
subagent is never also emitted as a stray top-level turn.
``expanding`` holds the spawn ids of the subagents being folded around
this pass. A ``fork`` subagent's transcript opens with a copy of the parent
turn that spawned it, its own spawn call included; that copy is dropped,
since the enclosing subagent block already stands for it. Expanding it
again recursed without end (RecursionError, every daily sync from
2026-09-24).
"""
subagents = subagents or {}
messages: list[dict] = []
# Pending collapsed tool activity: name → call count, plus total bytes.
pending_tools: Counter = Counter()
pending_bytes = 0
def flush_pending() -> None:
nonlocal pending_bytes
if not pending_tools:
return
shown = ", ".join(
f"{name} ×{count}" for name, count in pending_tools.most_common(_TOOL_NAMES_SHOWN)
)
if len(pending_tools) > _TOOL_NAMES_SHOWN:
shown += ", …"
calls = sum(pending_tools.values())
block = make_collapsed_block(
origin=f"{calls} calls: {shown}",
content_type="tool_activity",
size_bytes=pending_bytes,
kind=COLLAPSED_KIND_TOOL_DUMP,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(block)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None, "blocks": [block]}
)
pending_tools.clear()
pending_bytes = 0
for rec in records:
if rec.get("type") not in ("user", "assistant"):
continue
if rec.get("isMeta"):
continue
if rec.get("isSidechain") and not include_sidechain:
continue
msg = rec.get("message") or {}
role = msg.get("role") or rec["type"]
content = msg.get("content")
blocks: list[dict] = []
# This record's own tool traffic — merged into pending only after the
# record's message is appended, so the placeholder lands on it (not on
# the previous message).
local_tools: Counter = Counter()
local_bytes = 0
if isinstance(content, str):
text = _strip_harness_noise(content)
block = make_text_block(text)
if block:
blocks.append(block)
elif isinstance(content, list):
for item in content:
if not isinstance(item, dict):
continue
item_type = item.get("type", "")
if item_type == "text":
block = make_text_block(_strip_harness_noise(item.get("text", "")))
if block:
blocks.append(block)
elif item_type in ("thinking", "redacted_thinking"):
if policy == HIDDEN_CONTENT_FULL:
block = make_thinking_block(
item.get("thinking") or item.get("text") or ""
)
if block:
blocks.append(block)
else:
# Decision 2026-06-12: thinking is dropped without a
# placeholder, but stays visible in the run summary.
report.record_collapsed(
"thinking", len(json.dumps(item, default=str))
)
elif item_type == "tool_use":
name = item.get("name") or "tool"
tool_id = item.get("id")
if tool_id in expanding:
continue
if name in _SUBAGENT_TOOL_NAMES and tool_id in subagents:
# A Task/Agent spawn: fold its separate transcript inline
# instead of collapsing it. Its own tool traffic is
# collapsed by the recursive pass under the same policy.
sub = subagents[tool_id]
meta = sub.get("meta") or {}
sub_msgs = _extract_messages(
sub.get("records") or [],
conv_id,
report,
policy,
subagents,
include_sidechain=True,
expanding=expanding | {tool_id},
)
blocks.append(
make_subagent_block(
agent_type=meta.get("agentType") or name,
description=meta.get("description") or "",
messages=sub_msgs,
)
)
elif policy == HIDDEN_CONTENT_FULL:
blocks.append(
make_tool_use_block(
item.get("name", ""), item.get("input"), item.get("id")
)
)
else:
size = len(json.dumps(item, default=str))
local_tools[name] += 1
local_bytes += size
report.record_collapsed(name, size)
elif item_type == "tool_result":
if policy == HIDDEN_CONTENT_FULL:
blocks.append(
make_tool_result_block(
_stringify_tool_result(item.get("content")),
is_error=bool(item.get("is_error")),
)
)
else:
size = len(json.dumps(item, default=str))
local_bytes += size
report.record_collapsed("tool_result", size)
elif item_type == "image":
blocks.append(
make_image_placeholder(ref="embedded image", source="user_upload")
)
else:
logger.warning(
"[claude-code] Unknown content block type %r in session %s",
item_type,
conv_id[:8],
)
report.record_unknown(f"claude-code.{item_type or '?'}")
blocks.append(
make_unknown_block(
raw_type=f"claude-code.{item_type or '?'}",
observed_keys=list(item.keys()),
reason=UNKNOWN_REASON_UNKNOWN_TYPE,
)
)
if not blocks:
# Tool-result-only records (and similar): traffic accumulates and
# is attached to the message that initiated it.
pending_tools.update(local_tools)
pending_bytes += local_bytes
continue
# Attach previous records' tool activity to the previous message
# before starting a new one, so the placeholder lands between the
# dialogue turns it actually occurred between.
flush_pending()
messages.append(
{
"role": role,
"content_type": "text",
"timestamp": rec.get("timestamp"),
"blocks": blocks,
}
)
pending_tools.update(local_tools)
pending_bytes += local_bytes
flush_pending()
return messages
def _stringify_tool_result(content) -> str:
"""tool_result content may be a string or a list of text blocks."""
if isinstance(content, str):
return content
if isinstance(content, list):
parts = []
for item in content:
if isinstance(item, dict) and item.get("type") == "text":
parts.append(item.get("text", ""))
else:
parts.append(json.dumps(item, default=str))
return "\n".join(parts)
return json.dumps(content, default=str) if content is not None else ""
+881
View File
@@ -0,0 +1,881 @@
"""Codex CLI session provider — archives local agent transcripts.
Reads JSONL rollout files from ``~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl``
(override with ``CODEX_DIR``). Like Claude Code: no tokens, no rate limits, no
ToS risk — the data is local and Codex may prune it.
Local-only, deliberately
------------------------
Codex Cloud tasks (``codex cloud``) live server-side at
``https://chatgpt.com/backend-api/api/codex/tasks{,/list}`` — the same host and
``/backend-api`` root the ChatGPT provider already speaks. **CLI sessions are
never uploaded there**, so the cloud API is not an alternative source for these
transcripts and this provider does not talk to the network. Verified 2026-08-18
against Codex 0.147.0. If ``codex cloud exec`` ever enters regular use, those
transcripts *would* be cloud-only and would need a separate provider.
Two representations, one file (measured 2026-08-18 over 7 sessions / 0.147.0)
----------------------------------------------------------------------------
Every rollout line is ``{timestamp, ordinal, type, payload}``. Dialogue appears
twice, in two different shapes, and we parse the **typed** one:
* ``response_item`` — the model-facing wire format (mirrors the OpenAI Responses
API). Tool calls arrive as *JavaScript source* because Codex's ``exec`` tool is
code-mode::
const r = await tools.exec_command({"cmd":"git status","workdir":"/x", …});
text(r.output);
Exactly one ``tools.*`` call per invocation; three functions observed:
``exec_command`` (180), ``web__run`` (15), ``apply_patch`` (13).
* ``event_msg`` / ``item_completed`` — Codex's own typed items, already decoded:
``UserMessage``, ``AgentMessage``, ``Reasoning``, ``CommandExecution``,
``FileChange``, ``Extension``, ``ContextCompaction``.
The typed layer wins on every axis that matters here. It is 1:1 with the raw
layer for prose (91 ``AgentMessage`` ↔ 91 assistant messages, same ids; 345
``Reasoning`` ↔ 345), it hands us structured command/exit-code/output fields
instead of JS we would have to regex, and it pre-filters harness plumbing for
free: all 51 ``developer``-role messages (skills manifests, ``<multi_agent_mode>``,
"Approved command prefix saved") plus the 7 ``# AGENTS.md instructions…``
injections and 1 ``<environment_context>`` have no typed item. That is the same
noise ``claude_code._HARNESS_TAG_RE`` strips by hand.
Its one weakness: it records what *ran*, not what was *attempted*. 26 of 180
``exec_command`` calls produced no ``CommandExecution`` item — 14 sandbox launch
failures (``bwrap: loopback: Failed RTM_NEWADDR``), 6 user aborts ("aborted by
user after 504.6s"), ~5 still running at turn end, 1 "Script failed". So we read
the raw layer *only* to count attempts, and the collapsed placeholder reports the
shortfall ("3 did not complete") rather than silently under-reporting. ``wait``
function calls (50) are process polls, not attempts, and are not counted.
Reasoning is unrecoverable
--------------------------
All 345 reasoning items carry ``encrypted_content``; ``summary`` is ``[]`` in the
raw layer and ``summary_text``/``raw_content`` are empty in the typed layer, in
every session. Codex does not persist readable reasoning locally. Thinking is
therefore always dropped and counted — the same end state as the Claude Code
policy (decision 2026-06-12), but by necessity rather than by choice, so even
``full`` cannot surface it.
Why the sidecar SQLite is not read
----------------------------------
``~/.codex/state_5.sqlite`` carries a ``threads`` table (title, cwd, model,
tokens_used, rollout_path) and ``thread_history_1.sqlite`` a projection of the
items — but ``thread_history_projection_state`` tracks a byte offset *into the
rollout file*, i.e. the JSONL is canonical and SQLite is derived. Its ``title``
is just the first user message truncated (identical to ``first_user_message`` and
``preview``), so it offers nothing the JSONL lacks, and its filename carries a
schema version that will churn. We read the files.
Rendering follows the EXPORTER_HIDDEN_CONTENT policy: prose-only by default
(dialogue kept, tool traffic collapsed to one grouped placeholder per activity
run, reasoning dropped); ``full`` keeps the decoded tool calls and their output.
Subagents: Codex 0.147.0's ``thread_spawn_edges`` table exists but is empty and no
sub-transcripts were observed, so there is no subagent folding here (contrast
``claude_code._load_subagents``). If spawned agents start appearing they will
arrive as new item types and land in the loss report as unknowns.
"""
import json
import logging
import os
import re
from collections import Counter
from datetime import datetime, timezone
from pathlib import Path
from src.blocks import (
COLLAPSED_KIND_HIDDEN_CONTEXT,
COLLAPSED_KIND_TOOL_DUMP,
UNKNOWN_REASON_UNKNOWN_TYPE,
make_collapsed_block,
make_text_block,
make_tool_result_block,
make_tool_use_block,
make_unknown_block,
)
from src.loss_report import LossReport
from src.providers.base import (
BaseProvider,
HIDDEN_CONTENT_FULL,
ProviderError,
VALID_HIDDEN_CONTENT_POLICIES,
resolve_hidden_content_policy,
)
from src.utils import git_root_name
logger = logging.getLogger(__name__)
DEFAULT_SESSIONS_DIR = "~/.codex/sessions"
# rollout-<ISO-ish timestamp>-<uuid>.jsonl — the trailing UUID is the thread id.
_ROLLOUT_RE = re.compile(
r"^rollout-\d{4}-\d{2}-\d{2}T[\d-]+-"
r"([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12})$"
)
# The single tools.* call inside an `exec` custom_tool_call's JavaScript body.
_TOOLS_CALL_RE = re.compile(r"tools\.([A-Za-z_][\w.]*)\s*\(")
# Typed items that represent tool traffic, mapped to the label used in the
# collapsed placeholder. These labels must match the ones derived from the raw
# layer in _attempt_label, or the attempted-vs-completed delta is meaningless.
_TOOL_ITEM_LABELS = {
"CommandExecution": "exec_command",
"FileChange": "apply_patch",
# Extension is labelled by its `kind` (e.g. "web.search"); see _tool_label.
}
# function_call names that poll an already-running process rather than starting
# new work. Counting them as attempts would inflate the shortfall.
_POLLING_FUNCTIONS = {"wait"}
# How many distinct tool names to list in a collapsed-activity placeholder.
_TOOL_NAMES_SHOWN = 4
# Harness-injected user text. The typed layer already omits these (they have no
# UserMessage item), so this is a belt-and-braces guard for other Codex versions.
_HARNESS_USER_RE = re.compile(
r"^\s*(?:#\s*AGENTS\.md instructions for\b|<environment_context>)",
)
def resolve_roots(sessions_dir=None) -> list[Path]:
"""Resolve the ordered list of Codex ``sessions/`` roots to scan.
Precedence:
1. Explicit ``sessions_dir`` (a single path or a list) — used by tests.
2. ``CODEX_DIR`` env, split on ``os.pathsep`` (``:``) so multiple roots can
be given; a single path (no separator) stays backward compatible. Falls
back to the default root when unset.
3. Additionally, if ``CODEX_HOME`` is set (Codex's own name for its state
directory), its ``sessions/`` subdir is appended.
Roots are expanded and de-duplicated (order preserved); existence is checked
by the caller / ``_scan``.
"""
raw: list[str]
if sessions_dir is not None:
raw = (
[str(p) for p in sessions_dir]
if isinstance(sessions_dir, (list, tuple))
else [str(sessions_dir)]
)
else:
env = os.getenv("CODEX_DIR")
raw = env.split(os.pathsep) if env else [DEFAULT_SESSIONS_DIR]
codex_home = os.getenv("CODEX_HOME")
if codex_home:
raw.append(str(Path(codex_home) / "sessions"))
roots: list[Path] = []
for p in raw:
p = p.strip()
if not p:
continue
path = Path(p).expanduser()
if path not in roots:
roots.append(path)
return roots
class CodexProvider(BaseProvider):
"""Local-file provider over Codex CLI rollout transcripts."""
provider_name = "codex"
def __init__(
self,
sessions_dir: str | Path | None = None,
hidden_content: str | None = None,
) -> None:
super().__init__()
self._sessions_dirs = resolve_roots(sessions_dir)
self._hidden_content = (
hidden_content
if hidden_content in VALID_HIDDEN_CONTENT_POLICIES
else resolve_hidden_content_policy()
)
# conv_id → rollout file path, populated by _scan()
self._path_map: dict[str, Path] = {}
# ------------------------------------------------------------------
# BaseProvider interface
# ------------------------------------------------------------------
def list_conversations(self, offset: int = 0, limit: int = 100) -> list[dict]:
full = self._scan()
return full[offset : offset + limit]
def fetch_all_conversations(self, since: datetime | None = None) -> list[dict]:
convs = self._scan()
if since is not None:
since_aware = since if since.tzinfo else since.replace(tzinfo=timezone.utc)
convs = [
c for c in convs
if datetime.fromisoformat(c["updated_at"]) >= since_aware
]
logger.info(
"[codex] Found %d session(s) under %s",
len(convs),
", ".join(str(d) for d in self._sessions_dirs),
)
return convs
def get_conversation(self, conv_id: str) -> dict:
path = self._path_map.get(conv_id)
if path is None:
# Direct call without a prior listing (e.g. tests) — scan first.
self._scan()
path = self._path_map.get(conv_id)
if path is None or not path.exists():
raise ProviderError(
self.provider_name,
f"get_conversation({conv_id[:8]})",
FileNotFoundError(f"No rollout file for id {conv_id}"),
)
records = _parse_jsonl(path)
mtime = datetime.fromtimestamp(path.stat().st_mtime, tz=timezone.utc)
return {
"id": conv_id,
"_path": str(path),
"_records": records,
# Listing and normalized updated_at must match, or the cache
# staleness comparison would re-export every session every run.
"_mtime_iso": mtime.isoformat(),
}
def normalize_conversation(self, raw: dict, loss_report: LossReport | None = None) -> dict:
report = loss_report if loss_report is not None else LossReport()
policy = getattr(self, "_hidden_content", None) or resolve_hidden_content_policy()
conv_id = raw.get("id") or ""
records: list[dict] = raw.get("_records") or []
title = _extract_title(records)
# Append the repos this session touched, e.g. "… [repo-a, repo-b]".
# Codex sessions are commonly all launched from one workspace root, so
# without this every session lands in the same notebook with no way to
# tell them apart.
launch_cwd = _extract_launch_cwd(records)
if launch_cwd:
repos = _repos_touched(records, launch_cwd)
if repos:
title = f"{title} [{', '.join(repos)}]"
project = _extract_project(records)
created_at = _extract_created_at(records)
updated_at = raw.get("_mtime_iso") or next(
(r.get("timestamp") for r in reversed(records) if r.get("timestamp")), ""
)
messages = _extract_messages(records, conv_id, report, policy)
for _ in messages:
report.record_message()
report.record_conversation()
return {
"id": conv_id,
"title": title,
"provider": self.provider_name,
"project": project,
"created_at": created_at or "",
"updated_at": updated_at or "",
"message_count": len(messages),
"messages": messages,
}
# ------------------------------------------------------------------
# Scanning
# ------------------------------------------------------------------
def _scan(self) -> list[dict]:
existing = [d for d in self._sessions_dirs if d.is_dir()]
if not existing:
logger.warning(
"[codex] No sessions directory exists: %s",
", ".join(str(d) for d in self._sessions_dirs),
)
return []
# conv_id → (mtime, conv dict). If the same thread id appears under two
# roots (e.g. a live dir and a backup), the newer-mtime copy wins.
by_id: dict[str, tuple[float, dict]] = {}
for root in existing:
# Rollouts are filed under YYYY/MM/DD; rglob keeps us agnostic to
# that layout in case Codex reorganises it.
for session_file in sorted(root.rglob("rollout-*.jsonl")):
match = _ROLLOUT_RE.match(session_file.stem)
if not match:
logger.debug("[codex] Skipping unrecognised filename %s", session_file.name)
continue
try:
stat = session_file.stat()
except OSError:
continue
if stat.st_size == 0:
continue
conv_id = match.group(1)
prev = by_id.get(conv_id)
if prev is not None and prev[0] >= stat.st_mtime:
logger.debug(
"[codex] Duplicate session %s in %s; keeping newer copy",
conv_id[:8], root,
)
continue
self._path_map[conv_id] = session_file
title, project, created = _read_session_meta(session_file)
by_id[conv_id] = (
stat.st_mtime,
{
"id": conv_id,
"title": title,
"project": project,
# The --project filter and dry-run table read the
# listing dict, not the normalized conversation.
"_project_name": project,
"created_at": created,
"updated_at": datetime.fromtimestamp(
stat.st_mtime, tz=timezone.utc
).isoformat(),
"_path": str(session_file),
},
)
# Deterministic order: by creation date bucket, then thread id.
return sorted(
(conv for _, conv in by_id.values()),
key=lambda c: (c["created_at"], c["id"]),
)
# ---------------------------------------------------------------------------
# Internal helpers
# ---------------------------------------------------------------------------
def _parse_jsonl(path: Path) -> list[dict]:
"""Read a JSONL file into records, tolerating (and logging) bad lines."""
records: list[dict] = []
bad_lines = 0
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
logger.warning("[codex] Could not read %s: %s", path, e)
return records
# split("\n"), not splitlines(): splitlines() also breaks on U+0085,
# U+2028/9 and friends, which are legal *inside* a JSON string. A NEL in
# captured command output shreds one record into unparseable fragments and
# loses it silently (observed 2026-08-18 in a real Codex rollout).
for line in text.split("\n"):
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError:
bad_lines += 1
if bad_lines:
logger.warning("[codex] %s: skipped %d unparseable line(s)", path.name, bad_lines)
return records
def _item(rec: dict) -> dict | None:
"""Return the typed item from an ``event_msg``/``item_completed`` record."""
if rec.get("type") != "event_msg":
return None
payload = rec.get("payload") or {}
if payload.get("type") != "item_completed":
return None
item = payload.get("item")
return item if isinstance(item, dict) else None
def _item_kind(item: dict) -> str:
"""Typed item discriminator. 0.147.0 uses ``type``; ``item_type`` is a hedge."""
return str(item.get("item_type") or item.get("type") or "")
def _item_text(item: dict) -> str:
"""Concatenate the text of a UserMessage / AgentMessage item.
The two disagree on case — ``UserMessage`` blocks are ``"text"`` and
``AgentMessage`` blocks are ``"Text"`` — so the comparison is case-folded.
"""
content = item.get("content")
if isinstance(content, str):
return content.strip()
if not isinstance(content, list):
return ""
parts = []
for block in content:
if isinstance(block, dict) and str(block.get("type", "")).lower() == "text":
text = block.get("text")
if isinstance(text, str):
parts.append(text)
return "".join(parts).strip()
def _read_session_meta(path: Path) -> tuple[str, str | None, str]:
"""Light single-pass scan for listing metadata: (title, project, created_at).
Substring guards keep this cheap — only candidate lines are JSON-parsed, and
the scan stops as soon as the title is found (the first UserMessage is
usually within the first few dozen lines of a multi-megabyte file).
"""
title = ""
project: str | None = None
created = ""
try:
with path.open(encoding="utf-8") as fh:
for line in fh:
if not created and '"session_meta"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
payload = rec.get("payload") or {}
created = str(payload.get("timestamp") or rec.get("timestamp") or "")
cwd = payload.get("cwd")
if cwd:
project = Path(cwd).name or None
continue
if not title and '"UserMessage"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
item = _item(rec)
if item is None or _item_kind(item) != "UserMessage":
continue
text = _item_text(item)
if text and not _HARNESS_USER_RE.match(text):
title = text[:80]
break
except OSError as e:
logger.warning("[codex] Could not read %s: %s", path, e)
return title or "Untitled session", project, created
def _extract_title(records: list[dict]) -> str:
"""Codex has no AI-generated title — the first real user prompt is the title.
(``state_5.sqlite``'s ``title`` column is this same string truncated; see the
module docstring on why the sidecar DB is not consulted.)
"""
for rec in records:
item = _item(rec)
if item is None or _item_kind(item) != "UserMessage":
continue
text = _item_text(item)
if text and not _HARNESS_USER_RE.match(text):
return text[:80]
return "Untitled session"
def _session_meta_payload(records: list[dict]) -> dict:
for rec in records:
if rec.get("type") == "session_meta":
payload = rec.get("payload")
if isinstance(payload, dict):
return payload
return {}
def _extract_launch_cwd(records: list[dict]) -> str | None:
"""The session's working directory (constant per session)."""
cwd = _session_meta_payload(records).get("cwd")
return cwd if isinstance(cwd, str) and cwd else None
def _extract_project(records: list[dict]) -> str | None:
"""Project = basename of the session's working directory."""
cwd = _extract_launch_cwd(records)
if cwd:
return Path(cwd).name or None
return None
def _extract_created_at(records: list[dict]) -> str:
"""Session start time.
Prefers ``session_meta.payload.timestamp`` — the outer line ``timestamp`` is
when the record was *flushed*, which can trail the true start by minutes.
"""
payload = _session_meta_payload(records)
ts = payload.get("timestamp")
if isinstance(ts, str) and ts:
return ts
return next((r.get("timestamp") for r in records if r.get("timestamp")), "") or ""
def _ignored_repos() -> set[str]:
"""Optional ignore-list: git repos to never tag (comma-separated names)."""
env = os.getenv("CODEX_REPO_TAG_IGNORE", "")
return {s.strip() for s in env.split(",") if s.strip()}
def _strip_file_uri(path: str) -> str:
"""``CommandExecution.cwd`` is a ``file://`` URI; every other path is plain."""
return path[7:] if path.startswith("file://") else path
# Absolute path-like tokens inside command strings.
_ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+")
def _repos_touched(
records: list[dict], launch_cwd: str, cap: int = 3, ignore: set[str] | None = None
) -> list[str]:
"""Repos a session touched, for the title's ``[repo-a, repo-b]`` tag.
A file's repo is the git repository it lives in (nearest ancestor with a
``.git``). Paths come from the typed items: ``FileChange.changes`` keys
(absolute), and ``CommandExecution``'s argv plus its ``cwd``. Ordered by
touch frequency, capped with a trailing ``…``.
"""
ignore = _ignored_repos() if ignore is None else ignore
counts: Counter = Counter()
seen: set[str] = set()
def note(p) -> None:
if not isinstance(p, str) or not p:
return
p = _strip_file_uri(p)
if not p.startswith("/"): # resolve relative paths against the launch cwd
if not launch_cwd:
return
p = str(Path(launch_cwd) / p)
if p in seen:
return
seen.add(p)
name = git_root_name(Path(p))
if name and not name.startswith(".") and name not in ignore:
counts[name] += 1
for rec in records:
item = _item(rec)
if item is None:
continue
kind = _item_kind(item)
if kind == "FileChange":
changes = item.get("changes")
if isinstance(changes, dict):
for file_path in changes:
note(file_path)
elif kind == "CommandExecution":
note(item.get("cwd"))
command = item.get("command")
if isinstance(command, list) and command:
tail = command[-1]
if isinstance(tail, str):
for m in _ABS_PATH_RE.finditer(tail):
note(m.group(0))
ordered = [name for name, _ in counts.most_common()]
if len(ordered) > cap:
return ordered[:cap] + ["…"]
return ordered
def _tool_label(item: dict, kind: str) -> str:
"""Placeholder label for a completed tool item.
``Extension`` covers everything routed through the model's own extensions
(``web.search`` so far), so its ``kind`` field is the useful name.
"""
if kind == "Extension":
return str(item.get("kind") or "extension")
return _TOOL_ITEM_LABELS.get(kind, kind)
def _attempt_label(payload: dict) -> str | None:
"""Label for a raw tool call, or None if it should not count as an attempt.
``exec`` custom_tool_calls carry JavaScript; the inner ``tools.<fn>`` name is
what lines up with the typed items' labels. ``wait`` polls an already-running
process and starts no new work.
"""
ptype = payload.get("type")
name = payload.get("name") or ""
if ptype == "custom_tool_call":
if name == "exec":
match = _TOOLS_CALL_RE.search(payload.get("input") or "")
if match:
# tools.web__run → the Extension item calls itself "web.search";
# both are "one web call", which is all the count claims.
return match.group(1)
return "exec"
return name or "tool"
if ptype == "function_call":
if name in _POLLING_FUNCTIONS:
return None
return name or "function"
return None
def _command_string(item: dict) -> str:
"""Human-readable command from a ``CommandExecution`` argv list.
The argv is ``["/bin/bash", "-lc", "<script>"]``; the script is the part
worth showing.
"""
command = item.get("command")
if isinstance(command, list):
if len(command) >= 3 and command[0].endswith("sh") and command[1] in ("-lc", "-c"):
return str(command[-1])
return " ".join(str(c) for c in command)
return str(command or "")
def _extract_messages(
records: list[dict],
conv_id: str,
report: LossReport,
policy: str,
) -> list[dict]:
"""Normalize Codex rollout records into messages.
Reads the typed ``item_completed`` layer for content and the raw
``response_item`` layer only to count attempted tool calls, so a collapsed
placeholder can report calls that never produced an item (sandbox failures,
user aborts, still-running processes). See the module docstring.
"""
messages: list[dict] = []
# Pending collapsed tool activity for the current run.
pending_tools: Counter = Counter()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
def flush_pending() -> None:
nonlocal pending_bytes, pending_attempted, pending_completed
# An activity run with only failed attempts still deserves a placeholder
# — silence would imply nothing happened.
shortfall = max(pending_attempted - pending_completed, 0)
if not pending_tools and not shortfall:
return
if policy == HIDDEN_CONTENT_FULL:
# Under `full` the completed calls are already rendered as tool
# blocks, so only the shortfall is left to report — and a *collapsed*
# block would render the "omitted (set full to keep)" suffix, which
# contradicts the policy in force. Say it plainly instead.
report.record_collapsed("tool_call_incomplete", 0)
note = make_tool_result_block(
f"{shortfall} tool call(s) produced no result — the process failed to "
"launch, was aborted, or was still running when the turn ended.",
tool_name="incomplete",
is_error=True,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(note)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None,
"blocks": [note]}
)
pending_tools.clear()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
return
if pending_tools:
shown = ", ".join(
f"{name} ×{count}"
for name, count in pending_tools.most_common(_TOOL_NAMES_SHOWN)
)
if len(pending_tools) > _TOOL_NAMES_SHOWN:
shown += ", …"
origin = f"{sum(pending_tools.values())} calls: {shown}"
else:
origin = "0 calls"
if shortfall:
origin += f" (+{shortfall} did not complete)"
report.record_collapsed("tool_call_incomplete", 0)
block = make_collapsed_block(
origin=origin,
content_type="tool_activity",
size_bytes=pending_bytes,
kind=COLLAPSED_KIND_TOOL_DUMP,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(block)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None, "blocks": [block]}
)
pending_tools.clear()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
def append_message(role: str, blocks: list[dict], timestamp) -> None:
# Attach the preceding run's tool activity to the previous message
# before starting a new one, so the placeholder lands between the
# dialogue turns it actually occurred between.
flush_pending()
messages.append(
{
"role": role,
"content_type": "text",
"timestamp": timestamp,
"blocks": blocks,
}
)
for rec in records:
rec_type = rec.get("type")
payload = rec.get("payload") or {}
timestamp = rec.get("timestamp")
# ── Raw layer: attempt counting only ──────────────────────────────
if rec_type == "response_item":
if payload.get("type") in ("custom_tool_call", "function_call"):
if _attempt_label(payload) is not None:
pending_attempted += 1
continue
item = _item(rec)
if item is None:
continue
kind = _item_kind(item)
# ── Dialogue ──────────────────────────────────────────────────────
if kind in ("UserMessage", "AgentMessage"):
text = _item_text(item)
if kind == "UserMessage" and _HARNESS_USER_RE.match(text):
# Defensive: 0.147.0 gives these no typed item at all.
report.record_filtered_role("codex.harness_injection")
continue
block = make_text_block(text)
if block:
append_message(
"user" if kind == "UserMessage" else "assistant", [block], timestamp
)
continue
# ── Reasoning: encrypted at rest, nothing to keep under any policy ─
if kind == "Reasoning":
report.record_collapsed("reasoning", len(json.dumps(item, default=str)))
continue
# ── Context compaction: a visible marker, not a silent drop ────────
if kind == "ContextCompaction":
flush_pending()
marker = make_collapsed_block(
origin="context compacted — earlier turns dropped from the model's context",
content_type="context_compaction",
size_bytes=0,
kind=COLLAPSED_KIND_HIDDEN_CONTEXT,
)
report.record_collapsed("context_compaction", 0)
messages.append(
{"role": "tool", "content_type": "text", "timestamp": timestamp,
"blocks": [marker]}
)
continue
# ── Tool traffic ──────────────────────────────────────────────────
if kind in ("CommandExecution", "FileChange", "Extension"):
pending_completed += 1
label = _tool_label(item, kind)
size = len(json.dumps(item, default=str))
if policy == HIDDEN_CONTENT_FULL:
use_block, result_block = _full_tool_blocks(item, kind, label)
target = messages[-1] if messages and messages[-1]["role"] == "assistant" else None
if target is None:
messages.append(
{"role": "tool", "content_type": "text",
"timestamp": timestamp, "blocks": []}
)
target = messages[-1]
target["blocks"].append(use_block)
if result_block:
target["blocks"].append(result_block)
else:
pending_tools[label] += 1
pending_bytes += size
report.record_collapsed(label, size)
continue
# ── Anything new in a future Codex version ────────────────────────
logger.warning("[codex] Unknown item type %r in session %s", kind, conv_id[:8])
report.record_unknown(f"codex.{kind or '?'}")
append_message(
"tool",
[
make_unknown_block(
raw_type=f"codex.{kind or '?'}",
observed_keys=list(item.keys()),
reason=UNKNOWN_REASON_UNKNOWN_TYPE,
)
],
timestamp,
)
flush_pending()
return messages
def _full_tool_blocks(item: dict, kind: str, label: str) -> tuple[dict, dict | None]:
"""Decoded tool_use / tool_result blocks for EXPORTER_HIDDEN_CONTENT=full.
The typed item already carries structured fields, so this reads the command,
exit code and output directly rather than parsing the raw layer's JavaScript.
"""
if kind == "CommandExecution":
exit_code = item.get("exit_code")
use = make_tool_use_block(
label,
{
"command": _command_string(item),
"cwd": _strip_file_uri(str(item.get("cwd") or "")),
"status": item.get("status"),
"exit_code": exit_code,
},
item.get("id"),
)
output = item.get("aggregated_output") or item.get("formatted_output") or ""
result = make_tool_result_block(
str(output),
tool_name=label,
is_error=bool(exit_code not in (0, None)),
)
return use, result
if kind == "FileChange":
changes = item.get("changes") if isinstance(item.get("changes"), dict) else {}
use = make_tool_use_block(
label,
{
"files": {
path: (change.get("type") if isinstance(change, dict) else "?")
for path, change in changes.items()
},
"status": item.get("status"),
},
item.get("id"),
)
output = "\n".join(
part for part in (item.get("stdout"), item.get("stderr")) if part
)
result = make_tool_result_block(
output or f"{len(changes)} file(s) changed",
tool_name=label,
is_error=item.get("status") not in (None, "completed", "success"),
)
return use, result
# Extension (web.search and anything else routed through an extension).
use = make_tool_use_block(
label,
{"query": item.get("query"), "action": item.get("action")},
item.get("id"),
)
results = item.get("results")
result = make_tool_result_block(
json.dumps(results, default=str, indent=1) if results else "",
tool_name=label,
)
return use, result
+61 -6
View File
@@ -50,7 +50,7 @@ def build_export_path(
created_at: ISO8601 creation timestamp (used for year folder).
filename: Already-generated filename from generate_filename().
structure: OUTPUT_STRUCTURE value. One of:
"provider/project/year" (default)
"provider/project/year" (default) — project and year combined, e.g. no-project.2025/
"provider/project"
"provider/year"
@@ -64,23 +64,39 @@ def build_export_path(
parts: list[str] = [provider]
if structure == "provider/project/year":
parts += [project_slug, year]
parts += [f"{project_slug}.{year}"]
elif structure == "provider/project":
parts += [project_slug]
elif structure == "provider/year":
parts += [year]
else:
# Unknown structure — fall back to default
parts += [project_slug, year]
parts += [f"{project_slug}.{year}"]
return base_dir.joinpath(*parts) / filename
def _is_sensitive_key(key: object) -> bool:
"""True if a mapping key names a secret.
Matches the whole key and each of its underscore/dash-separated words, so
compound names carry too: exact-match alone let ``access_token`` and
``api_key`` through into logged response bodies. Word-level matching keeps
innocent keys ("keywords", "monkey") intact.
"""
if not isinstance(key, str):
return False
lowered = key.lower()
if lowered in _SENSITIVE_KEYS:
return True
return any(part in _SENSITIVE_KEYS for part in re.split(r"[^a-z0-9]+", lowered))
def redact_secrets(data: object) -> object:
"""Recursively redact sensitive values from a dict/list for safe logging.
Keys matching _SENSITIVE_KEYS (case-insensitive) have their values
replaced with "[REDACTED]".
Keys naming a secret (see _is_sensitive_key) have their values replaced
with "[REDACTED]".
Args:
data: Any JSON-serializable object.
@@ -90,7 +106,7 @@ def redact_secrets(data: object) -> object:
"""
if isinstance(data, dict):
return {
k: "[REDACTED]" if k.lower() in _SENSITIVE_KEYS else redact_secrets(v)
k: "[REDACTED]" if _is_sensitive_key(k) else redact_secrets(v)
for k, v in data.items()
}
if isinstance(data, list):
@@ -155,3 +171,42 @@ def _parse_dt(ts: str) -> datetime:
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
# ---------------------------------------------------------------------------
# Git-repo resolution (shared by the local agent-transcript providers)
# ---------------------------------------------------------------------------
# dir Path → git-repo name it belongs to (or None). Process-wide; the working
# tree doesn't change under us mid-run, so caching walked dirs is safe.
_GIT_ROOT_CACHE: dict[Path, str | None] = {}
_CACHE_MISS = object()
def git_root_name(path: Path, max_steps: int = 25) -> str | None:
"""Name of the git repo ``path`` lives in — nearest ancestor with ``.git``.
Walks up from ``path`` until a ``.git`` entry is found (returns that dir's
basename) or the filesystem root is reached (returns ``None``). Disk-based:
a path in no git repo, or a repo no longer on disk, yields ``None``.
Used by the Claude Code and Codex providers to tag a session title with the
repos it touched, so sessions launched from a shared workspace root stay
distinguishable.
"""
cur = path
for _ in range(max_steps):
cached = _GIT_ROOT_CACHE.get(cur, _CACHE_MISS)
if cached is not _CACHE_MISS:
return cached
try:
if (cur / ".git").exists():
_GIT_ROOT_CACHE[cur] = cur.name
return cur.name
except OSError:
break
if cur.parent == cur: # filesystem root
break
cur = cur.parent
_GIT_ROOT_CACHE[path] = None
return None
+39
View File
@@ -0,0 +1,39 @@
"""Shared test fixtures.
The autouse fixture here exists because of a real incident (2026-08-18): the
`sync` CLI tests invoke the actual command, which calls `load_config()`, which
calls `load_dotenv()` — so the developer's real `.env` was loaded and its
`NTFY_TOPIC` used. Every `pytest` run fired real push notifications at the
developer's phone, including a fabricated "3 conversations failed to export"
from a fixture. Nothing appeared in the log to explain it, because the tests
pass `--no-log-file`.
The lesson generalises past ntfy: any test that exercises a command end to end
inherits whatever is in `.env` unless the environment is neutralised first.
"""
import pytest
@pytest.fixture(autouse=True)
def _no_outbound_side_effects(monkeypatch):
"""Neutralise every environment variable that could reach a real service.
Set to empty/unroutable values rather than deleted: `load_dotenv` is called
with ``override=False``, which only skips keys **already present** in the
environment. Deleting a key would let the developer's `.env` put it back.
Individual tests may still `monkeypatch.setenv` these — that is how the
notification tests point at a dead local port on purpose.
"""
# Notifications: an empty topic disables sending outright (`is_configured`
# strips and checks truthiness), and the server is pointed at a closed port
# so even a test that sets its own topic cannot reach the internet.
monkeypatch.setenv("NTFY_TOPIC", "")
monkeypatch.setenv("NTFY_SERVER", "http://127.0.0.1:9")
monkeypatch.setenv("NTFY_TOKEN", "")
# Joplin: a test that reached the developer's running desktop instance would
# create or overwrite real notes in their archive.
monkeypatch.setenv("JOPLIN_API_URL", "http://127.0.0.1:9")
monkeypatch.setenv("JOPLIN_API_TOKEN", "")
+148 -9
View File
@@ -8,12 +8,30 @@
"node-root": {
"id": "node-root",
"parent": null,
"children": ["node-1"],
"children": ["node-uec"],
"message": null
},
"node-uec": {
"id": "node-uec",
"parent": "node-root",
"children": ["node-1"],
"message": {
"id": "node-uec",
"author": {"role": "user"},
"create_time": null,
"content": {
"content_type": "user_editable_context",
"user_profile": "Preferred name: Jesse",
"user_instructions": "The user provided the additional info about how they would like you to respond:\n```Always cite sources.```"
},
"metadata": {
"is_visually_hidden_from_conversation": true
}
}
},
"node-1": {
"id": "node-1",
"parent": "node-root",
"parent": "node-uec",
"children": ["node-2"],
"message": {
"id": "node-1",
@@ -28,7 +46,7 @@
"node-2": {
"id": "node-2",
"parent": "node-1",
"children": ["node-3"],
"children": ["node-mm-user"],
"message": {
"id": "node-2",
"author": {"role": "assistant"},
@@ -39,19 +57,140 @@
}
}
},
"node-3": {
"id": "node-3",
"node-mm-user": {
"id": "node-mm-user",
"parent": "node-2",
"children": [],
"children": ["node-mm-assistant"],
"message": {
"id": "node-3",
"id": "node-mm-user",
"author": {"role": "user"},
"create_time": 1704067300.0,
"content": {
"content_type": "image_asset_pointer",
"parts": [{"content_type": "image_asset_pointer", "asset_pointer": "file://some-image"}]
"content_type": "multimodal_text",
"parts": [
{"content_type": "audio_transcription", "text": "What is the capital of France?", "direction": "in", "decoding_id": null},
{"content_type": "real_time_user_audio_video_asset_pointer", "frames_asset_pointers": [], "video_container_asset_pointer": null, "audio_asset_pointer": {"content_type": "audio_asset_pointer", "asset_pointer": "sediment://file_user001", "size_bytes": 50000, "format": "wav", "metadata": {"start": 0.0, "end": 2.5}}, "audio_start_timestamp": 1.0}
]
},
"metadata": {"voice_mode_message": true}
}
},
"node-mm-assistant": {
"id": "node-mm-assistant",
"parent": "node-mm-user",
"children": ["node-mm-user-rev"],
"message": {
"id": "node-mm-assistant",
"author": {"role": "assistant"},
"create_time": 1704067305.0,
"content": {
"content_type": "multimodal_text",
"parts": [
{"content_type": "audio_transcription", "text": "The capital of France is Paris.", "direction": "out", "decoding_id": null},
{"content_type": "audio_asset_pointer", "asset_pointer": "sediment://file_assistant001", "size_bytes": 80000, "format": "wav", "metadata": {"start": 0.0, "end": 3.2}}
]
}
}
},
"node-mm-user-rev": {
"id": "node-mm-user-rev",
"parent": "node-mm-assistant",
"children": ["node-image-only"],
"message": {
"id": "node-mm-user-rev",
"author": {"role": "user"},
"create_time": 1704067400.0,
"content": {
"content_type": "multimodal_text",
"parts": [
{"content_type": "real_time_user_audio_video_asset_pointer", "frames_asset_pointers": [], "video_container_asset_pointer": null, "audio_asset_pointer": {"content_type": "audio_asset_pointer", "asset_pointer": "sediment://file_user002", "size_bytes": 30000, "format": "wav", "metadata": {"start": 0.0, "end": 1.5}}, "audio_start_timestamp": 5.0},
{"content_type": "audio_transcription", "text": "Tell me more please.", "direction": "in", "decoding_id": null}
]
}
}
},
"node-image-only": {
"id": "node-image-only",
"parent": "node-mm-user-rev",
"children": ["node-exec-output"],
"message": {
"id": "node-image-only",
"author": {"role": "user"},
"create_time": 1704067500.0,
"content": {
"content_type": "multimodal_text",
"parts": [
{"content_type": "image_asset_pointer", "asset_pointer": "file-service://image001"}
]
}
}
},
"node-exec-output": {
"id": "node-exec-output",
"parent": "node-image-only",
"children": ["node-exec-output-empty"],
"message": {
"id": "node-exec-output",
"author": {"role": "tool", "name": "container.exec", "metadata": {}},
"create_time": 1704067600.0,
"content": {
"content_type": "execution_output",
"text": "Hello from container.exec\nLine 2 of output"
},
"metadata": {
"aggregate_result": {"status": "success", "messages": []},
"reasoning_title": "Reading skill documentation"
}
}
},
"node-exec-output-empty": {
"id": "node-exec-output-empty",
"parent": "node-exec-output",
"children": ["node-system-error"],
"message": {
"id": "node-exec-output-empty",
"author": {"role": "tool", "name": "python", "metadata": {}},
"create_time": 1704067610.0,
"content": {
"content_type": "execution_output",
"text": ""
},
"metadata": {}
}
},
"node-system-error": {
"id": "node-system-error",
"parent": "node-exec-output-empty",
"children": ["node-tether-spinner"],
"message": {
"id": "node-system-error",
"author": {"role": "tool", "name": "web", "metadata": {}},
"create_time": 1704067620.0,
"content": {
"content_type": "system_error",
"name": "tool_error",
"text": "Error: Error from browse service: Error calling browse service: 503"
},
"metadata": {}
}
},
"node-tether-spinner": {
"id": "node-tether-spinner",
"parent": "node-system-error",
"children": [],
"message": {
"id": "node-tether-spinner",
"author": {"role": "tool", "name": "file_search", "metadata": {}},
"create_time": 1704067630.0,
"content": {
"content_type": "tether_browsing_display",
"result": "",
"summary": "",
"assets": null,
"tether_id": null
},
"metadata": {"command": "spinner", "status": "running"}
}
}
}
}
+9
View File
@@ -30,6 +30,15 @@
"sender": "human",
"created_at": "2024-06-10T14:45:00.000Z",
"content": "Thank you, that helped!"
},
{
"uuid": "msg-004",
"sender": "human",
"created_at": "2024-06-10T14:50:00.000Z",
"content": [
{"type": "text", "text": "What about this image?"},
{"type": "image", "source": {"file_uuid": "claude-image-uuid-1", "media_type": "image/png"}}
]
}
]
}
+103
View File
@@ -151,3 +151,106 @@ class TestGetNewOrUpdated:
convs = [{"id": "a", "updated_at": "2024-06-01T00:00:00Z"}]
result = tmp_cache.get_new_or_updated("claude", convs)
assert len(result) == 1
def test_force_returns_all_cached_and_unchanged(self, tmp_cache):
tmp_cache.mark_exported("claude", "a", {"updated_at": "2024-01-01T00:00:00Z"})
convs = [
{"id": "a", "updated_at": "2024-01-01T00:00:00Z"}, # cached, unchanged
{"id": "b", "updated_at": "2024-01-02T00:00:00Z"}, # new
]
assert len(tmp_cache.get_new_or_updated("claude", convs)) == 1
assert len(tmp_cache.get_new_or_updated("claude", convs, force=True)) == 2
def test_force_orders_least_recently_exported_first(self, tmp_cache):
# 'a' was exported long ago; 'b' just now; 'c' never.
tmp_cache.mark_exported("claude", "a", {"updated_at": "2024-01-01T00:00:00Z"})
tmp_cache._data["claude"]["a"]["exported_at"] = "2024-01-01T00:00:00Z"
tmp_cache.mark_exported("claude", "b", {"updated_at": "2024-01-01T00:00:00Z"})
tmp_cache._data["claude"]["b"]["exported_at"] = "2026-06-12T00:00:00Z"
convs = [{"id": "a"}, {"id": "b"}, {"id": "c"}]
ordered = [c["id"] for c in tmp_cache.get_new_or_updated("claude", convs, force=True)]
# never-exported ('c', exported_at "") and oldest ('a') come before 'b'
assert ordered == ["c", "a", "b"]
def test_force_capped_run_makes_progress(self, tmp_cache):
"""The bug: force + cap repeated the same head every run. Now each
capped run re-exports the next-oldest batch and converges."""
convs = [{"id": f"c{i}", "updated_at": "2024-01-01T00:00:00Z"} for i in range(5)]
seen = set()
for _ in range(3): # cap=2 over 3 runs should cover all 5
batch = tmp_cache.get_new_or_updated("claude", convs, force=True)[:2]
for c in batch:
seen.add(c["id"])
tmp_cache.mark_exported("claude", c["id"], {"file_path": f"/{c['id']}.md"})
assert seen == {f"c{i}" for i in range(5)}
def test_campaign_excludes_already_refreshed(self, tmp_cache):
"""With a campaign stamp, conversations re-exported during the campaign
(exported_at >= stamp) drop out, so the remaining count shrinks to 0."""
convs = [{"id": f"c{i}"} for i in range(5)]
campaign = "2026-06-12T00:00:00+00:00"
# Nothing exported yet → all 5 are candidates.
assert len(tmp_cache.get_new_or_updated("claude", convs, force=True,
campaign_at=campaign)) == 5
# Re-export 2 "during" the campaign (after the stamp).
for cid in ("c0", "c1"):
tmp_cache.mark_exported("claude", cid, {"file_path": f"/{cid}.md"})
tmp_cache._data["claude"][cid]["exported_at"] = "2026-06-12T01:00:00+00:00"
remaining = tmp_cache.get_new_or_updated("claude", convs, force=True,
campaign_at=campaign)
assert {c["id"] for c in remaining} == {"c2", "c3", "c4"}
# Finish them → none left.
for cid in ("c2", "c3", "c4"):
tmp_cache.mark_exported("claude", cid, {"file_path": f"/{cid}.md"})
tmp_cache._data["claude"][cid]["exported_at"] = "2026-06-12T02:00:00+00:00"
assert tmp_cache.get_new_or_updated("claude", convs, force=True,
campaign_at=campaign) == []
def test_force_campaign_marker_roundtrip(self, tmp_cache):
assert tmp_cache.get_force_campaign() is None
tmp_cache.set_force_campaign("2026-06-12T00:00:00+00:00")
assert tmp_cache.get_force_campaign() == "2026-06-12T00:00:00+00:00"
# Survives reload, and is not mistaken for a provider by stats().
reopened = Cache(tmp_cache._dir)
assert reopened.get_force_campaign() == "2026-06-12T00:00:00+00:00"
assert "force_campaign_at" not in reopened.stats()
tmp_cache.clear_force_campaign()
assert tmp_cache.get_force_campaign() is None
class TestReexportPreservesJoplinLink:
"""A re-export must not drop the Joplin note link, or re-syncing would
create duplicate notes instead of updating the existing ones."""
def test_joplin_fields_carried_across_reexport(self, tmp_cache):
tmp_cache.mark_exported("chatgpt", "c1", {
"title": "T", "updated_at": "2024-01-01T00:00:00Z", "file_path": "/x.md",
})
tmp_cache.mark_joplin_synced("chatgpt", "c1", "note-123")
tmp_cache.set_joplin_resources("chatgpt", "c1", {"media/a.png": "res-1"})
# Re-export the same conversation (new render, new file_path).
tmp_cache.mark_exported("chatgpt", "c1", {
"title": "T", "updated_at": "2024-01-01T00:00:00Z", "file_path": "/x2.md",
})
entry = tmp_cache.get_all_entries("chatgpt")["c1"]
assert entry["joplin_note_id"] == "note-123"
assert entry["joplin_resources"] == {"media/a.png": "res-1"}
assert entry["file_path"] == "/x2.md"
def test_reexport_marks_note_for_update_not_create(self, tmp_cache):
tmp_cache.mark_exported("chatgpt", "c1", {
"updated_at": "2024-01-01T00:00:00Z", "file_path": "/x.md",
})
tmp_cache.mark_joplin_synced("chatgpt", "c1", "note-123")
# Nothing pending right after sync.
assert tmp_cache.get_joplin_pending("chatgpt") == []
# Re-export bumps exported_at past joplin_synced_at → pending for update,
# carrying the existing note id so the sync updates rather than creates.
tmp_cache.mark_exported("chatgpt", "c1", {
"updated_at": "2024-01-01T00:00:00Z", "file_path": "/x.md",
})
pending = tmp_cache.get_joplin_pending("chatgpt")
assert len(pending) == 1
assert pending[0][1]["joplin_note_id"] == "note-123"
+586
View File
@@ -0,0 +1,586 @@
"""Unit tests for the Claude Code session provider."""
import json
import pytest
from src.blocks import (
BLOCK_TYPE_COLLAPSED,
BLOCK_TYPE_TEXT,
BLOCK_TYPE_THINKING,
BLOCK_TYPE_TOOL_RESULT,
BLOCK_TYPE_TOOL_USE,
render_blocks_to_markdown,
)
from src.loss_report import LossReport
from src.providers.claude_code import ClaudeCodeProvider
def _write_session(tmp_path, records, project_dir="-home-jesse-myproj", name="abc-123"):
proj = tmp_path / project_dir
proj.mkdir(parents=True, exist_ok=True)
f = proj / f"{name}.jsonl"
# ensure_ascii=False so raw U+0085/U+2028 reach the parser — see
# TestExoticLineBreaks.
f.write_text(
"\n".join(json.dumps(r, ensure_ascii=False) for r in records), encoding="utf-8"
)
return f
def _records():
"""A representative session: title, harness noise, dialogue, tool traffic."""
return [
{"type": "ai-title", "aiTitle": "First title", "sessionId": "abc-123"},
{
"type": "user",
"cwd": "/home/jesse/myproj",
"timestamp": "2026-05-01T10:00:00.000Z",
"message": {
"role": "user",
"content": (
"<local-command-caveat>Caveat: ignore</local-command-caveat>"
"<command-name>/model</command-name>"
"<local-command-stdout>Set model</local-command-stdout>"
"Please review my project."
),
},
},
{
"type": "assistant",
"timestamp": "2026-05-01T10:00:05.000Z",
"message": {
"role": "assistant",
"content": [
{"type": "thinking", "thinking": "Let me think about this."},
{"type": "text", "text": "I'll review it now."},
{"type": "tool_use", "id": "t1", "name": "Read", "input": {"file_path": "/x"}},
{"type": "tool_use", "id": "t2", "name": "Read", "input": {"file_path": "/y"}},
{"type": "tool_use", "id": "t3", "name": "Bash", "input": {"command": "ls"}},
],
},
},
{
"type": "user",
"timestamp": "2026-05-01T10:00:08.000Z",
"message": {
"role": "user",
"content": [
{"type": "tool_result", "tool_use_id": "t1", "content": "file contents " * 50},
],
},
},
{
"type": "assistant",
"timestamp": "2026-05-01T10:00:20.000Z",
"message": {
"role": "assistant",
"content": [{"type": "text", "text": "Here is my review: all good."}],
},
},
# Harness records — skipped
{"type": "file-history-snapshot", "snapshot": {"big": "blob"}},
{"type": "permission-mode", "mode": "plan"},
# Subagent and meta records — skipped
{
"type": "assistant",
"isSidechain": True,
"message": {"role": "assistant", "content": [{"type": "text", "text": "subagent noise"}]},
},
{
"type": "user",
"isMeta": True,
"message": {"role": "user", "content": "meta noise"},
},
# Later title wins
{"type": "ai-title", "aiTitle": "Review my project", "sessionId": "abc-123"},
]
class TestClaudeCodeProvider:
def _provider(self, tmp_path, policy="placeholder"):
return ClaudeCodeProvider(projects_dir=tmp_path, hidden_content=policy)
def test_listing_metadata(self, tmp_path):
_write_session(tmp_path, _records())
p = self._provider(tmp_path)
convs = p.fetch_all_conversations()
assert len(convs) == 1
c = convs[0]
assert c["id"] == "abc-123"
assert c["title"] == "Review my project" # last ai-title wins
assert c["_project_name"] == "myproj" # from cwd basename
assert c["created_at"] == "2026-05-01T10:00:00.000Z"
assert c["updated_at"] # file mtime
def test_empty_files_skipped(self, tmp_path):
proj = tmp_path / "-home-x"
proj.mkdir()
(proj / "empty.jsonl").write_text("")
p = self._provider(tmp_path)
assert p.fetch_all_conversations() == []
def test_normalize_prose_only_default(self, tmp_path):
_write_session(tmp_path, _records())
p = self._provider(tmp_path)
p.fetch_all_conversations()
raw = p.get_conversation("abc-123")
result = p.normalize_conversation(raw)
assert result["provider"] == "claude-code"
assert result["title"] == "Review my project"
assert result["project"] == "myproj"
roles = [m["role"] for m in result["messages"]]
assert roles == ["user", "assistant", "assistant"]
# Harness tags stripped, dialogue kept
user_text = result["messages"][0]["blocks"][0]["text"]
assert user_text == "Please review my project."
assert "Caveat" not in user_text
# Tool activity collapsed into one grouped placeholder on the
# assistant message that ran the tools (includes the tool_result
# bytes from the following user record).
first_assistant = result["messages"][1]
types = [b["type"] for b in first_assistant["blocks"]]
assert types == [BLOCK_TYPE_TEXT, BLOCK_TYPE_COLLAPSED]
collapsed = first_assistant["blocks"][1]
assert "3 calls" in collapsed["origin"]
assert "Read ×2" in collapsed["origin"]
assert "Bash ×1" in collapsed["origin"]
assert collapsed["size_bytes"] > 500
# Thinking dropped without placeholder; sidechain/meta absent
all_blocks = [b for m in result["messages"] for b in m["blocks"]]
assert not any(b["type"] == BLOCK_TYPE_THINKING for b in all_blocks)
assert not any(
"subagent noise" in (b.get("text") or "") or "meta noise" in (b.get("text") or "")
for b in all_blocks
)
def test_collapsed_counted_in_loss_report(self, tmp_path):
_write_session(tmp_path, _records())
p = self._provider(tmp_path)
p.fetch_all_conversations()
report = LossReport()
p.normalize_conversation(p.get_conversation("abc-123"), report)
assert report.collapsed["Read"] == 2
assert report.collapsed["Bash"] == 1
assert report.collapsed["tool_result"] == 1
assert report.collapsed["thinking"] == 1
assert report.collapsed_bytes > 500
def test_full_policy_keeps_everything(self, tmp_path):
_write_session(tmp_path, _records())
p = self._provider(tmp_path, policy="full")
p.fetch_all_conversations()
result = p.normalize_conversation(p.get_conversation("abc-123"))
all_blocks = [b for m in result["messages"] for b in m["blocks"]]
types = {b["type"] for b in all_blocks}
assert BLOCK_TYPE_THINKING in types
assert BLOCK_TYPE_TOOL_USE in types
assert BLOCK_TYPE_TOOL_RESULT in types
assert BLOCK_TYPE_COLLAPSED not in types
def test_updated_at_consistent_with_listing(self, tmp_path):
"""Listing and normalized updated_at must match or the cache would
consider every session stale on every run (mtime vs last timestamp)."""
_write_session(tmp_path, _records())
p = self._provider(tmp_path)
listing = p.fetch_all_conversations()[0]
normalized = p.normalize_conversation(p.get_conversation("abc-123"))
assert normalized["updated_at"] == listing["updated_at"]
def test_title_falls_back_to_first_prompt(self, tmp_path):
records = [r for r in _records() if r.get("type") != "ai-title"]
_write_session(tmp_path, records)
p = self._provider(tmp_path)
p.fetch_all_conversations()
result = p.normalize_conversation(p.get_conversation("abc-123"))
assert result["title"] == "Please review my project."
def test_unknown_block_type_surfaces(self, tmp_path):
records = [
{
"type": "assistant",
"timestamp": "2026-05-01T10:00:00.000Z",
"cwd": "/home/jesse/myproj",
"message": {
"role": "assistant",
"content": [
{"type": "text", "text": "hello"},
{"type": "future_block_xyz", "data": 1},
],
},
},
]
_write_session(tmp_path, records)
p = self._provider(tmp_path)
p.fetch_all_conversations()
report = LossReport()
result = p.normalize_conversation(p.get_conversation("abc-123"), report)
assert report.unknown_blocks["claude-code.future_block_xyz"] == 1
rendered = render_blocks_to_markdown(result["messages"][0]["blocks"])
assert "Unsupported content" in rendered
def test_collapsed_placeholder_renders_grouped_line(self, tmp_path):
_write_session(tmp_path, _records())
p = self._provider(tmp_path)
p.fetch_all_conversations()
result = p.normalize_conversation(p.get_conversation("abc-123"))
rendered = render_blocks_to_markdown(result["messages"][1]["blocks"])
assert "> 🔧 **Tool output** — `3 calls: Read ×2, Bash ×1`" in rendered
assert "omitted (EXPORTER_HIDDEN_CONTENT=full to keep)" in rendered
# ---------------------------------------------------------------------------
# Subagent folding (Part B)
# ---------------------------------------------------------------------------
def _write_subagent(
session_file, tool_use_id, records, agent_type="Explore",
description="do a thing", name="agent-x",
):
"""Write a <session>/subagents/<name>.jsonl + .meta.json sidecar."""
subdir = session_file.parent / session_file.stem / "subagents"
subdir.mkdir(parents=True, exist_ok=True)
(subdir / f"{name}.jsonl").write_text(
"\n".join(json.dumps(r) for r in records), encoding="utf-8"
)
(subdir / f"{name}.meta.json").write_text(
json.dumps({
"agentType": agent_type,
"description": description,
"toolUseId": tool_use_id,
}),
encoding="utf-8",
)
def _parent_with_task(tool_use_id="toolu_sub1"):
return [
{"type": "ai-title", "aiTitle": "Parent session"},
{
"type": "user", "cwd": "/home/jesse/myproj",
"timestamp": "2026-05-01T10:00:00.000Z",
"message": {"role": "user", "content": "delegate the research"},
},
{
"type": "assistant", "timestamp": "2026-05-01T10:00:05.000Z",
"message": {"role": "assistant", "content": [
{"type": "text", "text": "I'll delegate this."},
{"type": "tool_use", "id": tool_use_id, "name": "Task",
"input": {"description": "research"}},
]},
},
]
def _subagent_records():
return [
{
"type": "user", "isSidechain": True,
"timestamp": "2026-05-01T10:00:06.000Z",
"message": {"role": "user", "content": "Go research X"},
},
{
"type": "assistant", "isSidechain": True,
"timestamp": "2026-05-01T10:00:07.000Z",
"message": {"role": "assistant", "content": [
{"type": "text", "text": "Here is the research result."},
{"type": "tool_use", "id": "st1", "name": "Grep",
"input": {"pattern": "x"}},
]},
},
]
class TestSubagentFold:
def _provider(self, tmp_path, policy="placeholder"):
return ClaudeCodeProvider(projects_dir=tmp_path, hidden_content=policy)
def _normalize(self, tmp_path, policy="placeholder"):
f = _write_session(tmp_path, _parent_with_task(), name="parent-1")
_write_subagent(
f, "toolu_sub1", _subagent_records(),
agent_type="Explore", description="research X",
)
p = self._provider(tmp_path, policy)
p.fetch_all_conversations()
return p.normalize_conversation(p.get_conversation("parent-1"))
def test_subagent_folded_into_parent(self, tmp_path):
conv = self._normalize(tmp_path)
blocks = [b for m in conv["messages"] for b in m["blocks"]]
subs = [b for b in blocks if b["type"] == "subagent"]
assert len(subs) == 1
assert subs[0]["agent_type"] == "Explore"
assert subs[0]["description"] == "research X"
inner_text = [
bb.get("text") for m in subs[0]["messages"] for bb in m["blocks"]
]
assert "Go research X" in inner_text
assert "Here is the research result." in inner_text
def test_subagent_tools_collapsed(self, tmp_path):
conv = self._normalize(tmp_path)
sub = next(
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == "subagent"
)
inner_blocks = [bb for m in sub["messages"] for bb in m["blocks"]]
assert any(b["type"] == BLOCK_TYPE_COLLAPSED for b in inner_blocks)
assert not any(b["type"] == BLOCK_TYPE_TOOL_USE for b in inner_blocks)
collapsed = next(b for b in inner_blocks if b["type"] == BLOCK_TYPE_COLLAPSED)
assert "Grep" in collapsed["origin"]
def test_subagent_not_listed_as_conversation(self, tmp_path):
f = _write_session(tmp_path, _parent_with_task(), name="parent-1")
_write_subagent(f, "toolu_sub1", _subagent_records())
p = self._provider(tmp_path)
convs = p.fetch_all_conversations()
assert [c["id"] for c in convs] == ["parent-1"]
def test_renders_as_details_block(self, tmp_path):
conv = self._normalize(tmp_path)
spawning = conv["messages"][1] # the assistant turn with the Task call
rendered = render_blocks_to_markdown(spawning["blocks"])
assert "<details>" in rendered
assert "<summary>🤖 Subagent: Explore — research X</summary>" in rendered
assert "</details>" in rendered
assert "Here is the research result." in rendered
def test_main_pass_still_skips_stray_sidechain(self, tmp_path):
# A sidechain record with no matching subagent file must NOT leak into
# the top-level dialogue.
records = _parent_with_task() + [
{"type": "assistant", "isSidechain": True,
"message": {"role": "assistant",
"content": [{"type": "text", "text": "stray sidechain"}]}},
]
f = _write_session(tmp_path, records, name="parent-2")
_write_subagent(f, "toolu_sub1", _subagent_records())
p = self._provider(tmp_path)
p.fetch_all_conversations()
conv = p.normalize_conversation(p.get_conversation("parent-2"))
top_text = [
b.get("text") for m in conv["messages"] for b in m["blocks"]
]
assert "stray sidechain" not in top_text
def test_fork_containing_its_own_spawn_call(self, tmp_path):
# The shape Claude Code writes for a `fork` subagent: a context-ref
# record, a copy of the parent turn holding the fork's own spawn call,
# then the directive behind <fork-boilerplate>. Folding that copied
# call recursed until RecursionError.
fork_records = [
{"type": "fork-context-ref", "agentId": "a1", "parentLastUuid": "u0"},
{
"type": "assistant", "isSidechain": True,
"message": {"role": "assistant", "content": [
{"type": "tool_use", "id": "toolu_sub1", "name": "Agent",
"input": {"description": "research"}},
]},
},
{
"type": "user", "isSidechain": True,
"timestamp": "2026-05-01T10:00:06.000Z",
"message": {"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_sub1",
"content": [{"type": "text", "text": "Fork started"}]},
{"type": "text", "text":
"<fork-boilerplate>\nYou are a worker fork.\n"
"</fork-boilerplate>\n\nYour directive: research X"},
]},
},
{
"type": "assistant", "isSidechain": True,
"timestamp": "2026-05-01T10:00:07.000Z",
"message": {"role": "assistant", "content": [
{"type": "text", "text": "Fork result."},
]},
},
]
f = _write_session(tmp_path, _parent_with_task(), name="parent-3")
_write_subagent(
f, "toolu_sub1", fork_records,
agent_type="fork", description="research X",
)
p = self._provider(tmp_path)
p.fetch_all_conversations()
conv = p.normalize_conversation(p.get_conversation("parent-3"))
subs = [
b for m in conv["messages"] for b in m["blocks"] if b["type"] == "subagent"
]
assert len(subs) == 1
inner = [bb for m in subs[0]["messages"] for bb in m["blocks"]]
assert not any(b["type"] == "subagent" for b in inner)
inner_text = [b.get("text") for b in inner]
assert "Your directive: research X" in inner_text
assert "Fork result." in inner_text
assert not any("worker fork" in (t or "") for t in inner_text)
# ---------------------------------------------------------------------------
# Repo tags in title (Part C)
# ---------------------------------------------------------------------------
class TestRepoTags:
@staticmethod
def _mkrepo(base, name):
"""Create a git repo dir (bare .git marker) and return its path."""
(base / name / ".git").mkdir(parents=True, exist_ok=True)
return base / name
def _title(self, tmp_path, content):
(tmp_path / "ws").mkdir(exist_ok=True)
records = [
{"type": "ai-title", "aiTitle": "Work session"},
{"type": "user", "cwd": str(tmp_path / "ws"),
"timestamp": "2026-05-01T10:00:00.000Z",
"message": {"role": "user", "content": "go"}},
{"type": "assistant",
"message": {"role": "assistant", "content": content}},
]
_write_session(tmp_path, records, name="ws-1")
p = ClaudeCodeProvider(projects_dir=tmp_path)
p.fetch_all_conversations()
return p.normalize_conversation(p.get_conversation("ws-1"))["title"]
@staticmethod
def _read(path):
return {"type": "tool_use", "id": str(path), "name": "Read",
"input": {"file_path": str(path)}}
def test_repos_by_git_root_ordered_by_frequency(self, tmp_path):
a = self._mkrepo(tmp_path, "repo-a")
b = self._mkrepo(tmp_path, "repo-b")
content = [
self._read(a / "src/x.py"), self._read(a / "src/y.py"),
self._read(b / "z.py"),
# a file not inside any git repo → excluded
self._read(tmp_path / "CLAUDE.md"),
]
assert self._title(tmp_path, content) == "Work session [repo-a, repo-b]"
def test_cross_workspace_repo_tagged(self, tmp_path):
# Launched in ws/, but touched a repo in a sibling tree — still tagged.
other = self._mkrepo(tmp_path / "elsewhere", "faraway-repo")
content = [self._read(other / "deep/nested/file.py")]
assert self._title(tmp_path, content) == "Work session [faraway-repo]"
def test_no_repos_no_bracket(self, tmp_path):
content = [self._read(tmp_path / "loose/file.txt")] # no .git anywhere
assert self._title(tmp_path, content) == "Work session"
def test_cap_three_with_ellipsis(self, tmp_path):
content = [
self._read(self._mkrepo(tmp_path, f"repo-{c}") / "f/x.py")
for c in "abcd"
]
title = self._title(tmp_path, content)
assert title.startswith("Work session [repo-a, repo-b, repo-c, …]")
assert "repo-d" not in title
def test_paths_inside_bash_commands_count(self, tmp_path):
x = self._mkrepo(tmp_path, "repo-x")
content = [
{"type": "tool_use", "id": "a", "name": "Bash",
"input": {"command": f"cat {x / 'main.go'}"}},
]
assert self._title(tmp_path, content) == "Work session [repo-x]"
def test_ignore_list_via_env(self, tmp_path, monkeypatch):
a = self._mkrepo(tmp_path, "repo-a")
b = self._mkrepo(tmp_path, "repo-b")
monkeypatch.setenv("CLAUDE_CODE_REPO_TAG_IGNORE", "repo-b, notes")
content = [self._read(a / "x.py"), self._read(b / "y.py")]
assert self._title(tmp_path, content) == "Work session [repo-a]"
# ---------------------------------------------------------------------------
# Multiple projects roots (Part D)
# ---------------------------------------------------------------------------
class TestMultiRoot:
def _titled(self, aititle, cwd="/home/jesse/p"):
return [
{"type": "ai-title", "aiTitle": aititle},
{"type": "user", "cwd": cwd, "timestamp": "2026-05-01T10:00:00.000Z",
"message": {"role": "user", "content": "hi"}},
]
def test_two_roots_merged(self, tmp_path):
r1, r2 = tmp_path / "root1", tmp_path / "root2"
_write_session(r1, self._titled("From root1"), name="s1")
_write_session(r2, self._titled("From root2"), name="s2")
p = ClaudeCodeProvider(projects_dir=[r1, r2])
ids = sorted(c["id"] for c in p.fetch_all_conversations())
assert ids == ["s1", "s2"]
def test_duplicate_uuid_newer_wins(self, tmp_path):
import os
import time
r1, r2 = tmp_path / "root1", tmp_path / "root2"
_write_session(r1, self._titled("Old copy"), name="dup")
f2 = _write_session(r2, self._titled("New copy"), name="dup")
os.utime(f2, (time.time() + 10, time.time() + 10))
p = ClaudeCodeProvider(projects_dir=[r1, r2])
convs = p.fetch_all_conversations()
assert len(convs) == 1
assert convs[0]["title"] == "New copy"
assert p._path_map["dup"] == f2
def test_single_path_backward_compatible(self, tmp_path):
_write_session(tmp_path, _records())
p = ClaudeCodeProvider(projects_dir=tmp_path)
assert len(p.fetch_all_conversations()) == 1
def test_env_pathsep_list(self, tmp_path, monkeypatch):
import os
from src.providers.claude_code import resolve_roots
r1, r2 = tmp_path / "root1", tmp_path / "root2"
_write_session(r1, self._titled("From root1"), name="s1")
_write_session(r2, self._titled("From root2"), name="s2")
monkeypatch.setenv("CLAUDE_CODE_DIR", f"{r1}{os.pathsep}{r2}")
assert r1 in resolve_roots() and r2 in resolve_roots()
p = ClaudeCodeProvider()
assert len(p.fetch_all_conversations()) == 2
def test_config_dir_projects_included(self, tmp_path, monkeypatch):
from src.providers.claude_code import resolve_roots
monkeypatch.delenv("CLAUDE_CODE_DIR", raising=False)
monkeypatch.setenv("CLAUDE_CONFIG_DIR", str(tmp_path / "cfg"))
roots = resolve_roots()
assert (tmp_path / "cfg" / "projects") in roots
class TestExoticLineBreaks:
"""Same U+0085 hazard as the Codex provider — see its TestExoticLineBreaks."""
# C0 controls (\x0b, \x0c) are excluded: JSON requires them escaped, so they
# never reach the splitter literally. These three do.
@pytest.mark.parametrize("sep", ["\x85", "
", "
"])
def test_record_with_exotic_break_survives(self, tmp_path, sep):
records = [
{
"type": "user",
"cwd": "/home/jesse/myproj",
"timestamp": "2026-05-01T10:00:00.000Z",
"message": {"role": "user", "content": f"before{sep}after"},
},
]
_write_session(tmp_path, records)
prov = ClaudeCodeProvider(projects_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation("abc-123"))
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert f"before{sep}after" in texts
+534
View File
@@ -1,5 +1,7 @@
"""CLI-level tests using Click's CliRunner — no live API calls required."""
from pathlib import Path
import pytest
from click.testing import CliRunner
@@ -124,6 +126,538 @@ class TestExportSinceValidation:
"CHATGPT_SESSION_TOKEN": "eyJtesttoken",
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
# This test hits the real (failing) auth endpoint with retries;
# don't add politeness pacing on top of the backoff sleeps.
"REQUEST_DELAY": "0",
},
)
assert "Invalid --since date" not in result.output
def test_max_conversations_zero_rejected(self, tmp_path):
"""--max-conversations uses IntRange(min=1); 0 must be rejected by click."""
self._pre_populated_cache(tmp_path)
runner = CliRunner(mix_stderr=True)
result = runner.invoke(
cli,
["--no-log-file", "export", "--max-conversations", "0"],
env={
"CHATGPT_SESSION_TOKEN": "eyJtesttoken",
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
},
)
assert result.exit_code == 2
assert "max-conversations" in result.output
# ---------------------------------------------------------------------------
# LossReport summary
# ---------------------------------------------------------------------------
class TestLossReportSummary:
"""The LossReport's format_summary() pinned format covers zero, top-5, and overflow cases."""
def test_zero_summary_uses_none_sentinel(self):
from src.loss_report import LossReport
report = LossReport()
out = report.format_summary()
assert "[export] Run summary:" in out
assert "conversations: 0" in out
assert "messages rendered: 0" in out
# All three "(none)" sentinels present — never empty parens
# (unknown blocks, extraction failures, collapsed by policy)
assert out.count("(none)") == 3
def test_top_5_breakdown(self):
from src.loss_report import LossReport
report = LossReport()
for raw_type in ("a", "b", "c", "d", "e", "f", "g"):
report.record_unknown(raw_type)
if raw_type == "a":
# Make 'a' the most common
for _ in range(4):
report.record_unknown("a")
out = report.format_summary()
# Top entry shown
assert "a=5" in out
# Overflow line present (7 types, top 5 + 2 more)
assert "+ 2 more types" in out
def test_messages_and_conversations_recorded(self):
from src.loss_report import LossReport
report = LossReport()
report.record_conversation()
report.record_message()
report.record_message()
out = report.format_summary()
assert "conversations: 1" in out
assert "messages rendered: 2" in out
# ---------------------------------------------------------------------------
# prune command
# ---------------------------------------------------------------------------
class TestPrune:
"""prune deletes export files not referenced by the manifest."""
def _setup(self, tmp_path):
"""Cache with one referenced file; one stale file + empty-dir candidate."""
cache = Cache(tmp_path / "cache")
cache.acknowledge_tos()
export_dir = tmp_path / "exports"
keep = export_dir / "chatgpt" / "proj.2026" / "keep.md"
keep.parent.mkdir(parents=True)
keep.write_text("kept")
stale = export_dir / "chatgpt" / "proj" / "2026" / "old-layout.md"
stale.parent.mkdir(parents=True)
stale.write_text("stale")
cache.mark_exported("chatgpt", "conv-1", {"file_path": str(keep)})
return export_dir, keep, stale
def _invoke(self, tmp_path, *args):
runner = CliRunner(mix_stderr=True)
return runner.invoke(
cli,
["--no-log-file", "prune", *args],
env={
"CACHE_DIR": str(tmp_path / "cache"),
"EXPORT_DIR": str(tmp_path / "exports"),
},
)
def test_dry_run_lists_but_keeps_files(self, tmp_path):
export_dir, keep, stale = self._setup(tmp_path)
result = self._invoke(tmp_path, "--dry-run")
assert result.exit_code == 0
assert "old-layout.md" in result.output
assert "Dry run" in result.output
assert stale.exists() and keep.exists()
def test_yes_deletes_stale_keeps_referenced_sweeps_dirs(self, tmp_path):
export_dir, keep, stale = self._setup(tmp_path)
result = self._invoke(tmp_path, "--yes")
assert result.exit_code == 0
assert not stale.exists()
assert keep.exists()
# Old-layout dirs are now empty and swept
assert not (export_dir / "chatgpt" / "proj").exists()
def test_refuses_with_empty_manifest(self, tmp_path):
"""Footgun guard: after cache --clear, prune must not wipe the archive."""
cache = Cache(tmp_path / "cache")
cache.acknowledge_tos()
export_dir = tmp_path / "exports"
f = export_dir / "chatgpt" / "a.md"
f.parent.mkdir(parents=True)
f.write_text("data")
result = self._invoke(tmp_path, "--yes")
assert result.exit_code == 1
assert "Refusing to prune" in result.output
assert f.exists()
def test_aborts_without_confirmation(self, tmp_path):
export_dir, keep, stale = self._setup(tmp_path)
runner = CliRunner(mix_stderr=True)
result = runner.invoke(
cli,
["--no-log-file", "prune"],
input="n\n",
env={
"CACHE_DIR": str(tmp_path / "cache"),
"EXPORT_DIR": str(tmp_path / "exports"),
},
)
assert result.exit_code == 0
assert "Aborted" in result.output
assert stale.exists()
class TestCanaryCommand:
"""`canary` wiring — no live API calls (no tokens configured)."""
def test_no_tokens_exits_nonzero(self, tmp_path):
cache = Cache(tmp_path / "cache")
cache.acknowledge_tos()
runner = CliRunner(mix_stderr=True)
result = runner.invoke(
cli,
["--no-log-file", "canary"],
env={
"CACHE_DIR": str(tmp_path / "cache"),
"EXPORT_DIR": str(tmp_path / "exports"),
# Empty so .env (override=False) can't repopulate real tokens.
"CHATGPT_SESSION_TOKEN": "",
"CLAUDE_SESSION_KEY": "",
},
)
assert result.exit_code == 1
assert "No web-API provider tokens" in result.output
class TestProjectsCommand:
"""`projects` discovers project IDs missing from CHATGPT_PROJECT_IDS."""
def _patch_provider(self, monkeypatch, summaries, details=None, names=None):
import src.providers.chatgpt as chatgpt_mod
names = names or {}
details = details or {}
class FakeProvider:
def __init__(self, **kwargs):
self._project_ids = kwargs.get("project_ids") or []
def fetch_all_conversations(self, since=None):
return summaries
def get_conversation(self, conv_id):
return details.get(conv_id, {})
def _fetch_project_name(self, gizmo_id):
return names.get(gizmo_id, gizmo_id)
monkeypatch.setattr(chatgpt_mod, "ChatGPTProvider", FakeProvider)
return FakeProvider
def _env(self, tmp_path, **extra):
# Clears the ToS gate and the first-run doctor check.
cache = Cache(tmp_path)
cache.acknowledge_tos()
cache.mark_exported("chatgpt", "dummy", {"updated_at": "2024-01-01T00:00:00Z"})
env = {
"CHATGPT_SESSION_TOKEN": "eyJtesttoken",
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
}
env.update(extra)
return env
def test_reports_project_absent_from_config(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-p-missing"}],
names={"g-p-missing": "Tech Questions"},
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects"], env=self._env(tmp_path)
)
assert result.exit_code == 0
assert "Tech Questions" in result.output
assert "g-p-missing" in result.output
assert "CHATGPT_PROJECT_IDS=g-p-missing" in result.output
def test_quiet_when_everything_is_configured(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-p-known"}],
names={"g-p-known": "Known"},
)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "projects"],
env=self._env(tmp_path, CHATGPT_PROJECT_IDS="g-p-known"),
)
assert result.exit_code == 0
assert "already configured" in result.output
def test_custom_gpt_ids_are_ignored(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-notaproject"}],
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects"], env=self._env(tmp_path)
)
assert "g-notaproject" not in result.output
def test_deep_reads_conversation_details(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1"}], # listing carries no gizmo_id
details={"c1": {"gizmo_id": "g-p-fromdetail"}},
names={"g-p-fromdetail": "Found Deep"},
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects", "--deep"], env=self._env(tmp_path)
)
assert "Found Deep" in result.output
assert "CHATGPT_PROJECT_IDS=g-p-fromdetail" in result.output
def test_without_deep_says_so_rather_than_reporting_nothing(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1"}],
details={"c1": {"gizmo_id": "g-p-fromdetail"}},
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects"], env=self._env(tmp_path)
)
assert "--deep" in result.output
def test_write_updates_env(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-p-new"}],
names={"g-p-new": "New Project"},
)
runner = CliRunner(mix_stderr=True)
with runner.isolated_filesystem(temp_dir=tmp_path) as fs:
result = runner.invoke(
cli, ["--no-log-file", "projects", "--write"], env=self._env(tmp_path)
)
assert result.exit_code == 0
env_text = (Path(fs) / ".env").read_text(encoding="utf-8")
assert "CHATGPT_PROJECT_IDS=g-p-new" in env_text
# ---------------------------------------------------------------------------
# sync command + non-interactive ToS gate
# ---------------------------------------------------------------------------
class TestSyncCommand:
"""`sync` chains export → joplin for schedulers, with a real exit code."""
def _cache(self, tmp_path) -> Cache:
cache = Cache(tmp_path)
cache.acknowledge_tos()
# Non-empty last_run so the first-run doctor gate stays out of the way.
cache.mark_exported("codex", "dummy", {"updated_at": "2024-01-01T00:00:00Z"})
return cache
def _env(self, tmp_path) -> dict:
"""A real (minimal) codex session — `export` exits 1 on no providers at
all, so an empty directory would test the wrong failure."""
import json
day = tmp_path / "sessions" / "2026" / "08" / "17"
day.mkdir(parents=True, exist_ok=True)
sid = "01a00e3f-a309-74a3-bf32-06c2cd87faa3"
records = [
{
"timestamp": "2026-08-17T05:44:06.666Z",
"type": "session_meta",
"payload": {"session_id": sid, "timestamp": "2026-08-17T05:44:06.666Z",
"cwd": str(tmp_path / "ws")},
},
{
"timestamp": "2026-08-17T05:44:09.000Z",
"type": "event_msg",
"payload": {"type": "item_completed", "item": {
"type": "UserMessage", "id": "u1",
"content": [{"type": "text", "text": "hello"}]}},
},
]
(day / f"rollout-2026-08-17T01-44-06-{sid}.jsonl").write_text(
"\n".join(json.dumps(r) for r in records), encoding="utf-8"
)
return {
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
"CODEX_DIR": str(tmp_path / "sessions"),
}
def test_skip_joplin_exits_zero(self, tmp_path):
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--skip-joplin"],
env=self._env(tmp_path),
)
assert result.exit_code == 0
assert "Skipping Joplin sync" in result.output
assert "Sync complete" in result.output
def test_joplin_optional_survives_unreachable_joplin(self, tmp_path):
"""Joplin being closed must not fail a scheduled run — the export is done."""
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--joplin-optional"],
env={**self._env(tmp_path), "JOPLIN_API_URL": "http://127.0.0.1:9"},
)
assert result.exit_code == 0
assert "Joplin sync skipped" in result.output
def test_unreachable_joplin_fails_without_the_flag(self, tmp_path):
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex"],
env={**self._env(tmp_path), "JOPLIN_API_URL": "http://127.0.0.1:9"},
)
assert result.exit_code == 1
def test_export_failures_set_nonzero_exit(self, tmp_path, monkeypatch):
"""A scheduler must be able to tell a real run from a silent no-op."""
self._cache(tmp_path)
import src.main as main_mod
real_export = main_mod.export.callback
def fake_export(*args, **kwargs):
import click
ctx = click.get_current_context()
ctx.obj["last_export_summary"] = {
"codex": {"exported": 0, "skipped": 0, "failed": 3}
}
monkeypatch.setattr(main_mod.export, "callback", fake_export)
try:
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--skip-joplin"],
env=self._env(tmp_path),
)
finally:
monkeypatch.setattr(main_mod.export, "callback", real_export)
assert result.exit_code == 1
assert "3 conversation(s) failed to export" in result.output
class TestNonInteractiveTosGate:
"""Without a TTY the gate must fail loudly, not exit 0 having done nothing."""
def test_no_tty_exits_one_with_explanation(self, tmp_path, monkeypatch):
Cache(tmp_path) # fresh cache: ToS not acknowledged
monkeypatch.setattr("sys.stdin.isatty", lambda: False)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "doctor"],
env={"CACHE_DIR": str(tmp_path), "EXPORT_DIR": str(tmp_path / "exports")},
)
assert result.exit_code == 1
assert "no terminal to" in result.output
# ---------------------------------------------------------------------------
# ntfy notifications
# ---------------------------------------------------------------------------
class TestNotifyFormatting:
"""The payload must stay counts-only: a public ntfy topic is world-readable."""
def test_success_summary(self):
from src.notify import format_summary
title, body, tags, priority = format_summary(
{"codex": {"exported": 2, "skipped": 5, "failed": 0}}, None, []
)
assert title.startswith("AI archive OK")
assert "codex: 2 exported, 5 up to date" in body
assert tags == "white_check_mark"
assert priority == "default"
def test_quiet_run_is_low_priority(self):
from src.notify import format_summary
_, _, tags, priority = format_summary(
{"codex": {"exported": 0, "skipped": 7, "failed": 0}}, None, []
)
assert tags == "zzz"
assert priority == "low"
def test_failure_summary_is_high_priority(self):
from src.notify import format_summary
title, body, tags, priority = format_summary(
{"chatgpt": {"exported": 0, "skipped": 0, "failed": 12}},
None,
["chatgpt: 12 conversation(s) failed to export"],
)
assert "FAILED" in title
assert "12 FAILED" in body
assert tags == "rotating_light"
assert priority == "high"
def test_hostname_present(self):
"""Two machines share one topic — counts are meaningless without it."""
from src.notify import format_summary, machine_name
title, _, _, _ = format_summary({}, None, [])
assert machine_name() in title
class TestNotifyHeaderEncoding:
"""HTTP headers are latin-1; an em dash in a title loses the notification."""
def test_smart_punctuation_flattened(self):
from src.notify import _ascii_header
out = _ascii_header("AI archive — don’t “fail”…")
assert out == 'AI archive - don\'t "fail"...'
out.encode("ascii") # must not raise
def test_arbitrary_unicode_survives_as_ascii(self):
from src.notify import _ascii_header
_ascii_header("héllo — 世界").encode("ascii")
class TestNotifySend:
def test_no_topic_is_a_no_op(self, monkeypatch):
from src import notify as notify_mod
monkeypatch.delenv("NTFY_TOPIC", raising=False)
assert notify_mod.send("t", "m") is False
assert notify_mod.is_configured() is False
def test_network_failure_never_raises(self, monkeypatch):
"""A down ntfy server must not fail a run that captured data."""
from src import notify as notify_mod
monkeypatch.setenv("NTFY_TOPIC", "unit-test-topic")
monkeypatch.setenv("NTFY_SERVER", "http://127.0.0.1:9")
assert notify_mod.send("t", "m") is False
def test_off_policy_disables(self, monkeypatch):
from src import notify as notify_mod
monkeypatch.setenv("NTFY_TOPIC", "unit-test-topic")
monkeypatch.setenv("NTFY_NOTIFY", "off")
assert notify_mod.is_configured() is False
class TestNoRealNotificationsDuringTests:
"""Regression guard for the 2026-08-18 incident: the suite pushed to the
developer's real ntfy topic because `sync` loads `.env` via `load_dotenv`.
Asserts the conftest neutralisation holds even though a real `.env` with a
live NTFY_TOPIC sits beside the tests.
"""
def test_notifications_are_disabled(self):
from src import notify as notify_mod
assert notify_mod.is_configured() is False
def test_dotenv_cannot_reintroduce_a_topic(self):
"""`load_dotenv(override=False)` skips keys already present — including
empty ones. Deleting the var instead of emptying it would reopen this."""
import os
from dotenv import load_dotenv
from src import notify as notify_mod
load_dotenv(override=False)
assert os.getenv("NTFY_TOPIC", "").strip() == ""
assert notify_mod.is_configured() is False
def test_send_cannot_reach_the_public_server(self, monkeypatch):
"""Even a test that sets its own topic is pinned to a dead local port."""
import os
from src import notify as notify_mod
monkeypatch.setenv("NTFY_TOPIC", "some-topic")
assert "127.0.0.1" in os.getenv("NTFY_SERVER", "")
assert notify_mod.send("t", "m") is False
+410
View File
@@ -0,0 +1,410 @@
"""Unit tests for the Codex CLI session provider.
Fixtures mirror the real 0.147.0 rollout shape observed on 2026-08-18: dialogue
carried twice (typed ``item_completed`` items plus raw ``response_item``s),
code-mode ``exec`` calls whose input is JavaScript, encrypted reasoning, and
harness-injected user messages that exist only in the raw layer.
"""
import json
import pytest
from src.blocks import (
BLOCK_TYPE_COLLAPSED,
BLOCK_TYPE_TEXT,
BLOCK_TYPE_TOOL_RESULT,
BLOCK_TYPE_TOOL_USE,
COLLAPSED_KIND_HIDDEN_CONTEXT,
)
from src.loss_report import LossReport
from src.providers.codex import CodexProvider, resolve_roots
SESSION_ID = "01a00e3f-a309-74a3-bf32-06c2cd87faa3"
FILENAME = f"rollout-2026-08-17T01-44-06-{SESSION_ID}.jsonl"
def _write_session(tmp_path, records, name=FILENAME, day="2026/08/17"):
day_dir = tmp_path / day
day_dir.mkdir(parents=True, exist_ok=True)
f = day_dir / name
# ensure_ascii=False: Codex (Rust) writes non-ASCII literally, so real
# rollouts contain raw U+0085/U+2028 inside JSON strings. Escaping them here
# would hide exactly the hazard TestExoticLineBreaks exists to catch.
f.write_text(
"\n".join(json.dumps(r, ensure_ascii=False) for r in records), encoding="utf-8"
)
return f
def _item(item_type, ts="2026-08-17T05:44:10.000Z", **fields):
# `item_type`, not `kind` — Extension items carry their own `kind` field.
return {
"timestamp": ts,
"type": "event_msg",
"payload": {"type": "item_completed", "item": {"type": item_type, **fields}},
}
def _exec_call(cmd, fn="exec_command"):
"""A raw code-mode custom_tool_call — its input is JavaScript, not JSON."""
return {
"timestamp": "2026-08-17T05:44:11.000Z",
"type": "response_item",
"payload": {
"type": "custom_tool_call",
"name": "exec",
"call_id": "call_1",
"input": (
f'const r = await tools.{fn}({{"cmd":{json.dumps(cmd)},'
f'"workdir":"/home/jesse/ws","yield_time_ms":30000}});\ntext(r.output);'
),
},
}
def _records(cwd="/home/jesse/ws"):
"""A representative session: meta, harness noise, dialogue, tool traffic."""
return [
{
"timestamp": "2026-08-17T05:49:13.629Z",
"ordinal": 0,
"type": "session_meta",
"payload": {
"session_id": SESSION_ID,
"timestamp": "2026-08-17T05:44:06.666Z",
"cwd": cwd,
"cli_version": "0.147.0",
"source": "cli",
},
},
# Harness plumbing: present in the raw layer only, exactly as 0.147.0
# writes it. Must not become a message.
{
"timestamp": "2026-08-17T05:44:07.000Z",
"type": "response_item",
"payload": {
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "# AGENTS.md instructions for /x"}],
},
},
{
"timestamp": "2026-08-17T05:44:08.000Z",
"type": "response_item",
"payload": {
"type": "message",
"role": "developer",
"content": [{"type": "input_text", "text": "<skills_instructions>…"}],
},
},
_item(
"UserMessage",
ts="2026-08-17T05:44:09.000Z",
id="u1",
content=[{"type": "text", "text": "Write the backup guide.", "text_elements": []}],
),
# Encrypted reasoning — nothing recoverable in either layer.
{
"timestamp": "2026-08-17T05:44:09.500Z",
"type": "response_item",
"payload": {"type": "reasoning", "summary": [], "encrypted_content": "gAAAA…"},
},
_item("Reasoning", id="r1", summary_text=[], raw_content=[]),
# AgentMessage content blocks are "Text" (capital T), unlike UserMessage.
_item(
"AgentMessage",
id="a1",
phase="commentary",
content=[{"type": "Text", "text": "I'll inspect the repo first."}],
),
_exec_call("ls -la"),
_item(
"CommandExecution",
id="exec-1",
process_id="123",
command=["/bin/bash", "-lc", "ls -la"],
cwd="file:///home/jesse/ws",
source="unified_exec_startup",
status="completed",
exit_code=0,
stdout="total 4\n",
aggregated_output="total 4\n",
formatted_output="total 4\n",
),
_exec_call("cat missing"),
_item(
"CommandExecution",
id="exec-2",
command=["/bin/bash", "-lc", "cat missing"],
cwd="file:///home/jesse/ws",
status="completed",
exit_code=1,
stderr="No such file\n",
aggregated_output="No such file\n",
),
# An attempt that never produced an item (sandbox failure / abort).
_exec_call("npm test"),
# A `wait` poll: not an attempt, must not inflate the shortfall.
{
"timestamp": "2026-08-17T05:44:12.000Z",
"type": "response_item",
"payload": {"type": "function_call", "name": "wait", "call_id": "call_w"},
},
_item(
"AgentMessage",
ts="2026-08-17T05:45:00.000Z",
id="a2",
phase="final_answer",
content=[{"type": "Text", "text": "Done — the guide is written."}],
),
]
class TestCodexProvider:
def test_scan_lists_session_with_metadata(self, tmp_path):
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
convs = prov.list_conversations()
assert len(convs) == 1
assert convs[0]["id"] == SESSION_ID
assert convs[0]["title"] == "Write the backup guide."
assert convs[0]["project"] == "ws"
# session_meta payload timestamp, not the (later) flush timestamp.
assert convs[0]["created_at"] == "2026-08-17T05:44:06.666Z"
def test_ignores_non_rollout_files(self, tmp_path):
_write_session(tmp_path, _records())
(tmp_path / "2026/08/17/notes.jsonl").write_text("{}", encoding="utf-8")
(tmp_path / "2026/08/17/rollout-garbage.jsonl").write_text("{}", encoding="utf-8")
assert len(CodexProvider(sessions_dir=tmp_path).list_conversations()) == 1
def test_empty_file_skipped(self, tmp_path):
_write_session(tmp_path, _records())
(tmp_path / "2026/08/16").mkdir(parents=True)
(tmp_path / "2026/08/16" / FILENAME.replace("17T01", "16T01")).write_text("")
assert len(CodexProvider(sessions_dir=tmp_path).list_conversations()) == 1
def test_missing_root_returns_empty(self, tmp_path):
prov = CodexProvider(sessions_dir=tmp_path / "nope")
assert prov.list_conversations() == []
def test_get_conversation_unknown_id_raises(self, tmp_path):
from src.providers.base import ProviderError
prov = CodexProvider(sessions_dir=tmp_path)
with pytest.raises(ProviderError):
prov.get_conversation("does-not-exist")
class TestNormalize:
def _normalized(self, tmp_path, policy="placeholder", records=None):
_write_session(tmp_path, records if records is not None else _records())
prov = CodexProvider(sessions_dir=tmp_path, hidden_content=policy)
prov.list_conversations()
report = LossReport()
return prov.normalize_conversation(prov.get_conversation(SESSION_ID), report), report
def test_dialogue_only_by_default(self, tmp_path):
conv, _ = self._normalized(tmp_path)
roles = [m["role"] for m in conv["messages"]]
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert roles[0] == "user"
assert texts == [
"Write the backup guide.",
"I'll inspect the repo first.",
"Done — the guide is written.",
]
def test_harness_injections_never_become_messages(self, tmp_path):
conv, _ = self._normalized(tmp_path)
blob = json.dumps(conv)
assert "AGENTS.md instructions" not in blob
assert "skills_instructions" not in blob
def test_reasoning_is_dropped_and_counted(self, tmp_path):
conv, report = self._normalized(tmp_path)
assert "thinking" not in json.dumps(conv)
assert "reasoning" in report.format_summary()
def test_reasoning_stays_dropped_under_full(self, tmp_path):
# Unlike Claude Code, `full` cannot surface it — it is encrypted at rest.
conv, _ = self._normalized(tmp_path, policy="full")
assert "encrypted" not in json.dumps(conv)
def test_tool_traffic_collapses_with_shortfall(self, tmp_path):
conv, _ = self._normalized(tmp_path)
collapsed = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
]
assert len(collapsed) == 1
origin = collapsed[0]["origin"]
# 2 completed of 3 attempted; the `wait` poll is not an attempt.
assert "2 calls: exec_command ×2" in origin
assert "+1 did not complete" in origin
def test_full_policy_emits_decoded_tool_blocks(self, tmp_path):
conv, _ = self._normalized(tmp_path, policy="full")
uses = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TOOL_USE
]
results = [
b for m in conv["messages"] for b in m["blocks"]
# The "incomplete" note is asserted by TestFullPolicyShortfall.
if b["type"] == BLOCK_TYPE_TOOL_RESULT and b.get("tool_name") != "incomplete"
]
assert len(uses) == 2 and len(results) == 2
# The command is read from the typed item, not parsed out of the JS.
assert uses[0]["input"]["command"] == "ls -la"
assert uses[0]["input"]["cwd"] == "/home/jesse/ws" # file:// stripped
assert results[0]["is_error"] is False
assert results[1]["is_error"] is True # exit_code 1
def test_message_count_matches(self, tmp_path):
conv, _ = self._normalized(tmp_path)
assert conv["message_count"] == len(conv["messages"])
def test_updated_at_matches_listing(self, tmp_path):
"""Cache staleness compares these two; a mismatch re-exports every run."""
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
listed = prov.list_conversations()[0]
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["updated_at"] == listed["updated_at"]
def test_context_compaction_is_visible(self, tmp_path):
records = _records() + [_item("ContextCompaction", id="c1")]
conv, report = self._normalized(tmp_path, records=records)
markers = [
b for m in conv["messages"] for b in m["blocks"]
if b.get("kind") == COLLAPSED_KIND_HIDDEN_CONTEXT
]
assert len(markers) == 1
assert "compacted" in markers[0]["origin"]
assert "context_compaction" in report.format_summary()
def test_unknown_item_type_is_reported(self, tmp_path):
records = _records() + [_item("QuantumMessage", id="q1", mystery=True)]
conv, report = self._normalized(tmp_path, records=records)
unknowns = [
b for m in conv["messages"] for b in m["blocks"] if b["type"] == "unknown"
]
assert len(unknowns) == 1
assert unknowns[0]["raw_type"] == "codex.QuantumMessage"
assert "codex.QuantumMessage" in report.format_summary()
def test_extension_labelled_by_kind(self, tmp_path):
records = _records() + [
_item("Extension", id="e1", kind="web.search", query="hsts", results=[]),
]
conv, _ = self._normalized(tmp_path, records=records)
origins = " ".join(
b["origin"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
)
assert "web.search ×1" in origins
class TestRepoTags:
def test_title_tagged_with_repos_touched(self, tmp_path):
repo = tmp_path / "ws" / "myrepo"
(repo / ".git").mkdir(parents=True)
records = _records(cwd=str(tmp_path / "ws")) + [
_item(
"FileChange",
id="fc1",
status="completed",
changes={str(repo / "README.md"): {"type": "add", "content": "x"}},
)
]
_write_session(tmp_path, records)
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["title"].endswith("[myrepo]")
def test_no_tag_when_nothing_touched(self, tmp_path):
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["title"] == "Write the backup guide."
class TestResolveRoots:
def test_default_root(self, monkeypatch):
monkeypatch.delenv("CODEX_DIR", raising=False)
monkeypatch.delenv("CODEX_HOME", raising=False)
assert resolve_roots()[0].name == "sessions"
def test_codex_dir_splits_on_pathsep(self, monkeypatch):
monkeypatch.setenv("CODEX_DIR", "/a/sessions:/b/sessions")
monkeypatch.delenv("CODEX_HOME", raising=False)
assert [str(p) for p in resolve_roots()] == ["/a/sessions", "/b/sessions"]
def test_codex_home_appended_and_deduped(self, monkeypatch):
monkeypatch.setenv("CODEX_DIR", "/a/sessions")
monkeypatch.setenv("CODEX_HOME", "/a")
# /a/sessions is already listed — must not appear twice.
assert [str(p) for p in resolve_roots()] == ["/a/sessions"]
class TestExoticLineBreaks:
"""U+0085 (NEL) and friends are legal inside a JSON string.
``str.splitlines()`` breaks on them, shredding one record into unparseable
fragments and losing it silently. Observed 2026-08-18 in a real rollout,
where captured command output contained two NELs.
"""
# C0 controls (\x0b, \x0c) are excluded: JSON requires them escaped, so they
# never reach the splitter literally. These three do.
@pytest.mark.parametrize("sep", ["\x85", "
", "
"])
def test_record_with_exotic_break_survives(self, tmp_path, sep):
records = _records()
records.append(
_item(
"AgentMessage",
ts="2026-08-17T05:46:00.000Z",
id="a3",
phase="final_answer",
content=[{"type": "Text", "text": f"before{sep}after"}],
)
)
_write_session(tmp_path, records)
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert f"before{sep}after" in texts
class TestFullPolicyShortfall:
def test_shortfall_note_is_not_a_collapsed_block(self, tmp_path):
"""Under `full`, a collapsed block would advise setting the policy that
is already in force. The shortfall is stated plainly instead."""
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path, hidden_content="full")
prov.list_conversations()
report = LossReport()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID), report)
collapsed = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
]
assert collapsed == []
notes = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TOOL_RESULT and b.get("tool_name") == "incomplete"
]
assert len(notes) == 1
assert "1 tool call(s) produced no result" in notes[0]["output"]
assert "tool_call_incomplete" in report.format_summary()
+68
View File
@@ -54,3 +54,71 @@ class TestValidateChatGPTToken:
result = _validate_chatgpt_token("notajwttoken")
assert any("does not look like a JWT" in r.message for r in caplog.records)
assert result is None
class TestSessionLimiterConfig:
"""MAX_CONVERSATIONS_PER_RUN and REQUEST_DELAY parsing in load_config."""
def _load(self, monkeypatch, tmp_path, **env):
from src import config as config_module
from src.config import load_config
# load_config() calls load_dotenv(override=False), which re-populates
# any variable this test just deleted from the developer's real .env —
# so test_defaults only saw defaults on a machine without one. Stub it:
# these tests are about parsing the environment, not discovering .env.
monkeypatch.setattr(config_module, "load_dotenv", lambda *a, **k: False)
monkeypatch.setenv("EXPORT_DIR", str(tmp_path / "exports"))
monkeypatch.setenv("CACHE_DIR", str(tmp_path / "cache"))
for key in (
"MAX_CONVERSATIONS_PER_RUN",
"REQUEST_DELAY",
"EXPORTER_HIDDEN_CONTENT",
"EXPORTER_DOWNLOAD_MEDIA",
):
monkeypatch.delenv(key, raising=False)
for key, value in env.items():
monkeypatch.setenv(key, value)
return load_config()
def test_defaults(self, monkeypatch, tmp_path):
cfg = self._load(monkeypatch, tmp_path)
assert cfg.max_conversations is None
assert cfg.request_delay == 1.0
assert cfg.hidden_content == "placeholder"
assert cfg.download_media == "images"
def test_download_media_valid(self, monkeypatch, tmp_path):
cfg = self._load(monkeypatch, tmp_path, EXPORTER_DOWNLOAD_MEDIA="all")
assert cfg.download_media == "all"
def test_download_media_invalid_raises(self, monkeypatch, tmp_path):
from src.config import ConfigError
with pytest.raises(ConfigError, match="EXPORTER_DOWNLOAD_MEDIA"):
self._load(monkeypatch, tmp_path, EXPORTER_DOWNLOAD_MEDIA="sometimes")
def test_valid_values(self, monkeypatch, tmp_path):
cfg = self._load(
monkeypatch, tmp_path,
MAX_CONVERSATIONS_PER_RUN="25", REQUEST_DELAY="0.5",
)
assert cfg.max_conversations == 25
assert cfg.request_delay == 0.5
def test_zero_delay_allowed(self, monkeypatch, tmp_path):
cfg = self._load(monkeypatch, tmp_path, REQUEST_DELAY="0")
assert cfg.request_delay == 0.0
def test_non_integer_cap_raises(self, monkeypatch, tmp_path):
from src.config import ConfigError
with pytest.raises(ConfigError, match="MAX_CONVERSATIONS_PER_RUN"):
self._load(monkeypatch, tmp_path, MAX_CONVERSATIONS_PER_RUN="lots")
def test_zero_cap_raises(self, monkeypatch, tmp_path):
from src.config import ConfigError
with pytest.raises(ConfigError, match="at least 1"):
self._load(monkeypatch, tmp_path, MAX_CONVERSATIONS_PER_RUN="0")
def test_negative_delay_raises(self, monkeypatch, tmp_path):
from src.config import ConfigError
with pytest.raises(ConfigError, match="REQUEST_DELAY"):
self._load(monkeypatch, tmp_path, REQUEST_DELAY="-1")
+312 -2
View File
@@ -1,4 +1,4 @@
"""Unit tests for src/exporters/."""
"""Unit tests for src/exporters/ and src/blocks.py."""
import json
import os
@@ -7,6 +7,24 @@ from pathlib import Path
import pytest
from src.blocks import (
BLOCK_TYPE_TEXT,
UNKNOWN_REASON_EXTRACTION_FAILED,
UNKNOWN_REASON_UNKNOWN_TYPE,
_blockquote_prefix,
_safe_fence,
make_code_block,
make_file_placeholder,
make_hidden_context_marker,
make_image_placeholder,
make_subagent_block,
make_text_block,
make_thinking_block,
make_tool_result_block,
make_tool_use_block,
make_unknown_block,
render_blocks_to_markdown,
)
from src.exporters.markdown import MarkdownExporter, _yaml_escape, _format_timestamp
from src.exporters.json_export import JSONExporter
@@ -102,6 +120,30 @@ class TestMarkdownFrontmatter:
assert "```python" in content
assert "print('hello')" in content
def test_subagent_block_renders_as_details(self, tmp_path):
sub = make_subagent_block(
agent_type="Explore",
description="find the thing",
messages=[
{"role": "user", "blocks": [make_text_block("Go find it")]},
{"role": "assistant", "blocks": [make_text_block("Found it here.")]},
],
)
conv = {
**SAMPLE_CONV,
"provider": "claude-code",
"messages": [
{"role": "assistant", "content_type": "text", "timestamp": None,
"blocks": [make_text_block("Delegating."), sub]},
],
}
content = MarkdownExporter(tmp_path).export(conv).read_text()
assert "<details>" in content
assert "<summary>🤖 Subagent: Explore — find the thing</summary>" in content
assert "Go find it" in content
assert "Found it here." in content
assert "</details>" in content
class TestMarkdownFilenameGeneration:
def test_filename_format(self, tmp_path):
@@ -122,7 +164,7 @@ class TestMarkdownFilenameGeneration:
def test_year_in_path(self, tmp_path):
exp = MarkdownExporter(tmp_path)
path = exp.export(SAMPLE_CONV)
assert "/2024/" in str(path)
assert ".2024/" in str(path)
def test_output_structure_provider_project(self, tmp_path):
exp = MarkdownExporter(tmp_path, output_structure="provider/project")
@@ -250,3 +292,271 @@ class TestFormatTimestamp:
def test_empty_string(self):
assert _format_timestamp("") == ""
# ---------------------------------------------------------------------------
# Block helpers and rendering
# ---------------------------------------------------------------------------
class TestSafeFence:
def test_minimum_three_backticks(self):
assert _safe_fence("plain text") == "```"
def test_four_backticks_when_three_in_content(self):
assert _safe_fence("here ``` is a fence") == "````"
def test_five_backticks_when_four_in_content(self):
assert _safe_fence("here ```` is four") == "`````"
def test_handles_empty_string(self):
assert _safe_fence("") == "```"
def test_handles_run_at_end(self):
# Trailing run still counted
assert _safe_fence("text ending in ```") == "````"
class TestBlockquotePrefix:
def test_single_line(self):
assert _blockquote_prefix("hello") == "> hello"
def test_multi_line(self):
assert _blockquote_prefix("a\nb\nc") == "> a\n> b\n> c"
def test_empty_lines_become_naked_quote_marker(self):
assert _blockquote_prefix("a\n\nb") == "> a\n>\n> b"
def test_empty_string(self):
assert _blockquote_prefix("") == ">"
class TestBlockConstructors:
def test_make_text_block_returns_none_for_empty(self):
assert make_text_block("") is None
assert make_text_block(" ") is None
def test_make_text_block_returns_dict(self):
b = make_text_block("hello")
assert b == {"type": "text", "text": "hello"}
def test_make_code_block_returns_none_for_empty(self):
assert make_code_block("") is None
def test_make_thinking_block_returns_none_for_empty(self):
assert make_thinking_block("") is None
class TestRenderBlocks:
def test_text_block_renders_as_paragraph(self):
out = render_blocks_to_markdown([make_text_block("Hello world")])
assert out == "Hello world"
def test_blocks_separated_by_blank_line(self):
out = render_blocks_to_markdown(
[make_text_block("first"), make_text_block("second")]
)
assert out == "first\n\nsecond"
def test_code_block_with_language(self):
out = render_blocks_to_markdown([make_code_block("print(1)", language="python")])
assert "```python" in out
assert "print(1)" in out
def test_thinking_block_uses_blockquote(self):
out = render_blocks_to_markdown([make_thinking_block("step 1\nstep 2")])
assert "**💭 Reasoning**" in out
assert "> step 1" in out
assert "> step 2" in out
def test_tool_use_renders_as_blockquote_with_safe_fence(self):
out = render_blocks_to_markdown(
[make_tool_use_block("search", {"query": "test"})]
)
assert "> 🔧 **Tool: search**" in out
# Every line of the body is blockquote-prefixed
assert "> ```json" in out
assert "> }" in out
def test_tool_use_with_multiline_input(self):
out = render_blocks_to_markdown(
[make_tool_use_block("complex", {"a": 1, "b": [{"x": "y"}]})]
)
# Prefix every line of multi-line JSON
for line in out.split("\n"):
assert line.startswith(">") or line == ""
def test_tool_result_success_uses_outbox_icon(self):
out = render_blocks_to_markdown([make_tool_result_block("OK")])
assert "📤 **Result**" in out
assert "❌" not in out
def test_tool_result_error_uses_x_icon(self):
out = render_blocks_to_markdown([make_tool_result_block("oops", is_error=True)])
assert "❌ **Result (error)**" in out
assert "📤" not in out
def test_tool_result_with_tool_name_in_header(self):
out = render_blocks_to_markdown(
[make_tool_result_block("done", tool_name="container.exec")]
)
assert "📤 **Result: container.exec**" in out
def test_tool_result_error_with_tool_name(self):
out = render_blocks_to_markdown(
[make_tool_result_block("503", tool_name="web", is_error=True)]
)
assert "❌ **Result (error): web**" in out
def test_tool_result_summary_renders_as_italic_line(self):
out = render_blocks_to_markdown(
[
make_tool_result_block(
"output",
tool_name="container.exec",
summary="Reading skill documentation",
)
]
)
# Summary line is italic, lives between header and fence,
# all inside the blockquote prefix.
assert "> *Reading skill documentation*" in out
# Order: header before summary before fence
header_idx = out.index("Result: container.exec")
summary_idx = out.index("Reading skill documentation")
fence_idx = out.index("output")
assert header_idx < summary_idx < fence_idx
def test_image_placeholder_rendering(self):
out = render_blocks_to_markdown(
[make_image_placeholder(ref="file-123", source="user_upload")]
)
assert "🖼️ **Image attached**" in out
assert "`file-123`" in out
assert "user_upload" in out
assert "content not preserved" in out
def test_file_placeholder_with_metadata(self):
out = render_blocks_to_markdown(
[make_file_placeholder(ref="sediment://x", mime="audio/wav", size_bytes=10240, duration_seconds=2.5)]
)
assert "📎 **File attached**" in out
assert "audio/wav" in out
assert "KB" in out
assert "2.50s" in out
def test_unknown_block_renders_with_keys(self):
out = render_blocks_to_markdown(
[
make_unknown_block(
raw_type="future_x",
observed_keys=["foo", "bar"],
reason=UNKNOWN_REASON_UNKNOWN_TYPE,
)
]
)
assert "⚠️ **Unsupported content**" in out
assert "future_x" in out
assert "`foo`" in out
assert "`bar`" in out
def test_unknown_extraction_failed_includes_summary(self):
out = render_blocks_to_markdown(
[
make_unknown_block(
raw_type="audio_transcription",
observed_keys=["asset_pointer"],
reason=UNKNOWN_REASON_EXTRACTION_FAILED,
summary="expected key 'text' not found",
)
]
)
assert "extraction_failed" in out
assert "expected key 'text' not found" in out
def test_hidden_context_marker(self):
out = render_blocks_to_markdown(
[make_hidden_context_marker("user_editable_context")]
)
assert "ℹ️ **Hidden context**" in out
assert "`user_editable_context`" in out
def test_safe_fence_prevents_runaway_code_block(self):
# Content contains an unbalanced opening fence — without _safe_fence
# this would corrupt downstream rendering.
evil_content = "before\n```Follow\ntext\nraw is: \"```"
block = make_code_block(evil_content)
out = render_blocks_to_markdown([block, make_text_block("after")])
# The 4-backtick wrap should be present
assert "````" in out
# The "after" text should appear OUTSIDE any code block — it follows
# the closing ```` fence.
assert out.endswith("after")
def test_block_order_preserved(self):
blocks = [
make_text_block("a"),
make_image_placeholder(ref="r1", source="user_upload"),
make_text_block("b"),
]
out = render_blocks_to_markdown(blocks)
assert out.index("a") < out.index("Image attached")
assert out.index("Image attached") < out.index("b")
# ---------------------------------------------------------------------------
# Markdown exporter with blocks
# ---------------------------------------------------------------------------
SAMPLE_CONV_BLOCKS = {
"id": "blocks12345",
"title": "Blocks Conversation",
"provider": "claude",
"project": None,
"created_at": "2024-06-10T14:32:00Z",
"updated_at": "2024-06-10T15:00:00Z",
"message_count": 1,
"messages": [
{
"role": "assistant",
"content_type": "text",
"timestamp": None,
"blocks": [
{"type": "text", "text": "Here is the answer."},
{"type": "tool_use", "name": "search", "input": {"q": "x"}, "tool_id": "t1"},
],
}
],
}
class TestMarkdownExporterWithBlocks:
def test_renders_blocks(self, tmp_path):
exp = MarkdownExporter(tmp_path)
path = exp.export(SAMPLE_CONV_BLOCKS)
body = path.read_text()
assert "Here is the answer." in body
assert "🔧 **Tool: search**" in body
def test_falls_back_to_content_when_blocks_missing(self, tmp_path):
# Backward-compat: messages with `content` only (no `blocks`) still render.
exp = MarkdownExporter(tmp_path)
path = exp.export(SAMPLE_CONV) # SAMPLE_CONV has content only, no blocks
body = path.read_text()
assert "Hello, how are you?" in body
def test_skips_messages_with_neither_blocks_nor_content(self, tmp_path):
conv = {
**SAMPLE_CONV_BLOCKS,
"messages": [
{"role": "user", "content_type": "text", "timestamp": None, "blocks": []},
{"role": "assistant", "content_type": "text", "timestamp": None, "blocks": [
{"type": "text", "text": "I am here."}
]},
],
}
exp = MarkdownExporter(tmp_path)
path = exp.export(conv)
body = path.read_text()
assert "I am here." in body
+130 -16
View File
@@ -5,7 +5,7 @@ from unittest.mock import MagicMock, patch
import pytest
import requests
from src.joplin import JoplinClient, JoplinError, _http_error_message, _timeout_message, notebook_title
from src.joplin import JoplinClient, JoplinError, _http_error_message, _timeout_message, notebook_path
# ---------------------------------------------------------------------------
@@ -31,25 +31,32 @@ def _mock_response(json_data=None, text="", status_code=200):
# ---------------------------------------------------------------------------
# notebook_title helper
# notebook_path helper
# ---------------------------------------------------------------------------
class TestNotebookTitle:
class TestNotebookPath:
def test_no_project(self):
assert notebook_title("chatgpt", None) == "ChatGPT - No Project"
assert notebook_path("chatgpt", None) == ("AI-ChatGPT", "No Project")
def test_no_project_string(self):
assert notebook_title("chatgpt", "no-project") == "ChatGPT - No Project"
assert notebook_path("chatgpt", "no-project") == ("AI-ChatGPT", "No Project")
def test_project_with_hyphens(self):
assert notebook_title("chatgpt", "my-project") == "ChatGPT - My Project"
assert notebook_path("chatgpt", "my-project") == ("AI-ChatGPT", "My Project")
def test_claude_provider(self):
assert notebook_title("claude", "budget-tracker") == "Claude - Budget Tracker"
assert notebook_path("claude", "budget-tracker") == ("AI-Claude", "Budget Tracker")
def test_claude_code_gets_own_top_level_notebook(self):
assert notebook_path("claude-code", "services") == ("AI-ClaudeCode", "Services")
def test_multi_word_project(self):
assert notebook_title("claude", "ai-research-notes") == "Claude - Ai Research Notes"
assert notebook_path("claude", "ai-research-notes") == ("AI-Claude", "Ai Research Notes")
def test_returns_tuple(self):
result = notebook_path("chatgpt", "some-project")
assert isinstance(result, tuple) and len(result) == 2
# ---------------------------------------------------------------------------
@@ -57,6 +64,24 @@ class TestNotebookTitle:
# ---------------------------------------------------------------------------
class TestUpdateNote:
def test_parent_id_moves_note(self):
client = _make_client()
client._put = MagicMock()
client.update_note("nid", "Title", "body", parent_id="nb1")
args, _ = client._put.call_args
assert args[0] == "/notes/nid"
assert args[1]["parent_id"] == "nb1"
assert args[1]["title"] == "Title"
def test_no_parent_id_omits_field(self):
client = _make_client()
client._put = MagicMock()
client.update_note("nid", "Title", "body")
args, _ = client._put.call_args
assert "parent_id" not in args[1]
class TestPing:
def test_ping_success(self):
client = _make_client()
@@ -236,18 +261,30 @@ class TestListNotebooks:
class TestGetOrCreateNotebook:
def test_returns_existing_notebook_id(self):
def test_returns_existing_root_notebook_id(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={
"items": [{"id": "nb-existing", "title": "ChatGPT - No Project"}],
"items": [{"id": "nb-existing", "title": "AI-ChatGPT", "parent_id": ""}],
"has_more": False,
}
)
nb_id = client.get_or_create_notebook("ChatGPT - No Project")
nb_id = client.get_or_create_notebook("AI-ChatGPT")
assert nb_id == "nb-existing"
def test_returns_existing_child_notebook_id(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={
"items": [{"id": "nb-child", "title": "No Project", "parent_id": "nb-parent"}],
"has_more": False,
}
)
nb_id = client.get_or_create_notebook("No Project", parent_id="nb-parent")
assert nb_id == "nb-child"
def test_creates_new_notebook_when_not_found(self):
client = _make_client()
with patch("requests.get") as mock_get, patch("requests.post") as mock_post:
@@ -255,26 +292,103 @@ class TestGetOrCreateNotebook:
json_data={"items": [], "has_more": False}
)
mock_post.return_value = _mock_response(
json_data={"id": "nb-new", "title": "ChatGPT - New Project"}
json_data={"id": "nb-new", "title": "AI-ChatGPT"}
)
nb_id = client.get_or_create_notebook("ChatGPT - New Project")
nb_id = client.get_or_create_notebook("AI-ChatGPT")
assert nb_id == "nb-new"
mock_post.assert_called_once()
def test_creates_child_notebook_with_parent_id(self):
client = _make_client()
with patch("requests.get") as mock_get, patch("requests.post") as mock_post:
mock_get.return_value = _mock_response(
json_data={"items": [], "has_more": False}
)
mock_post.return_value = _mock_response(
json_data={"id": "nb-child", "title": "My Project"}
)
nb_id = client.get_or_create_notebook("My Project", parent_id="nb-parent")
assert nb_id == "nb-child"
_, kwargs = mock_post.call_args
assert kwargs["json"]["parent_id"] == "nb-parent"
def test_does_not_include_parent_id_for_root(self):
client = _make_client()
with patch("requests.get") as mock_get, patch("requests.post") as mock_post:
mock_get.return_value = _mock_response(json_data={"items": [], "has_more": False})
mock_post.return_value = _mock_response(json_data={"id": "nb-root", "title": "AI-Claude"})
client.get_or_create_notebook("AI-Claude")
_, kwargs = mock_post.call_args
assert "parent_id" not in kwargs["json"]
def test_caches_notebook_after_first_load(self):
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={
"items": [{"id": "nb1", "title": "Claude - No Project"}],
"items": [{"id": "nb1", "title": "AI-Claude", "parent_id": ""}],
"has_more": False,
}
)
# Call twice — GET /folders should only happen once
client.get_or_create_notebook("Claude - No Project")
client.get_or_create_notebook("Claude - No Project")
client.get_or_create_notebook("AI-Claude")
client.get_or_create_notebook("AI-Claude")
assert mock_get.call_count == 1
def test_different_parent_ids_are_distinct_cache_entries(self):
"""Same title under different parents are different notebooks."""
client = _make_client()
with patch("requests.get") as mock_get:
mock_get.return_value = _mock_response(
json_data={
"items": [
{"id": "nb-a", "title": "No Project", "parent_id": "parent-chatgpt"},
{"id": "nb-b", "title": "No Project", "parent_id": "parent-claude"},
],
"has_more": False,
}
)
id_a = client.get_or_create_notebook("No Project", parent_id="parent-chatgpt")
id_b = client.get_or_create_notebook("No Project", parent_id="parent-claude")
assert id_a == "nb-a"
assert id_b == "nb-b"
class TestGetOrCreateNotebookPath:
def test_creates_two_level_path(self):
client = _make_client()
with patch("requests.get") as mock_get, patch("requests.post") as mock_post:
mock_get.return_value = _mock_response(json_data={"items": [], "has_more": False})
mock_post.side_effect = [
_mock_response(json_data={"id": "nb-parent", "title": "AI-ChatGPT"}),
_mock_response(json_data={"id": "nb-child", "title": "No Project"}),
]
leaf_id = client.get_or_create_notebook_path(["AI-ChatGPT", "No Project"])
assert leaf_id == "nb-child"
assert mock_post.call_count == 2
# Second POST should use the parent's ID
_, kwargs = mock_post.call_args_list[1]
assert kwargs["json"]["parent_id"] == "nb-parent"
def test_reuses_existing_parent_for_new_child(self):
client = _make_client()
with patch("requests.get") as mock_get, patch("requests.post") as mock_post:
mock_get.return_value = _mock_response(
json_data={
"items": [{"id": "nb-parent", "title": "AI-Claude", "parent_id": ""}],
"has_more": False,
}
)
mock_post.return_value = _mock_response(
json_data={"id": "nb-child", "title": "Budget Tracker"}
)
leaf_id = client.get_or_create_notebook_path(["AI-Claude", "Budget Tracker"])
assert leaf_id == "nb-child"
# Only one POST — the parent already existed
assert mock_post.call_count == 1
_, kwargs = mock_post.call_args
assert kwargs["json"]["parent_id"] == "nb-parent"
# ---------------------------------------------------------------------------
# create_note
+264
View File
@@ -0,0 +1,264 @@
"""Tests for media downloads and Joplin resource rewriting."""
from pathlib import Path
import pytest
from src.blocks import (
make_file_placeholder,
make_image_placeholder,
render_blocks_to_markdown,
)
from src.loss_report import LossReport
from src.media import resolve_media, resolve_media_policy
from src.providers.base import ProviderError
from src.providers.chatgpt import parse_asset_file_id
# ---------------------------------------------------------------------------
# Asset reference parsing
# ---------------------------------------------------------------------------
class TestParseAssetFileId:
def test_plain_sediment(self):
assert parse_asset_file_id("sediment://file_00000000245c71fda5") == "file_00000000245c71fda5"
def test_generated_image_with_hash_and_page(self):
ref = "sediment://8456107fc383a53#file_00000000979c71f685#p_6.png"
assert parse_asset_file_id(ref) == "file_00000000979c71f685"
def test_file_service_scheme(self):
assert parse_asset_file_id("file-service://file-AbCdEf") == "file-AbCdEf"
def test_unrecognised(self):
assert parse_asset_file_id("https://example.com/x.png") is None
assert parse_asset_file_id("") is None
assert parse_asset_file_id(None) is None
# ---------------------------------------------------------------------------
# Media policy
# ---------------------------------------------------------------------------
class TestMediaPolicy:
def test_default(self, monkeypatch):
monkeypatch.delenv("EXPORTER_DOWNLOAD_MEDIA", raising=False)
assert resolve_media_policy() == "images"
def test_valid(self, monkeypatch):
monkeypatch.setenv("EXPORTER_DOWNLOAD_MEDIA", "all")
assert resolve_media_policy() == "all"
def test_invalid_falls_back(self, monkeypatch, caplog):
monkeypatch.setenv("EXPORTER_DOWNLOAD_MEDIA", "bogus")
assert resolve_media_policy() == "images"
# ---------------------------------------------------------------------------
# resolve_media
# ---------------------------------------------------------------------------
class _FakeProvider:
"""Minimal provider exposing download_asset + the ref parser."""
def __init__(self, assets=None, fail_refs=None):
self._assets = assets or {}
self._fail_refs = fail_refs or {}
self.calls = []
def parse_asset_file_id(self, ref):
return parse_asset_file_id(ref)
def download_asset(self, ref):
self.calls.append(ref)
if ref in self._fail_refs:
raise ProviderError("chatgpt", "download_asset", self._fail_refs[ref])
return self._assets[ref] # (content, mime, file_name)
def _conv_with(blocks):
return {
"id": "conv-1",
"title": "Has Media",
"provider": "chatgpt",
"project": None,
"created_at": "2026-05-20T00:00:00+00:00",
"messages": [{"role": "user", "blocks": blocks}],
}
class TestResolveMedia:
def test_downloads_image_and_inlines(self, tmp_path):
ref = "sediment://file_img1"
provider = _FakeProvider({ref: (b"\x89PNG\r\n", "image/png", "x.png")})
block = make_image_placeholder(ref=ref, source="user_upload")
conv = _conv_with([block])
report = LossReport()
n = resolve_media(conv, provider, tmp_path, "provider/project/year", "images", report)
assert n == 1
assert report.media_downloaded == 1
assert block["local_path"] == "media/file_img1.png"
rendered = render_blocks_to_markdown([block])
assert rendered == "![user_upload](media/file_img1.png)"
# File written under the conversation's media/ dir
written = list(tmp_path.rglob("media/file_img1.png"))
assert written and written[0].read_bytes() == b"\x89PNG\r\n"
def test_images_policy_skips_files(self, tmp_path):
ref = "sediment://file_audio1"
provider = _FakeProvider({ref: (b"RIFF", "audio/wav", "a.wav")})
block = make_file_placeholder(ref=ref, mime="audio/wav", size_bytes=1000)
conv = _conv_with([block])
report = LossReport()
n = resolve_media(conv, provider, tmp_path, "provider/project/year", "images", report)
assert n == 0
assert "local_path" not in block
assert provider.calls == []
def test_all_policy_downloads_files(self, tmp_path):
ref = "sediment://file_audio1"
provider = _FakeProvider({ref: (b"RIFFdata", "audio/wav", "a.wav")})
block = make_file_placeholder(ref=ref, mime="audio/wav", size_bytes=8)
conv = _conv_with([block])
report = LossReport()
n = resolve_media(conv, provider, tmp_path, "provider/project/year", "all", report)
assert n == 1
assert block["local_path"] == "media/file_audio1.wav"
rendered = render_blocks_to_markdown([block])
assert "media/file_audio1.wav" in rendered and rendered.startswith("> 📎")
def test_off_policy_noop(self, tmp_path):
provider = _FakeProvider({"sediment://file_x": (b"x", "image/png", None)})
block = make_image_placeholder(ref="sediment://file_x", source="user_upload")
conv = _conv_with([block])
report = LossReport()
assert resolve_media(conv, provider, tmp_path, "provider/project/year", "off", report) == 0
assert provider.calls == []
def test_idempotent_uses_disk(self, tmp_path):
ref = "sediment://file_img1"
provider = _FakeProvider({ref: (b"\x89PNG", "image/png", "x.png")})
conv = _conv_with([make_image_placeholder(ref=ref, source="user_upload")])
report = LossReport()
resolve_media(conv, provider, tmp_path, "provider/project/year", "images", report)
assert len(provider.calls) == 1
# Second run, fresh blocks: file already on disk → no new API call.
conv2 = _conv_with([make_image_placeholder(ref=ref, source="user_upload")])
resolve_media(conv2, provider, tmp_path, "provider/project/year", "images", LossReport())
assert len(provider.calls) == 1 # unchanged
assert conv2["messages"][0]["blocks"][0]["local_path"] == "media/file_img1.png"
def test_failure_keeps_placeholder_and_counts(self, tmp_path):
ref = "sediment://file_gone"
provider = _FakeProvider(fail_refs={ref: RuntimeError("Signed URL returned HTTP 404")})
block = make_image_placeholder(ref=ref, source="model_generated")
conv = _conv_with([block])
report = LossReport()
n = resolve_media(conv, provider, tmp_path, "provider/project/year", "images", report)
assert n == 0
assert "local_path" not in block
assert report.media_failed["expired-or-missing"] == 1
# Still renders as a placeholder, not a broken image link
assert render_blocks_to_markdown([block]).startswith("> 🖼️")
def test_forbidden_counted_separately_from_generic_error(self, tmp_path):
"""403 is a distinct bucket: the asset exists, we were refused."""
ref = "sediment://file_denied"
provider = _FakeProvider(
fail_refs={ref: RuntimeError("HTTP 403 — detail: unauthorized")}
)
block = make_image_placeholder(ref=ref, source="model_generated")
report = LossReport()
resolve_media(
_conv_with([block]), provider, tmp_path, "provider/project/year",
"images", report,
)
assert report.media_failed["forbidden"] == 1
assert "download-error" not in report.media_failed
def test_provider_without_download_asset(self, tmp_path):
"""claude-code has no remote assets — resolve_media must no-op."""
class NoDownload:
pass
block = make_image_placeholder(ref="sediment://file_x", source="user_upload")
conv = _conv_with([block])
assert resolve_media(conv, NoDownload(), tmp_path, "provider/project/year", "images", LossReport()) == 0
# ---------------------------------------------------------------------------
# Joplin media rewriting
# ---------------------------------------------------------------------------
class _FakeJoplin:
def __init__(self):
self.uploaded = []
self._n = 0
def create_resource(self, file_path, title=None):
self.uploaded.append(Path(file_path).name)
self._n += 1
return f"res{self._n}"
class TestUploadMediaAndRewrite:
def _note(self, tmp_path, body):
media = tmp_path / "media"
media.mkdir()
(media / "file_img1.png").write_bytes(b"\x89PNG")
(media / "clip.wav").write_bytes(b"RIFF")
return body
def test_rewrites_image_and_file_links(self, tmp_path):
from src.joplin import upload_media_and_rewrite
body = self._note(
tmp_path,
"Look: ![user_upload](media/file_img1.png)\n"
"> 📎 **File attached** — [clip.wav](media/clip.wav) (audio/wav)",
)
client = _FakeJoplin()
new_body, res_map = upload_media_and_rewrite(body, tmp_path, client, {})
assert "![user_upload](:/res1)" in new_body
assert "(:/res2)" in new_body
assert "media/" not in new_body
assert res_map == {"media/file_img1.png": "res1", "media/clip.wav": "res2"}
assert len(client.uploaded) == 2
def test_reuses_known_resource_ids(self, tmp_path):
from src.joplin import upload_media_and_rewrite
body = self._note(tmp_path, "![x](media/file_img1.png)")
client = _FakeJoplin()
existing = {"media/file_img1.png": "existing-res"}
new_body, res_map = upload_media_and_rewrite(body, tmp_path, client, existing)
assert "(:/existing-res)" in new_body
assert client.uploaded == [] # no re-upload
assert res_map == existing
def test_missing_file_left_as_link(self, tmp_path):
from src.joplin import upload_media_and_rewrite
(tmp_path / "media").mkdir()
body = "![x](media/gone.png)"
client = _FakeJoplin()
new_body, res_map = upload_media_and_rewrite(body, tmp_path, client, {})
assert new_body == body
assert res_map == {}
assert client.uploaded == []
def test_no_media_links_untouched(self, tmp_path):
from src.joplin import upload_media_and_rewrite
body = "Just text with [a link](https://example.com) and `media/foo` in code."
client = _FakeJoplin()
new_body, res_map = upload_media_and_rewrite(body, tmp_path, client, {})
assert new_body == body
assert res_map == {}
+1366 -67
View File
File diff suppressed because it is too large Load Diff
+20 -2
View File
@@ -57,13 +57,13 @@ class TestBuildExportPath:
path = build_export_path(
Path("/exports"), "claude", "my-project", "2024-06-01T00:00:00Z", "file.md"
)
assert str(path) == "/exports/claude/my-project/2024/file.md"
assert str(path) == "/exports/claude/my-project.2024/file.md"
def test_no_project_uses_no_project_slug(self):
path = build_export_path(
Path("/exports"), "chatgpt", None, "2024-06-01T00:00:00Z", "file.md"
)
assert "no-project" in str(path)
assert "no-project.2024" in str(path)
def test_provider_project_structure_omits_year(self):
path = build_export_path(
@@ -145,3 +145,21 @@ class TestFormatTokenStatus:
expiry = datetime.now(tz=timezone.utc) + timedelta(days=10, hours=12)
result = format_token_status("tok", expiry)
assert "10 days" in result
class TestRedactCompoundKeys:
"""Exact-match redaction let compound secret names through into logs."""
def test_compound_secret_keys_redacted(self):
result = redact_secrets(
{"access_token": "sk-abc", "api_key": "k1", "session-token": "s1"}
)
assert result == {
"access_token": "[REDACTED]",
"api_key": "[REDACTED]",
"session-token": "[REDACTED]",
}
def test_innocent_keys_containing_a_secret_word_kept(self):
result = redact_secrets({"keywords": ["a"], "monkey": "b", "tokenizer": "c"})
assert result == {"keywords": ["a"], "monkey": "b", "tokenizer": "c"}