22 Commits
Author SHA1 Message Date
JesseMarkowitz d3745e1de4 fix: test suite sent real push notifications via the developer's .env 2026-08-18 14:26:05 -04:00
JesseMarkowitz 2d5fcb26f5 release: v0.9.0, and rewrite FUTURE.md around what is actually planned
FUTURE.md had become a 648-line archaeological record: six shipped roadmap
items, two dropped, two implemented, an investigation trail for each, and a
backlog closed as not needed — with the two genuinely planned items buried
at the bottom. It is now 139 lines, planned work first.

- Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon
  behind decisions now recorded in one line each — why Brave cookie
  extraction is not viable, why the ChatGPT token's expiry cannot be read
  client-side, how the drift canary was designed — which would be expensive
  to rediscover.

- Roadmap is now two items: the StartOS service (including the 2026-08-18
  decision that clients upload to StartOS storage and the service owns the
  Joplin connection), and the README split.

- Everything shipped or dropped is off the roadmap. Decisions not to build
  survive as a one-line table so they are not re-proposed.

- Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0.

Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased]
section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the
launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy
notifications, gizmo_id project attribution and the `projects` command, and
the splitlines data-loss fix in both local providers.

Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose —
exported conversations are Markdown and may contain private content — with
each doc re-included by name, so a new doc is silently untracked rather than
rejected. The archive would not have been committed at all. Recorded in the
README-split roadmap item, since `docs/*.md` will hit exactly this.

Also refreshes the package description, which still named only ChatGPT and
Claude after two local providers were added.

Carries the documentation and table-of-contents work from earlier today.
2026-08-18 13:51:34 -04:00
JesseMarkowitz 55b9ce12f6 added ntfy support for scheduled runs 2026-08-18 11:21:26 -04:00
JesseMarkowitz b0aa05a2b1 fix sync bug for scheduled work 2026-08-18 10:31:28 -04:00
JesseMarkowitz 999429f61e add alias to avoid explicit python commands, scheduled runs and bug fixes 2026-08-18 09:59:15 -04:00
JesseMarkowitz fe5ed341ad added support for codex provider. fixed bug in claude code export (lines split when shouldn't be) 2026-08-18 08:01:34 -04:00
JesseMarkowitz d083aea135 feat: projects command — find the project IDs the config is missing
Reading gizmo_id during normalization fixed attribution, but it cannot
answer "which projects am I missing?": normalization only runs on
conversations being exported, and a normal run skips everything already
cached. Discovering the gaps would have meant --force re-exporting the
whole archive.

`ai-chat-exporter projects` does it directly. It lists conversations,
collects the g-p- ids they belong to, resolves display names, and prints a
table marking which are absent from .env plus a paste-ready
CHATGPT_PROJECT_IDS line. --write applies it; --deep falls back to one
detail request per conversation when the listing does not carry gizmo_id
(unverified which shape this account returns, so the command reports which
path it took rather than assuming).

This matters beyond tidiness: attribution is now self-correcting, but the
listing pass still needs the ids. Conversations that live only inside a
project never appear in the default listing, so an unlisted project's chats
are not merely misfiled — they are never fetched.

6 CLI tests: reporting, the already-configured case, the g-p- guard, --deep,
the hint when --deep is needed, and --write. 324 pass.
2026-08-17 13:19:34 -04:00
JesseMarkowitz 710889b65f fix: read project attribution from the conversation's gizmo_id
A conversation in a project absent from CHATGPT_PROJECT_IDS exported into
no-project/ even though its own payload names the project. Found while
investigating the media 403s: bpi-f3-case-options sits in
g-p-6a4edf5160848191927b05da20f49151, which is not among the 13 configured
ids, so it filed under no-project.2026.

Both existing sources are bounded by what the user configured. The detail
response is not: it carries gizmo_id. Use it as a third fallback, after the
listing annotation and the project map, and cache the result into the map.
Attribution now stays correct with no list to maintain, and moving a chat
into a new project stops silently misfiling it.

Only g-p- ids are treated as projects — a custom GPT is not a project and
must not become a folder.

CHATGPT_PROJECT_IDS still matters for the listing pass: conversations that
live only inside a project never appear in the default listing, so an
unconfigured project's chats can be missed entirely. Each one is now
reported once per run with the id to add, which turns "some chats are
missing" into a line to paste.

8 tests cover precedence, the g-p- guard, map caching and the once-per-run
report. 318 pass.
2026-08-17 12:49:04 -04:00
JesseMarkowitz dfa0645fba close the media 403 investigation: the bytes are gone
The render check settled it. Scrolling the whole conversation found sets of
images that do NOT display in ChatGPT — a set of 6 and a set of 4 — and they
are exactly our failures: the 6 are the records that 404 outright, the 4 are
the refused attachment batch we dumped (image.png, image(1).png,
image(2).png + 1). Every set that displays downloaded fine.

So the 403 was never OpenAI withholding something. It is the same failure
their own UI hits. All 19 are unrecoverable.

What the investigation established, now recorded in media.py and the
changelog so it is not repeated:
  - uploads do not expire (36/36 sampled, 2025-09 → 2026-08, still live), so
    export cadence was never the variable
  - 7 records 404; 12 report state=ready with a library_file_id
  - /files/{id}/download mints the signed estuary/content URL the UI fetches,
    and refuses the survivors regardless of headers, Authorization, gizmo_id,
    conversation_id, Referer or namespace; the Library id is file_not_found
  - all survivors were created 2026-07-14 within minutes of each other yet
    appear in conversations predating that date — a Library migration that
    kept the metadata and lost the bytes

Dropping the nine one-off probes; their findings live in the code comments
and changelog, and git history has the scripts if they are ever wanted.
2026-08-17 12:47:05 -04:00
JesseMarkowitz d30a9510bb tools: find what the browser sends that we don't
The decisive fact arrived: the images DO render in ChatGPT. So a valid
signature exists for files that hand us 403, and there is a route to find.

The rest is now pinned down. /files/{id}/download is the minting endpoint —
a working file's download_url is the same estuary/content URL the browser
uses. Signatures are per-file: swapping an id into a working URL returns
"Invalid signature or expired URL", and a ts without a sig gets the same.
So the difference is in the request, not the route.

Our call sends Authorization: Bearer plus cookies. A browser leans on
cookies and adds client headers a gated endpoint may check. This walks
those variants — no Authorization, conversation Referer, oai-device-id,
oai-language, sec-fetch-*, an image Accept, a conversation_id param —
against a file that is currently refused, and reports which mints a URL.

If none do, it prints the DevTools recipe for finding the minting call by
searching the Network panel for the file id.
2026-08-17 11:07:18 -04:00
JesseMarkowitz b1da1df986 tools: probe a refused file, not a deleted one
Confirmed by the last run: a working file's download_url is
chatgpt.com/backend-api/estuary/content?cid&id&p&sig&ts&v — the same route
the browser uses. So /files/{id}/download is the minting endpoint, and a
gizmo file's 403 is a refusal to mint the signature. That is why no amount
of scoping helped; we were turned away at the only door that issues them.
Nothing else in the estuary namespace serves files: 404 across the board.

But section B tested file_000000001b3071f5…, one of the seven DELETED
files, because it took the first failure without checking its class. Those
results say nothing about the eleven recoverable ones.

- Pick the target by probing /files/{id} and taking one that answers 200,
  skipping (and naming) the deleted ones.
- Add two experiments: estuary/content with a ts but no sig, since
  validation asked only for id/p/ts and never sig; and a working file's
  freshly minted URL with the refused id swapped in, which shows whether
  the signature is bound to the file.
2026-08-17 10:57:52 -04:00
JesseMarkowitz 726f57bdf9 tools: probe the estuary namespace
A URL copied from the browser gave us the route we never tried:

    /backend-api/estuary/content?id=…&ts=496382&p=fs&cid=1&sig=…&v=0

Fetched with no cookies it returns 403 {"detail":"File stream access
denied."}, so the signature rides on the session rather than replacing it.
Every earlier probe lived under /backend-api/files/*; estuary/* is new
ground.

estuary_probe.py checks two things:

A. What /files/{id}/download hands back for a file that works. If its
   download_url is an estuary URL, that endpoint is the minting step, and a
   gizmo file's 403 is a refusal to mint — which is why no amount of
   scoping helped.
B. Whether the estuary namespace exposes a route that serves a refused
   file directly.

It also replays a pasted URL through the exporter's session, and says
plainly whether the id in that URL is one of the refused files — the two
captured so far were working images, so they showed the shape without
telling us whether the broken ones have a URL at all.
2026-08-17 10:50:23 -04:00
JesseMarkowitz 024bfde030 tools: test whether a browser image URL works from the exporter
The endpoint hunt is done guessing. Results:

  /files/{libfile}/download    200 {"error_code":"file_not_found",
                                    "error_type":"GetDownloadLinkError"}
                               — twice, on two different Library ids, so the
                                 Library id is simply not valid there
  /library/*                   404 across every shape
  POST on the download route   405
  /content?asset_pointer=      422 requiring query params id, ts, p

That last one is the find: /backend-api/content is a signed route, and the
web UI has the values we cannot compute. So the question shifts from "which
endpoint" to "is the UI's URL reusable outside the browser".

try_pasted_url.py takes a URL copied from DevTools and fetches it three
ways — bare, through the exporter's authenticated session, and with the
Authorization header removed in case bearer and signature conflict. The
pattern says whether the exporter can fetch these at all, and therefore
whether it is worth hunting for what mints the signature.

Before any of that: check whether the images still render in ChatGPT's own
UI. If they show broken there, file_not_found is literally true, these 11
are lost like the other 7, and no route exists to find.
2026-08-17 10:27:57 -04:00
JesseMarkowitz ffd01ebcf3 tools: print the error body the Library endpoint returns
The pairing came back perfectly clean, and it is the whole diagnosis:

    has library_file_id  → refused (11/11)
    no  library_file_id  → deleted (7/7)

And /files/{libfile_id}/download answered 200 — not 403, not 404 — with
{status, error_code, error_type, error_message}. That endpoint accepts the
Library id; it is returning an application-level error inside an HTTP 200.
The probe printed only the keys and dropped the message, which is the one
thing that says what the call is missing.

- show() now prints the full body for every response, and detects a
  non-JSON 200 as possible raw bytes.
- Added routes worth ruling in or out: the same call as POST, with a
  conversation_id, /files/{lib}/content, /content?asset_pointer=, and four
  Library listing endpoints — a listing usually reveals both the id form
  the UI uses and the route that actually serves bytes.
- Probes a second Library id too, so one odd file cannot mislead us.
2026-08-17 10:14:18 -04:00
JesseMarkowitz 3f82b35a55 tools: try fetching refused images by their Library ID
Dumping the raw message metadata found the identity the asset pointer never
carried:

    "id":              "file_000000003454722f9481506b96aed510"   ← refused
    "library_file_id": "libfile_4eb82f478fe081919127e2eba9886e86"
    "source":          "local"

The exporter only knows the sediment id from the asset_pointer and asks
/files/{sediment}/download, which 403s for these. The Library is a separate
store with its own ids, so we have been asking for the conversation-scoped
copy of a file that now lives in the Library. That also fits the 2026-07-14
creation-time cluster: a Library migration would mint exactly these new
records.

Neither gizmo_id, conversation_id, a /gizmos path nor a project Referer
helped (403/404), so scope was never the missing piece — identity was.

library_download_probe.py pairs each refused sediment id with its
library_file_id from message.metadata.attachments and tries the endpoints
that could serve it.
2026-08-17 09:33:05 -04:00
JesseMarkowitz 03646009b9 tools: probe how to download a gizmo-scoped file
metadata_diff found the discriminator. Every refused image carries
use_case="gizmo" — ChatGPT's name for Projects and custom GPTs — while
everything that downloads is image_gen or multimodal. All are state=ready,
so nothing is damaged: the plain /files/{id}/download endpoint just will
not serve a project-scoped file.

Second clue: all 11 were created 2026-07-14 within minutes of each other,
yet appear in conversations dated 2025-11 through 2026-08, several
predating their own creation_time. Something on July 14 re-created them as
gizmo-scoped copies and repointed the conversations — which is also why
the losses start in July. Not a policy change, an event.

gizmo_download_probe.py dumps the raw conversation part carrying one of
these assets (it may simply name the scope the download wants) and then
tries the plausible calls — gizmo_id/conversation_id query params, a
/gizmos/{id}/files path, a project Referer. A 200 is the fix.
2026-08-17 09:22:48 -04:00
JesseMarkowitz a3ac279e39 tools: find what separates a refused image from a served one
branch_check killed the abandoned-branch hypothesis — every lost image is
on the live branch, and the 3 images that do sit on abandoned branches are
alive. But it turned up something the earlier probing missed by sampling
four IDs from one family and generalising: of the 19 failures, only 7 are
actually gone. The other 12 answer /files/{id} with 200 and full metadata
and refuse only /download. They exist, and may be recoverable.

Three states, then: gone (404), refused (200 meta + 403 download), working
(200 both). Since metadata comes back for the refused ones, the
discriminator can be read straight off — fetch it for every file in each
state and compare fields, flagging any field whose values never overlap
between states.

Also re-checks /download now: if a file refused during the export serves
today, those 403s were transient and a retry pass recovers them, which is
a completely different fix from anything permanent.
2026-08-17 09:18:32 -04:00
JesseMarkowitz 021a76c628 tools: test whether lost images sit on abandoned branches
Expiry is now ruled out: 36/36 sampled images from 2025-09 through
2026-08 are still live server-side, so nothing dies of age and export
cadence is not the variable. The loss is per-asset — 2026-07-09 kept 38
images and lost 12 in one conversation.

Next hypothesis: those images belong to messages that were edited or
regenerated. ChatGPT stores a conversation as a tree and editing forks it,
stranding the superseded messages on an abandoned branch. The exporter
walks every node (chatgpt.py:963), so it exports those branches too — and
an attachment unreachable from live history is a natural GC target.

branch_check.py fetches the raw conversation, derives the live branch by
walking parent links up from current_node, and cross-tabulates every image
against (on the live branch?) x (still downloadable?). If the lost images
are all off-branch and nothing on the live branch is missing, this is not
data loss at all — it is attachments to messages that were replaced.
2026-08-17 08:58:20 -04:00
JesseMarkowitz b9e8896b33 tools: distinguish "captured in time" from "still alive"
analyze_media_age.py claimed age was ruled out because an image from
2025-11 was saved while one from 2026-08 was not. That conclusion does not
follow. "Saved" means some earlier run downloaded it, not that ChatGPT
still holds it — an image captured in June looks saved forever after, even
if it died in July. A month with no losses shows the exports were timely,
not that the assets survived.

- probe_survival_by_month.py: sample images already on disk, grouped by
  their conversation's month, and ask /files/{id} whether each still
  exists today. Old months still 200 → no expiry, and cadence did not save
  them. Old months now 404 → uploads do expire and cadence is the whole
  ballgame.
- analyze_media_age.py: stop asserting the unsupported verdict; say what
  the number does and does not show, and point at the probe.

One signal there is immune to the confound and survives: 2026-07-09 kept
38 images and lost 12. Same conversation, same day, opposite outcomes —
no retention policy does that, so at least part of this is per-asset.
2026-08-17 08:53:54 -04:00
JesseMarkowitz 04191eed8c tools: analyze whether media loss is age-based
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
(![source](media/…) when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.

Offline, no token, no API calls.
2026-08-17 08:45:21 -04:00
JesseMarkowitz f40b25001a tools: add one-off probe for the media 403s
The improved error reporting landed, and the answer it produced is
{"detail":"Forbidden"} — generic, no reason. The cause has to be narrowed
by experiment instead, so collect the experiments in one script:

- classify the failed assets offline from the exported placeholders
  (user_upload vs model_generated) — no API call needed
- probe a known-good asset in the same session as a control, to rule the
  session in or out
- retry the 403 with ChatGPT-Account-Id, which the exporter never sends
  and which workspace-scoped resources can require
- compare /files/{id} against /files/{id}/download

Temporary: delete once the cause is known, or fold into `doctor` if the
check earns a permanent home.
2026-08-17 08:08:01 -04:00
JesseMarkowitz 395ea19ca8 fix: surface the response body on 4xx so 403s are diagnosable
Media downloads logged "HTTP Error 403:" with no reason. That string is
curl_cffi's raise_for_status() format, "HTTP Error {code}: {reason}", and
HTTP/2 carries no reason phrase — so the message said nothing, and
_make_request threw the response body away. The provider's JSON `detail`
is the only explanation available for a refused asset.

- base._make_request: end non-retryable statuses with a ProviderError
  carrying the body's detail/error/message (redacted, truncated to 300
  chars) instead of a bare raise_for_status().
- media: bucket 403 as `forbidden` in the run summary, separately from
  `download-error` — "the asset is gone" and "we were refused" are
  different problems.
- utils.redact_secrets: match secret key names per word. Exact matching
  let access_token, api_key, and session-token through into logged
  bodies; "keywords"/"monkey"/"tokenizer" stay intact.
- tests/test_config.py: test_defaults depended on the absence of a local
  .env — load_config() calls load_dotenv(override=False), which restored
  the variable the test had just deleted. Stub dotenv discovery.

305 tests pass.
2026-08-17 07:58:02 -04:00
28 changed files with 4377 additions and 602 deletions
+35
View File
@@ -38,6 +38,41 @@ CLAUDE_SESSION_KEY=
# list their names here (comma-separated).
#CLAUDE_CODE_REPO_TAG_IGNORE=some-repo,another-repo
# --- Codex (local agent sessions) ---
# The codex provider reads local Codex CLI rollout files. By default it scans
# ~/.codex/sessions/ (plus $CODEX_HOME/sessions when CODEX_HOME is set).
# To scan additional roots, set a ':'-separated list of sessions roots.
#CODEX_DIR=~/.codex/sessions:/mnt/backup/laptop/.codex/sessions
#
# As with Claude Code, session titles are tagged with the git repos they
# touched. To never tag specific repos, list their names here (comma-separated).
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
# --- Launcher ---
# Read by the ai-chat-exporter wrapper scripts, not by the Python code. The
# wrapper warns when run from outside the repo, because cache/ and exports/
# resolve against the current directory and the wrong one silently starts a
# separate archive. Set to 1 to silence that warning.
#AI_CHAT_EXPORTER_QUIET_CWD=1
# --- Notifications (ntfy) ---
# Push the result of a run to ntfy so an unattended archive reports back — the
# log file, the systemd journal and Task Scheduler's exit code are all pull-only.
# Unset NTFY_TOPIC disables notifications entirely.
#NTFY_TOPIC=my-archive-topic
#
# Self-hosting? Point at your own server.
#NTFY_SERVER=https://ntfy.sh
#
# Bearer token, for access-controlled topics. A topic on public ntfy.sh is
# readable by anyone who knows its name — notifications therefore carry counts
# and a machine name only, never conversation titles.
#NTFY_TOKEN=
#
# always (default) — notify on every run; failure — only when something failed;
# off — never.
#NTFY_NOTIFY=always
# --- Output ---
# Where exported Markdown files are written (default: ./exports)
EXPORT_DIR=./exports
+1
View File
@@ -22,6 +22,7 @@ exports/
!tests/fixtures/*.json
!README.md
!FUTURE.md
!FUTURE-ARCHIVE.md
!CHANGELOG.md
# Cache and logs
+44
View File
@@ -3,6 +3,50 @@
All notable changes to this project will be documented here.
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
## [Unreleased]
### Fixed
- **The test suite sent real push notifications to the developer's phone.** `TestSyncCommand` invokes the actual `sync` command, which calls `load_config()`, which calls `load_dotenv()` — so the real `.env` was loaded and its live `NTFY_TOPIC` used for the POST. Every `pytest` run fired three or four pushes, including a fabricated "codex: 3 conversation(s) failed to export" straight out of a fixture, which is worse than noise: it reports a failure that never happened. Nothing appeared in `cache/logs/exporter.log` to explain it, because every test invocation passes `--no-log-file`.
A `tests/conftest.py` autouse fixture now neutralises the environment for every test: `NTFY_TOPIC`/`NTFY_TOKEN` are emptied, and `NTFY_SERVER` and `JOPLIN_API_URL` are pointed at a closed local port, so a stray topic cannot reach the internet and a test cannot write notes into a real Joplin instance. The values are **emptied rather than deleted** — `load_dotenv(override=False)` skips only keys already present, so deleting one lets `.env` put it back. Verified by instrumenting `requests` across a full run: zero outbound requests, where the same instrumentation without the fixture records POSTs to the live ntfy topic.
## [0.9.0] - 2026-08-18
### Fixed
- **An em dash in a notification title silently dropped the notification.** HTTP header values are latin-1 at best and `requests` raises on anything outside it, so the first real send failed with `'latin-1' codec can't encode character '\u2014'`. Header values are now flattened to ASCII (smart punctuation mapped to its plain equivalent); the body is unaffected, being sent as UTF-8 bytes. Found by sending a test push rather than by reading the code.
- **The terms-of-service gate exited 0 without a terminal.** `click.prompt` raises `Abort` on a closed stdin, which the handler treated as a user Ctrl-C and exited 0 — so a scheduled run on a machine that had never acknowledged the notice would report success having archived nothing. Non-interactive invocations now exit 1 with an explanation of how to clear the gate once by hand. Found by running the new systemd unit rather than by reading the code.
- **A single U+0085 in a transcript silently dropped a whole record.** Both local providers split session files with `str.splitlines()`, which breaks not just on `\n` but on U+0085 (NEL), U+2028 and U+2029 — all of which are legal *inside* a JSON string and are written literally by Codex (Rust does not escape non-ASCII). One NEL in captured command output shredded one record into unparseable fragments; the parser logged "skipped 3 unparseable line(s)" and lost the record. Found while exporting a real rollout. Both providers now split on `\n` only, and both have regression tests that write their fixtures with `ensure_ascii=False` — with `json.dumps`' default the hazardous characters are escaped and the bug cannot reproduce.
- **Deleted uploads are no longer reported as permission errors.** ChatGPT's `/backend-api/files/{id}/download` answers a *missing* asset with `403 {"detail":"Forbidden"}`, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned `404 {"detail":"File not found"}` on `/files/{id}`, while assets that downloaded fine returned 200 on both in the same session, and `ChatGPT-Account-Id` made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as `expired-or-missing`. A 403 on an asset that *does* still exist is left alone as `forbidden` — that one would be a real problem.
- **4xx errors now report why.** `_make_request` ended non-retryable statuses with `raise_for_status()`, whose curl_cffi message is `HTTP Error {code}: {reason}` — and HTTP/2 carries no reason phrase, so a refused request logged as bare `HTTP Error 403:` and the response body (the only explanation the provider gives) was discarded. The body's `detail`/`error`/`message` is now carried into the `ProviderError`, redacted and truncated. This is what made the media 403s on `GET /backend-api/files/{id}/download` undiagnosable.
- **`redact_secrets` missed compound key names.** It matched keys exactly, so `access_token`, `api_key`, and `session-token` passed through un-redacted into debug-logged response bodies; matching now applies per word ("keywords", "monkey", "tokenizer" stay intact).
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
### Added
- **ntfy push notifications for unattended runs (`NTFY_TOPIC`).** A scheduled archive reports only to places you have to go and look at — the log file, the systemd journal, Task Scheduler's exit code — so a run that quietly failed every morning would stay quiet. `sync` now pushes its result: success carries per-provider counts at low/default priority, failure carries the reason at high priority with an alert tag, so the two are distinguishable at a glance on a phone. `ai-chat-exporter notify` shows the settings and `--test` sends a test push. `NTFY_NOTIFY=failure` limits it to failures, `off` disables, `--notify`/`--no-notify` override per run. Never fatal: an unreachable ntfy logs a warning and the run still reports its real exit code.
The payload is **counts and a machine name only, never conversation titles** — a topic on public ntfy.sh is readable by anyone who knows its name, and a title is the first line of what you asked. The machine name is included because several machines archive into one topic, where "2 exported" means nothing on its own. `NTFY_SERVER` points at a self-hosted instance and `NTFY_TOKEN` authenticates against an access-controlled topic.
- **`ai-chat-exporter` / `ai-chat-exporter.cmd` launchers — no virtualenv ceremony.** `cd` into the repo and run; the wrapper creates `.venv`, installs dependencies on first run, and reinstalls when `pyproject.toml` changes. A fresh clone goes from nothing to a working command in one step (measured: ~10s), on Linux/macOS and Windows alike, which matters for a tool meant to run on several machines. `cmd.exe` searches the current directory before `PATH`, so Windows needs no `.\` prefix. The working directory is deliberately not changed — `.env`, `cache/` and `exports/` still resolve against it, which is what lets one checkout archive different machines into different places — but the wrapper now warns when you run it from elsewhere, because a different `cache/manifest.json` silently starts a *second* archive rather than failing.
- **`sync` command — `export` then `joplin` in one invocation, with a real exit code.** Intended for schedulers (and the "trivial add-on" FUTURE.md §7 anticipated): it exits non-zero if any conversation failed to export or any note failed to sync, so a scheduled run that achieved nothing is distinguishable from one that had nothing to do. `--skip-joplin` exports only; `--joplin-optional` downgrades an unreachable Joplin to a warning, since the export has already captured the local transcripts and the notes rebuild from the cache on the next run that finds Joplin up.
- **Daily scheduling for both platforms.** `scheduling/install-systemd-timer.sh` (systemd user timer, `Persistent=true` so a machine that was off catches up at boot) and `scheduling/Register-AiChatSyncTask.ps1` (per-user Task Scheduler entry, `-StartWhenAvailable`). `--provider` is repeatable in both, because the right set differs per machine: the local providers need no credentials and always work unattended, while a web provider whose session token has expired would fail the job every single day and train you to ignore it.
- **Codex CLI provider (`--provider codex`).** Archives local Codex agent transcripts from `~/.codex/sessions/**/rollout-*.jsonl` — local-only, like `claude-code`: no tokens, no rate limits, no ToS exposure. Sessions land in their own top-level `AI-Codex` Joplin notebook, with the same prose-only default, repo tags (`CODEX_REPO_TAG_IGNORE`) and multi-root scanning (`CODEX_DIR`, plus `$CODEX_HOME/sessions`).
Codex writes each session twice in one file and the choice between the two layers is the whole design. `response_item` records are the model-facing wire format, where a tool call arrives as *JavaScript* (`tools.exec_command({...})`) because Codex's `exec` tool is code-mode; `event_msg`/`item_completed` records are Codex's own typed items, already decoded into `CommandExecution`/`FileChange`/`Extension` with argv, cwd, exit code and output as fields. Measured over 7 sessions on 0.147.0 (2026-08-18), the typed layer is 1:1 with the raw layer for prose (91 `AgentMessage` ↔ 91 assistant messages, sharing ids) and additionally omits every piece of harness plumbing — all 51 `developer`-role messages plus the 7 `# AGENTS.md instructions…` and 1 `<environment_context>` injections — which the Claude Code provider has to strip by regex. So the typed layer is parsed for content.
Its one gap is that it records what *ran*, not what was *attempted*: 26 of 180 `exec_command` calls produced no item (14 sandbox launch failures, 6 user aborts, ~5 still running at turn end, 1 failure). The raw layer is therefore read for a call count only, and placeholders report the shortfall — `3 calls: exec_command ×3 (+2 did not complete)` — instead of silently under-reporting. `wait` calls are process polls, not attempts, and are excluded.
**Reasoning is not exportable from Codex.** All 345 reasoning records carry `encrypted_content`, with `summary` empty in the raw layer and `summary_text`/`raw_content` empty in the typed layer, in every session. It is always dropped and counted; unlike Claude Code, `EXPORTER_HIDDEN_CONTENT=full` cannot surface it. The sidecar SQLite databases (`state_5.sqlite`, `thread_history_1.sqlite`) are deliberately not read: `thread_history_projection_state` tracks a byte offset into the rollout file, so the JSONL is canonical and the DB derived, and its `title` column is just the first user message truncated.
**Codex Cloud is out of scope, verified rather than assumed.** Cloud tasks are reachable at `chatgpt.com/backend-api/api/codex/tasks{,/list}` — the same host and `/backend-api` root the ChatGPT provider already uses — but local CLI sessions are never uploaded there, so it is not an alternative source for these transcripts. `codex cloud list` confirmed the account holds no cloud tasks. The provider makes no network calls.
- **`projects` command — discover the project IDs your config is missing.** `CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is invisible to the listing pass, so its conversations are never fetched. The command reports every project your conversations belong to, marks which are absent from `.env`, and prints a paste-ready line (`--write` applies it). It reads project ids from the conversation listing when they are there and falls back to `--deep`, one detail request per conversation, when they are not.
- **Project attribution now reads the conversation's own `gizmo_id`.** Previously the project name came only from `CHATGPT_PROJECT_IDS`, so a conversation in a project you had not listed exported into `no-project/` even though its payload names its project. The detail response carries `gizmo_id`, so it is used as a fallback after the listing annotation and the project map — attribution stays correct without maintaining a list, and moving a chat into a new project no longer silently misfiles it. Only `g-p-` ids count: a custom GPT is not a project and must not become a folder. Each unconfigured project is reported once per run, naming the id to add, because the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that live only inside a project never appear in the default listing.
### Changed
- Media download failures are bucketed as `forbidden` (403 — the file record survives) separately from `download-error`, so the run summary distinguishes it from `expired-or-missing` (404).
### Notes
- **ChatGPT media 403s: investigated and closed (2026-08-17).** 19 images across 7 conversations would not download. They are unrecoverable, and not because of anything the exporter or the export schedule did. Findings, recorded so this is not re-litigated: uploads do **not** expire (36/36 sampled images from 2025-09 through 2026-08 are still live, so export cadence is not a factor); the failures split into 7 records that 404 outright and 12 that report `state: "ready"` with a `library_file_id`; `/files/{id}/download` is the endpoint that mints the signed `estuary/content` URL the web UI fetches, and it refuses the survivors with a bare 403 regardless of headers, `Authorization`, `gizmo_id`, `conversation_id`, Referer, or namespace, while the Library id is rejected as `file_not_found`. Decisively, those same images render blank in ChatGPT's own UI — nothing is being withheld from the exporter. All the survivors were created 2026-07-14 within minutes of each other yet appear in conversations predating that date, pointing at a Library migration that kept the metadata and lost the bytes.
## [0.8.0] - 2026-07-06
Focused on the Claude Code provider after the tool's on-disk layout changed and
+660
View File
@@ -0,0 +1,660 @@
# FUTURE.md — archived 2026-08-18 (pre-v0.9.0 cleanup)
This is the full `FUTURE.md` as it stood before the v0.9.0 cleanup, kept
because it carries the investigation trails behind decisions that are now
recorded in one line each: why Brave cookie extraction is not viable, why the
ChatGPT token's expiry cannot be read client-side, how the drift canary was
designed, and the reasoning behind each closed backlog item.
Nothing here is planned work. The live roadmap is in `FUTURE.md`.
---
# Planned Future Work
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
>
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
> active roadmap is empty and the remaining backlog below has been **closed as
> not needed** — the tool does what it's needed to do as a local, manually-run
> backup CLI. Items are kept for reference only; revisit on demand if a real
> need shows up. Nothing here is planned work.
Items completed in each release are moved to the changelog. Items below the
roadmap were designed for but intentionally not implemented. The codebase is
structured to make each of these additions straightforward if ever revived.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
---
# Roadmap (decided 2026-06-12)
Priorities reflect the tool's primary purpose — a trustworthy backup so that
conversation data is not lost if a provider account is ever closed — plus the
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
local, manually-run CLI; the headless/StartOS direction was dropped
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
(manual refresh is sufficient).
**Now (in order):**
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
(full-archive `export --force` re-export completed 2026-06-13)
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
needs admin + SYSTEM impersonation that AV flags as credential theft, and
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
**Soon, but later:**
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
manually-run CLI, so no in-app polling loop is needed
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
sufficient (source convos live in the cloud and can be re-downloaded;
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
also retires the headless token-freshness problem — manual weekly refresh
is fine.
**Active:**
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
via the `/api/auth/session` `error` field (not `expires`). See §9.
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
`canary` command. See §10.
**Deprioritized** (entries kept at the bottom of this file; revisit on
demand): `--force` flags, per-conversation cache reset, official export-ZIP
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
scope — no significant usage to archive.
---
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, with one scoping correction
from live recon: the retrieval dumps are NOT flagged
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
the discriminator (the hidden flag only marks Custom Instructions and small
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
completed 2026-06-13.
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
the user's own attached documents, the same files dumped dozens of times
per conversation. These messages were invisible in the ChatGPT web UI;
they appear in exports because v0.4.0 lifted the role filter to fix silent
data loss. Custom Instructions hidden-context blocks are a minor secondary
case (~2KB, once per conversation) — the originally planned
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
file contains zero hidden-context blocks.
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
`placeholder`, plus a `--hidden-content` CLI override on `export`.
- `placeholder` renders affected messages as one line with type and size:
`> 🔧 Tool output (file_search, 24KB) — omitted
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
roughly halves; worst files shrink ~90%.
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
decision must happen in the provider, not the renderer); (b) messages
flagged `is_visually_hidden_from_conversation`, including
`user_editable_context` / `model_editable_context` (Custom
Instructions) — subsumes the old suppress-hidden-context idea.
- **Scope — keep at full size:** code-execution `tool_result` blocks and
web-search results; those are usually content the user wants.
- Count collapsed messages in the post-export summary so the omission
stays visible (mirror the LossReport presentation, but as intentional
policy, not loss).
Re-export workflow after shipping: `cache --clear` + `export` (same as
the v0.4.0 migration).
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
different mechanism (see below).
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
symptom hit in the field:
```
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
```
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
rights** and behavior that AV/EDR flags as infostealer activity. The one
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
this, so no off-disk reader is a reliable, non-invasive fit.
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
(launch the browser with `--remote-debugging-port`, read cookies via
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
hand. With the headless/StartOS direction dropped (#8), manual DevTools
refresh is the accepted approach — no automated extraction is needed.
## 3. Claude Code Session Provider — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
during implementation: `isSidechain` records are subagent transcripts (skipped),
`isMeta` marks harness-generated user records (skipped), and listing/normalized
`updated_at` must both use file mtime or the cache would re-export every
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
Archive local Claude Code session transcripts. No tokens, no rate limits,
no ToS risk — the data is already on disk but lives in a single JSONL per
session that Claude Code may clean up, and it contains deliverables
(reviews, plans, analyses) that exist nowhere else.
Decisions (2026-06-12):
- **Rendering: prose-only.** Keep user prompts and assistant text
(including full deliverable write-ups); collapse tool activity to
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
86KB — omitted`); **exclude thinking blocks**.
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
in v0.5.0).
Data facts (measured 2026-06-12):
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
Currently 29 sessions, 18.8MB total, largest 3.8MB.
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
activity is what makes these exports readable.
- Record types: `user` / `assistant` (Anthropic-style `message.content`
block arrays) plus harness records: `ai-title` (use for note title and
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
`permission-mode`, `system` (skip). Strip harness noise from user
messages (`<local-command-caveat>`, `<command-name>` blocks).
Implementation shape: new `src/providers/claude_code.py` implementing the
`BaseProvider` interface — `list_conversations` scans project dirs,
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
existing block schema (content is already block-shaped: text / tool_use /
tool_result / thinking). Incremental sync via file mtime/size recorded in
the existing manifest. Project name derives from the munged cwd dirname.
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
(refuses to prune right after `cache --clear`). First live run removed 420
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
manifest↔disk integrity (293/293 after the run).
A backup is only trustworthy if the on-disk tree matches the manifest.
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
v0.4.1) sit alongside current exports and would double-sync into Joplin.
- `prune` command: delete export files not referenced by the manifest.
`--dry-run` (default off, but always print the list before deleting)
shows what would be removed and why (old layout / orphan / unknown).
- `doctor` extension: verify every manifest entry's `file_path` exists on
disk; report missing files (re-export candidates) and unreferenced files
(prune candidates).
## 5. Binary Content Downloads — SHIPPED v0.6.0
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
`download_url`; a second GET yields the bytes (works for user uploads;
older AI-generated images 404 — expired server-side, handled gracefully).
Asset refs come in three shapes — `sediment://file_…`,
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
images-only is the default and audio is opt-in via `all`.
Original notes below.
**Priority note (2026-06-12): "later, but soon" — under the
backup-if-account-closes goal, embedded images are part of the data that
would be lost; placeholders alone don't preserve them.**
v0.4.0 ships placeholders for images and audio assets but does not download
the binary content. The `_safe_fence`-wrapped placeholders include the asset
reference (`sediment://...` or `file-service://...`), MIME type, size, and
duration where available; the actual bytes are not preserved.
Next steps:
- Download attached images alongside the Markdown export, save under a
`media/` sibling directory with a stable filename derived from the asset
reference.
- Replace `image_placeholder` rendering with an inline `![](relative/path)`
reference once the file is on disk.
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
rewrite the rendered Markdown to use `:/resourceId` references, and track
the resource ID in the cache manifest so re-syncs stay idempotent.
- DALL-E images on the assistant side: not observed in this user's data; the
code path exists (`source = "model_generated"`) but is untested.
The block-level schema is already in place — only the file-fetch + rewrite
layer needs to be added. See the `image_placeholder` and `file_placeholder`
block definitions in `src/blocks.py`.
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
the re-run picked it up.
Cap how many conversations are downloaded in a single `export` run so the
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
most importantly on the very first export, which otherwise fetches the
entire conversation history in one session. Because every run is resumable
(the manifest records each conversation immediately), a capped run simply
exports the first N pending conversations and the next run picks up where
it left off. This keeps traffic looking like a human-paced session rather
than a scraper, reducing the risk of rate limiting or account flags.
Two complementary pieces:
1. **Session cap** — `--max-conversations N` flag (and
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
`export` command, slice the pending list after the cache filter:
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
counts (reuse the message format from the 429 early-exit path) and
remind the user to re-run to continue.
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
random jitter) slept between per-conversation detail fetches in
`BaseProvider`, so even a capped run doesn't fire requests
back-to-back. The existing 429 backoff in `_request` stays as the
reactive safety net.
Note: the conversation *listing* (paginated, 100/page) still runs in full
each time so the cache comparison works — the cap applies to the heavy
per-conversation detail fetches, which dominate request volume.
This is a stepping stone to the StartOS service: a capped, politely-paced
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
traffic profile a headless deployment needs.
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
whatever hosts the tool: a user cron line locally, and on the long-term
StartOS target the platform's own scheduling. Either way the tool only needs
to do one capped, politely-paced `export` + `joplin` run and exit — which it
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
that chains `export` then `joplin` for a single cron line is a trivial
add-on, but the polling loop itself is off the roadmap.
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
Not pursuing a headless StartOS service. The local, manually-run CLI is
sufficient: the source conversations live in the providers' clouds and can be
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
so durability is covered without a server in the loop.
Dropping this also retires the one genuinely hard sub-problem it carried —
session-token freshness without a browser. There is no headless context to
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
extraction remains a dead end regardless — see #2.)
### REOPENED as a TODO (2026-08-18) — centralization, not durability
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
decision rested on "the source conversations live in the providers' clouds and
can be re-downloaded". **That is no longer true of half the providers.**
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
that exist *only* on the machine that produced them — Codex prunes its rollout
files, and neither is recoverable from any cloud. The re-download premise now
covers the web providers only.
The new motivation is consolidation rather than durability: work is split across
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
conversations from everywhere, instead of per-machine islands that only meet
inside Joplin.
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what it is uniquely able to do —
read that machine's local transcripts, and hold the browser session for the web
providers. What changes is where the output goes: instead of syncing to Joplin
itself, a local run **uploads its conversations to the StartOS storage area**,
and the StartOS service owns the Joplin connection for the whole corpus.
That inverts today's arrangement, where every machine talks to its own Joplin
desktop, and it removes two problems we already have:
- **The Joplin-availability race disappears from the clients.** A scheduled run
currently has to find Joplin desktop open on that same machine — the
2026-08-18 09:02 timer run exported fine and then skipped the sync because
Joplin did not start until 09:07. Uploading to a server that is always up has
no such window, and `--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note-update logic run once, server-side, against one manifest — rather than
each machine independently deciding what a notebook is called and racing to
update the same note.
What this would need, and what it would *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers can keep running interactively on the
machine that has the browser, pushing their output to the server. Only the
local providers need to run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`--upload`/`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter currently does not
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing would make that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should probably keep them rather than move them.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
Shipped: `doctor` now adds a "ChatGPT token active" check via
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
decode path was removed, and `_fetch_access_token` now fails fast on a set
`error` instead of returning a stale token. Tests in
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
below for the record.
Goal: if it's cheap to tell how much longer a token will work, show it on
`doctor`. Findings from live recon:
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
never yields an expiry for the real tokens (falls to the "not decodable"
branch).
- **`/api/auth/session` exposes an `expires`** (the provider already calls
this endpoint in `_fetch_access_token`; the response includes `expires`
alongside `accessToken`). Live value observed 2026-06-28:
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
assumption and the lived weekly-refresh cadence, so it is almost certainly
the NextAuth **rolling session window** (re-extended on every call), not
the point at which the pasted token actually 401s. Displaying it verbatim
would give false confidence.
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
was actually dead (conversations API → 401), `/api/auth/session` still
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
So `expires` is a **rolling session window that rolls forward on every call
even for a dead token**; displaying it would actively lie. The same response
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
token `error` is absent/null; when the session token is dead NextAuth can't
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
health check.
- **Design (ready to build):**
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
falsy and `accessToken` present; on `RefreshAccessTokenError` report
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
work) and do NOT surface `expires`.
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
without checking `error`, so it proceeds with a stale token and yields a
confusing downstream 401 instead of a clear "refresh your token" message.
Check `error` there and fail fast.
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
ChatGPT and Claude). It fetches one listing page + one conversation per
provider and asserts only the normalizer's load-bearing fields, emitting
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
trail below for the record.
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
change shape without notice. The worst failure for a backup tool is *silent*:
a response-schema change that makes the exporter skip or mis-parse content
without erroring. `doctor` currently checks token validity, reachability, and
manifest↔disk integrity — but not "does the provider's response still look
like what the parser expects."
To investigate: a lightweight schema/shape assertion on a known-good sample
of each provider's listing + conversation-detail responses (presence and type
of the fields the normalizers rely on), surfaced as a `doctor` check or a
dedicated canary.
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
- **Dependency surface (assert ONLY these — see below for why):**
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
`project.name`, `chat_messages[]`.
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
(list of typed blocks), `created_at`.
- **Key finding — full-shape diffing is the wrong design.** Claude's
`settings` object is full of volatile internal codenames that churn
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
canary would fire on every UI experiment. The canary MUST target the
normalizer's load-bearing fields only, not the whole response. (Aligns with
the drop-noise-don't-retain-it principle.)
- **Real Claude messages are flat `text`/`sender`** — in this archive every
message had a string `text` and NO `content` block list (0 rich blocks
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
thinking, image, …) is an **unexercised theoretical path**; drift there
can't be "caught" by a canary because it never runs on real data — it's a
safety net for if Claude ever switches to block content. The canary should
assert the flat shape and *warn if `content` ever appears as a list* (that
itself is the drift event that would activate the dormant code).
- **Possible silent-loss spot (separate from drift):** Claude messages carry
`attachments` and `files` arrays (empty in this sample) that the normalizer
ignores entirely. If a user ever attaches files in Claude, they'd be
dropped without a LossReport entry. Worth a follow-up check.
**Live recon — ChatGPT captured 2026-06-28:**
- **Dependency surface (assert ONLY these):**
- listing item: `id`, `title`, `update_time`/`create_time`.
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
`update_time`, `mapping` (non-empty).
- mapping node: `message`, `children` (the tree walk depends on both);
message: `author.role`, `author.name`, `content.content_type`,
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
- **content_type vocabulary observed (all currently handled):** `text`,
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
`execution_output`, `reasoning_recap`, `user_editable_context`,
`tether_browsing_display`. A *new* content_type already degrades gracefully
(visible `unknown` block + WARNING + LossReport tally) — so content_type
drift is **already non-silent**. The canary just needs to confirm the known
set still parses to non-empty blocks.
- **The genuinely silent drift risk — `author.name` collapse keys.**
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
confirms `file_search` is live (and `web.run`/`python` are correctly left
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
stops** and the archive re-bloats with no error or LossReport entry. This is
the top canary target: assert that retrieval-dump tool authors are still
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
- **Second silent risk — empty `content.parts`.** A `text` message whose
`parts` field is renamed/emptied yields zero blocks and is skipped with only
a debug log = silent loss. Canary should assert a sampled `text` message
produces a non-empty block.
**Recon complete for both providers. Canary design (ready to build):** a
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
listing page + one conversation and asserts the dependency-surface fields
above by presence+type — NOT full shape (Claude `settings` codenames prove
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
surface as a warning, never a hard error (a backup tool must still run).
---
# Deprioritized — CLOSED as not needed (2026-06-28)
These were considered and intentionally **not** built. Closed, not planned —
the tool is feature-complete for its purpose. Kept for reference in case a
real need ever revives one: Joplin `--force`, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
output, token-expiry notifications (also moot — see §9), and a search
command. (Also closed, outside this list: handling Claude `attachments`/
`files`, which the canary will flag if they ever appear in real data.)
## Export `--force` Flag — SHIPPED v0.6.0
Implemented 2026-06-12: `export --force` passes `force=True` to
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
preserves Joplin links across re-exports, so a forced re-render + `joplin`
updates existing notes instead of duplicating them.
## Joplin `--force` Flag
Similarly, add `--force` to the `joplin` command to re-sync all cached
conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
Implementation: in `get_joplin_pending()`, return all entries that have a
`file_path` when `force=True`, ignoring `joplin_synced_at`.
## Per-Conversation Cache Reset
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
single conversation without clearing the entire provider cache.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
## Official API Fallback
If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface
(`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON.
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
this user needs (project data), so it isn't a real fallback here.
## Reclassify o1/o3 Reasoning Subparts
v0.4.0 leaves dict parts inside `text` content_type messages with shape
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
shape was inferred from a code comment, not captured live). Once a real
reasoning conversation is captured, reclassify these as `thinking` blocks.
## Obsidian Vault Output
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
## Token Expiry Notifications
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
"proactively notify before expiry" idea is blocked by the same finding: a
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
window rather than the real refresh cadence). Any heads-up — a `doctor`
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
resolving the §9 open question of what signal is actually trustworthy.
## Search Command
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
## Split the README into Separate Documents
**TODO (2026-08-18).** The README is 830 lines / 5,562 words / 37 KB — about a
25-minute read, with 61 headings. An H2-only table of contents was added the
same day and helps navigation, but it treats the symptom: the file is doing at
least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
Not done yet, and not urgent, because it has a real cost the TOC does not: any
existing link into a README section (a bookmark, a note, another repo, a commit
message) breaks when that section moves to another file. Worth doing when the
README next needs substantial editing anyway, rather than as a change of its own.
Two things to decide when it happens:
- Whether `docs/` renders acceptably on the Gitea instance that hosts this repo
(relative links between Markdown files do work there, but worth confirming
before splitting rather than after).
- Whether the anchors in the split files stay stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors are derived from heading text, so they break silently on a reword —
the same fragility the TOC already carries.
+122 -535
View File
@@ -1,557 +1,144 @@
# Planned Future Work
> **Status 2026-07-06 (v0.8.0): Claude Code coverage reopened and shipped.**
> Claude Code changed its on-disk layout (subagent transcripts moved to separate
> `subagents/*.jsonl` files) and its sessions were hard to find in Joplin. v0.8.0
> addressed this: subagent capture (folded `<details>`), repo `[tags]` in titles,
> an own `AI-ClaudeCode` notebook with self-healing note moves, and multi-root
> scanning (`CLAUDE_CODE_DIR` list + `CLAUDE_CONFIG_DIR`). See the changelog.
>
> **Status 2026-06-28: feature-complete / done for now.** As of v0.7.0 the
> active roadmap is empty and the remaining backlog below has been **closed as
> not needed** — the tool does what it's needed to do as a local, manually-run
> backup CLI. Items are kept for reference only; revisit on demand if a real
> need shows up. Nothing here is planned work.
> **Status 2026-08-18 (v0.9.0).** The tool archives four providers — two web
> (`chatgpt`, `claude`) and two local agent-transcript (`claude-code`, `codex`)
> — on a schedule, on Linux and Windows, reporting results by push notification.
> Two items below are genuinely planned. Everything else has shipped or been
> decided against.
Items completed in each release are moved to the changelog. Items below the
roadmap were designed for but intentionally not implemented. The codebase is
structured to make each of these additions straightforward if ever revived.
**Completed:**
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- v0.4.0 — Rich content support: typed message blocks (text, code, thinking, tool_use, tool_result, image_placeholder, file_placeholder, unknown); ChatGPT voice transcripts as text + audio placeholders; Custom Instructions extraction; data-loss visibility via `LossReport` summary and visible `unknown` blocks
- v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- v0.6.0 — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT` policy; roadmap item 1); session limiter + request pacing (`MAX_CONVERSATIONS_PER_RUN`, `REQUEST_DELAY`; roadmap item 6)
Completed work moves to the changelog; this file holds only what is *not* built
yet. Decisions not to build something are recorded at the bottom in one line
each, so they are not re-proposed — the full investigation trails behind them
are in `FUTURE-ARCHIVE.md`.
---
# Roadmap (decided 2026-06-12)
# Roadmap
Priorities reflect the tool's primary purpose — a trustworthy backup so that
conversation data is not lost if a provider account is ever closed — plus the
day-to-day friction of the weekly ChatGPT token refresh. The tool stays a
local, manually-run CLI; the headless/StartOS direction was dropped
2026-06-28 (see #7 and #8), which also retires the token-freshness problem
(manual refresh is sufficient).
## 1. StartOS Service — one corpus, not per-machine islands
**Now (in order):**
1. ~~Collapse tool retrieval dumps & hidden context~~ — **shipped in v0.6.0**
(full-archive `export --force` re-export completed 2026-06-13)
2. ~~Brave cookie auto-extraction~~ — **shipped in v0.6.0, removed afterward.**
Not viable: modern Chromium App-Bound Encryption (Chrome 127+/current Brave)
needs admin + SYSTEM impersonation that AV flags as credential theft, and
fails on Brave specifically. Auth is manual (DevTools) — see entry below.
3. ~~Claude Code session provider~~ — **shipped in v0.6.0**
4. ~~Archive hygiene — `prune` + `doctor` integrity check~~ — **shipped in v0.6.0**
Each machine currently archives to its own `exports/` and its own Joplin. Work
is split across boxes — coding sessions (`claude-code`, `codex`) on the Linux
machine, web chats (`chatgpt`, `claude`) on the Windows one — so there is no
single place where all conversations exist together.
**Soon, but later:**
This was dropped on 2026-06-28 and **reopened 2026-08-18**, because the
reasoning behind the drop has gone stale. It rested on "the source conversations
live in the providers' clouds and can be re-downloaded". That is no longer true
of half the providers: `claude-code` and `codex` transcripts exist *only* on the
machine that produced them, Codex prunes its rollout files, and neither is
recoverable from any cloud. The motivation is also different from the one
weighed then — consolidation, not durability.
5. ~~Binary content downloads~~ — **shipped in v0.6.0**
6. ~~Per-session download limiter + polite pacing~~ — **shipped in v0.6.0**
7. ~~Scheduled / watch mode~~ — **dropped 2026-06-28**; the tool stays a
manually-run CLI, so no in-app polling loop is needed
8. ~~StartOS service packaging~~ — **dropped 2026-06-28**; the local CLI is
sufficient (source convos live in the cloud and can be re-downloaded;
Joplin already syncs encrypted to an offsite S3 provider). Dropping this
also retires the headless token-freshness problem — manual weekly refresh
is fine.
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what only it can do: read that
machine's local transcripts, and hold the browser session for the web providers.
What changes is where the output goes. Instead of syncing to Joplin itself, a
local run **uploads its conversations to the StartOS storage area**, and the
StartOS service owns the Joplin connection for the whole corpus.
**Active:**
9. ~~Surface remaining token validity on `doctor`~~ — **IMPLEMENTED 2026-06-28**
via the `/api/auth/session` `error` field (not `expires`). See §9.
10. ~~Provider API-drift detection~~ — **IMPLEMENTED 2026-06-28** as the
`canary` command. See §10.
That inverts today's arrangement and removes two problems we already have:
**Deprioritized** (entries kept at the bottom of this file; revisit on
demand): `--force` flags, per-conversation cache reset, official export-ZIP
fallback, o1/o3 reasoning reclassification, Obsidian output, search command.
Additional web providers (Gemini/Grok/Perplexity) are explicitly out of
scope — no significant usage to archive.
- **The Joplin-availability race leaves the clients.** A scheduled run currently
has to find Joplin desktop open on that same machine — the 09:02 timer run on
2026-08-18 exported fine and then skipped the sync because Joplin did not
start until 09:07. A server that is always up has no such window, and
`--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note updates run once, server-side, against one manifest — rather than each
machine independently deciding what a notebook is called and racing to update
the same note.
What this needs, and what it does *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers keep running interactively on the
machine that has the browser, and push their output to the server. Only the
local providers would run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter does not track today:
`claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing makes that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should keep them rather than hand them off.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 2. Split the README into separate documents
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61
headings. An H2-only table of contents was added 2026-08-18 and helps, but it
treats the symptom: the file is doing at least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
Not urgent, because it has a real cost the TOC did not: any existing link into a
README section — a bookmark, a note, another repo, a commit message — breaks
when that section moves to another file. Worth folding into the next substantial
README edit rather than doing as a change of its own.
Two things to settle when it happens:
- **`.gitignore` ignores `*.md` on purpose** — exported conversations are
Markdown and may contain private content — and re-includes each doc by name
(`!README.md`, `!FUTURE.md`, …). New files under `docs/` will be silently
ignored, with no error, until `!docs/*.md` is added. This already bit the
creation of `FUTURE-ARCHIVE.md` on 2026-08-18.
- Whether `docs/` renders acceptably on the Gitea instance hosting this repo.
Relative links between Markdown files do work there, but confirm before
splitting rather than after.
- Whether anchors in the split files are stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors derive from heading text and break silently on a reword — the same
fragility the TOC already carries.
---
## 1. Collapse Tool Retrieval Dumps & Hidden Context — SHIPPED v0.6.0
# Decided against
**Implemented 2026-06-12** as designed below, with one scoping correction
from live recon: the retrieval dumps are NOT flagged
`is_visually_hidden_from_conversation` — `author.name == "file_search"` is
the discriminator (the hidden flag only marks Custom Instructions and small
system stubs). Verified live: worst files shrink 93% (524KB → 36KB).
Full-archive re-export (the `export --force` campaign) + Joplin re-sync
completed 2026-06-13.
Recorded so they are not re-proposed. Full reasoning and the recon behind each
is in `FUTURE-ARCHIVE.md`.
**Problem (measured 2026-06-12 against a full fresh export):** 45% of the
entire 11.2MB archive (260 files) is tool-role messages; 29 files are
majority tool-dump. Worst case: a 524KB conversation where 495KB (94%, 64
messages) is ChatGPT's file-retrieval tool re-injecting the full text of
the user's own attached documents, the same files dumped dozens of times
per conversation. These messages were invisible in the ChatGPT web UI;
they appear in exports because v0.4.0 lifted the role filter to fix silent
data loss. Custom Instructions hidden-context blocks are a minor secondary
case (~2KB, once per conversation) — the originally planned
`EXPORTER_INCLUDE_HIDDEN_CONTEXT` toggle alone would not help: the worst
file contains zero hidden-context blocks.
**Fix: collapse, don't drop** (consistent with the no-silent-drop rule):
- `EXPORTER_HIDDEN_CONTENT=full|placeholder|omit` env var, default
`placeholder`, plus a `--hidden-content` CLI override on `export`.
- `placeholder` renders affected messages as one line with type and size:
`> 🔧 Tool output (file_search, 24KB) — omitted
(EXPORTER_HIDDEN_CONTENT=full to keep)`. Expected effect: archive
roughly halves; worst files shrink ~90%.
- **Scope — collapse:** (a) tool-role retrieval dumps, identified by raw
`author.name` (`file_search`, `myfiles_browser`, …) in the API response
(the rendered Markdown only shows a generic "🔧 Tool" label, so the
decision must happen in the provider, not the renderer); (b) messages
flagged `is_visually_hidden_from_conversation`, including
`user_editable_context` / `model_editable_context` (Custom
Instructions) — subsumes the old suppress-hidden-context idea.
- **Scope — keep at full size:** code-execution `tool_result` blocks and
web-search results; those are usually content the user wants.
- Count collapsed messages in the post-export summary so the omission
stays visible (mirror the LossReport presentation, but as intentional
policy, not loss).
Re-export workflow after shipping: `cache --clear` + `export` (same as
the v0.4.0 migration).
## 2. Brave Cookie Auto-Extraction — REMOVED (not viable)
**Shipped v0.6.0 (2026-06-12), removed 2026-06-27.** `auth --from-browser`
plus `src/browser_tokens.py` and the `browser-cookie3` dependency are gone.
Auth is manual (DevTools) only. Do not re-attempt without a fundamentally
different mechanism (see below).
**Why it doesn't work.** Modern Chromium browsers encrypt cookies on Windows
with **App-Bound Encryption** (Chrome 127+, July 2024; current Brave).
Cookies are written with a `v20` prefix and keyed off a secret wrapped in a
**SYSTEM-level** DPAPI layer plus app validation. `browser-cookie3` only
knows the legacy `v10`/DPAPI key, so its AES-GCM MAC check fails — the exact
symptom hit in the field:
```
ChatGPT: Could not read brave cookies for chatgpt.com: Unable to get key for cookie decryption.
```
Decrypting `v20` at all requires unwrapping the SYSTEM layer, which means
running as SYSTEM (e.g. a PsExec-style service) — i.e. **Administrator
rights** and behavior that AV/EDR flags as infostealer activity. The one
maintained Python option (`rookiepy`) needs admin from Chrome v130+, was
**archived 2026-06-07**, and has an unresolved bug where **Brave returns 0
cookies** even after the key is retrieved. ABE is *designed* to stop exactly
this, so no off-disk reader is a reliable, non-invasive fit.
**If ever revisited:** the only non-admin path is Chrome Remote Debugging
(launch the browser with `--remote-debugging-port`, read cookies via
`Network.getAllCookies` — the running browser decrypts for you). Heavier and
intrusive; not worth it for a weekly token refresh that takes 30 seconds by
hand. With the headless/StartOS direction dropped (#8), manual DevTools
refresh is the accepted approach — no automated extraction is needed.
## 3. Claude Code Session Provider — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below (`src/providers/claude_code.py`,
`--provider claude-code`, `CLAUDE_CODE_DIR` override). Additional findings
during implementation: `isSidechain` records are subagent transcripts (skipped),
`isMeta` marks harness-generated user records (skipped), and listing/normalized
`updated_at` must both use file mtime or the cache would re-export every
session every run. First export: 25 sessions, 19MB JSONL → 804KB Markdown.
Archive local Claude Code session transcripts. No tokens, no rate limits,
no ToS risk — the data is already on disk but lives in a single JSONL per
session that Claude Code may clean up, and it contains deliverables
(reviews, plans, analyses) that exist nowhere else.
Decisions (2026-06-12):
- **Rendering: prose-only.** Keep user prompts and assistant text
(including full deliverable write-ups); collapse tool activity to
one-line placeholders (`> 🔧 Tool activity — 14 calls (Read ×9, Bash ×2),
86KB — omitted`); **exclude thinking blocks**.
- **Joplin: sync enabled.** Each coding project becomes a notebook nested
under an **"AI-Claude"** parent notebook (nested-notebook support shipped
in v0.5.0).
Data facts (measured 2026-06-12):
- Source: `~/.claude/projects/<munged-cwd>/<session-uuid>.jsonl`.
Currently 29 sessions, 18.8MB total, largest 3.8MB.
- Representative 2.1MB session: tool_result 436KB, tool_use 122KB,
thinking 104KB, dialogue prose only ~24KB (~4%) — collapsing tool
activity is what makes these exports readable.
- Record types: `user` / `assistant` (Anthropic-style `message.content`
block arrays) plus harness records: `ai-title` (use for note title and
filename slug), `last-prompt`, `file-history-snapshot`, `attachment`,
`permission-mode`, `system` (skip). Strip harness noise from user
messages (`<local-command-caveat>`, `<command-name>` blocks).
Implementation shape: new `src/providers/claude_code.py` implementing the
`BaseProvider` interface — `list_conversations` scans project dirs,
`get_conversation` parses the JSONL, `normalize_conversation` maps onto the
existing block schema (content is already block-shaped: text / tool_use /
tool_result / thinking). Incremental sync via file mtime/size recorded in
the existing manifest. Project name derives from the munged cwd dirname.
## 4. Archive Hygiene: `prune` Command + Manifest Integrity — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below, plus an empty-manifest guard
(refuses to prune right after `cache --clear`). First live run removed 420
stale files (9.4 MB, old-layout trees + `_.md` orphans); doctor now reports
manifest↔disk integrity (293/293 after the run).
A backup is only trustworthy if the on-disk tree matches the manifest.
Observed 2026-06-12: pre-v0.5.0 layout trees (`tspc-expertcouncil/2025/`)
and `_.md` no-ID orphans (from the empty-conversation-id bug fixed in
v0.4.1) sit alongside current exports and would double-sync into Joplin.
- `prune` command: delete export files not referenced by the manifest.
`--dry-run` (default off, but always print the list before deleting)
shows what would be removed and why (old layout / orphan / unknown).
- `doctor` extension: verify every manifest entry's `file_path` exists on
disk; report missing files (re-export candidates) and unreferenced files
(prune candidates).
## 5. Binary Content Downloads — SHIPPED v0.6.0
**Implemented 2026-06-12** (`src/media.py`, `EXPORTER_DOWNLOAD_MEDIA`,
`download_asset`/`parse_asset_file_id` on the ChatGPT provider, Joplin
`create_resource` + `upload_media_and_rewrite`). Live recon settled the
download mechanism: `GET /backend-api/files/{id}/download` returns a signed
`download_url`; a second GET yields the bytes (works for user uploads;
older AI-generated images 404 — expired server-side, handled gracefully).
Asset refs come in three shapes — `sediment://file_…`,
`sediment://<hash>#file_…#p_N.png` (generated), `file-service://…`. Archive
scan: 14 images (4 uploads / 10 generated) + 556 audio clips ≈162MB, so
images-only is the default and audio is opt-in via `all`.
Original notes below.
**Priority note (2026-06-12): "later, but soon" — under the
backup-if-account-closes goal, embedded images are part of the data that
would be lost; placeholders alone don't preserve them.**
v0.4.0 ships placeholders for images and audio assets but does not download
the binary content. The `_safe_fence`-wrapped placeholders include the asset
reference (`sediment://...` or `file-service://...`), MIME type, size, and
duration where available; the actual bytes are not preserved.
Next steps:
- Download attached images alongside the Markdown export, save under a
`media/` sibling directory with a stable filename derived from the asset
reference.
- Replace `image_placeholder` rendering with an inline `![](relative/path)`
reference once the file is on disk.
- Joplin integration: upload binaries as Joplin resources via `POST /resources`,
rewrite the rendered Markdown to use `:/resourceId` references, and track
the resource ID in the cache manifest so re-syncs stay idempotent.
- DALL-E images on the assistant side: not observed in this user's data; the
code path exists (`source = "model_generated"`) but is untested.
The block-level schema is already in place — only the file-fetch + rewrite
layer needs to be added. See the `image_placeholder` and `file_placeholder`
block definitions in `src/blocks.py`.
## 6. Per-Session Download Limiter — SHIPPED v0.6.0
**Implemented 2026-06-12** as designed below: `--max-conversations N` /
`MAX_CONVERSATIONS_PER_RUN` session cap with deferred-count reporting, and
`REQUEST_DELAY` pacing (default 1.0s ±25% jitter) in `BaseProvider._request`.
Verified live: a 2-pending run with cap 1 exported one, deferred one, and
the re-run picked it up.
Cap how many conversations are downloaded in a single `export` run so the
tool never hammers the ChatGPT/Claude internal APIs with a large burst —
most importantly on the very first export, which otherwise fetches the
entire conversation history in one session. Because every run is resumable
(the manifest records each conversation immediately), a capped run simply
exports the first N pending conversations and the next run picks up where
it left off. This keeps traffic looking like a human-paced session rather
than a scraper, reducing the risk of rate limiting or account flags.
Two complementary pieces:
1. **Session cap** — `--max-conversations N` flag (and
`MAX_CONVERSATIONS_PER_RUN` env default). Implementation: in the
`export` command, slice the pending list after the cache filter:
`to_export = to_export[:n]`. On exit, print exported-vs-remaining
counts (reuse the message format from the 429 early-exit path) and
remind the user to re-run to continue.
2. **Polite pacing** — `REQUEST_DELAY` env var (seconds, with small
random jitter) slept between per-conversation detail fetches in
`BaseProvider`, so even a capped run doesn't fire requests
back-to-back. The existing 429 backoff in `_request` stays as the
reactive safety net.
Note: the conversation *listing* (paginated, 100/page) still runs in full
each time so the cache comparison works — the cap applies to the heavy
per-conversation detail fetches, which dominate request volume.
This is a stepping stone to the StartOS service: a capped, politely-paced
export — scheduled by the host (cron/StartOS), not an in-app loop — is the
traffic profile a headless deployment needs.
## 7. Scheduled / Watch Mode — DROPPED (2026-06-28)
An in-app `watch`/scheduler loop is not worth building. Scheduling belongs to
whatever hosts the tool: a user cron line locally, and on the long-term
StartOS target the platform's own scheduling. Either way the tool only needs
to do one capped, politely-paced `export` + `joplin` run and exit — which it
already does. If cron ergonomics ever feel clunky, a thin `sync` subcommand
that chains `export` then `joplin` for a single cron line is a trivial
add-on, but the polling loop itself is off the roadmap.
## 8. StartOS Service Packaging — DROPPED (2026-06-28)
Not pursuing a headless StartOS service. The local, manually-run CLI is
sufficient: the source conversations live in the providers' clouds and can be
re-downloaded, and Joplin already syncs (encrypted) to an offsite S3 provider,
so durability is covered without a server in the loop.
Dropping this also retires the one genuinely hard sub-problem it carried —
session-token freshness without a browser. There is no headless context to
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
extraction remains a dead end regardless — see #2.)
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
Shipped: `doctor` now adds a "ChatGPT token active" check via
`ChatGPTProvider.session_health()` (reads `/api/auth/session`, passes iff
`error` is falsy and `accessToken` is present), the never-working JWE/`exp`
decode path was removed, and `_fetch_access_token` now fails fast on a set
`error` instead of returning a stale token. Tests in
`tests/test_providers.py::TestChatGPTSessionHealth`. Investigation trail
below for the record.
Goal: if it's cheap to tell how much longer a token will work, show it on
`doctor`. Findings from live recon:
- **Not readable from the token itself.** ChatGPT's `CHATGPT_SESSION_TOKEN`
is a **JWE** (header `{"alg":"dir","enc":"A256GCM"}`, `eyJ…` prefix is just
the encrypted protected header) — the `exp` claim is AES-256-GCM encrypted
with an OpenAI-only key, so it cannot be decoded client-side. Claude's
`sk-…` key is fully opaque. The existing `doctor` JWT-decode path therefore
never yields an expiry for the real tokens (falls to the "not decodable"
branch).
- **`/api/auth/session` exposes an `expires`** (the provider already calls
this endpoint in `_fetch_access_token`; the response includes `expires`
alongside `accessToken`). Live value observed 2026-06-28:
`2026-09-26` — **~90 days out**. This contradicts both the code's ~7-day
assumption and the lived weekly-refresh cadence, so it is almost certainly
the NextAuth **rolling session window** (re-extended on every call), not
the point at which the pasted token actually 401s. Displaying it verbatim
would give false confidence.
- **RESOLVED 2026-06-28 by a live 401 data point.** When the ChatGPT token
was actually dead (conversations API → 401), `/api/auth/session` still
returned **HTTP 200** with `expires: 2026-09-26` (~90 days out) — and that
`expires` *advanced* between two calls seconds apart (`04:20:08` → `04:30:37`).
So `expires` is a **rolling session window that rolls forward on every call
even for a dead token**; displaying it would actively lie. The same response
carried `error: "RefreshAccessTokenError"` and a stale `accessToken`.
- **The real signal is `error`, not `expires` or `accessToken`.** On a healthy
token `error` is absent/null; when the session token is dead NextAuth can't
refresh and sets `error: "RefreshAccessTokenError"` while still echoing a
rolling `expires` and a stale `accessToken`. This is an exact, free, binary
health check.
- **Design (ready to build):**
1. `doctor` ChatGPT check → read `/api/auth/session`; pass iff `error` is
falsy and `accessToken` present; on `RefreshAccessTokenError` report
"token expired — refresh". Drop the JWE/`exp` decode path (it can never
work) and do NOT surface `expires`.
2. Latent bug to fix alongside: `_fetch_access_token` reads `accessToken`
without checking `error`, so it proceeds with a stale token and yields a
confusing downstream 401 instead of a clear "refresh your token" message.
Check `error` there and fail fast.
3. Claude stays a 401-only signal (opaque `sk-`, no equivalent endpoint).
## 10. Provider API-Drift Detection — IMPLEMENTED (2026-06-28)
Shipped: the `canary` command + `BaseProvider.check_drift()` (overridden by
ChatGPT and Claude). It fetches one listing page + one conversation per
provider and asserts only the normalizer's load-bearing fields, emitting
`DRIFT_OK/WARN/ERROR` findings (`src/providers/base.py`). Severity badges
print as a Rich table; ERROR exits non-zero, WARN is non-fatal so a backup
run is never blocked. Drift vocabularies (`_KNOWN_TOOL_AUTHORS`,
`_HANDLED_CONTENT_TYPES`) live in `chatgpt.py` next to the collapse set they
guard. Tests: `TestChatGPTDriftCanary`, `TestClaudeDriftCanary`,
`TestCanaryCommand`. Verified live 2026-06-28 — both providers OK. Recon
trail below for the record.
## 10b. Provider API-Drift Detection — investigation (2026-06-28)
The export depends on undocumented internal web APIs (ChatGPT/Claude) that can
change shape without notice. The worst failure for a backup tool is *silent*:
a response-schema change that makes the exporter skip or mis-parse content
without erroring. `doctor` currently checks token validity, reachability, and
manifest↔disk integrity — but not "does the provider's response still look
like what the parser expects."
To investigate: a lightweight schema/shape assertion on a known-good sample
of each provider's listing + conversation-detail responses (presence and type
of the fields the normalizers rely on), surfaced as a `doctor` check or a
dedicated canary.
**Live recon — Claude captured 2026-06-28 (ChatGPT pending a token refresh):**
- **Dependency surface (assert ONLY these — see below for why):**
- listing item: `uuid`, `name`, `updated_at`/`created_at`, `project.name`.
- conversation detail: `uuid`/`id`, `name`, `created_at`, `updated_at`,
`project.name`, `chat_messages[]`.
- message: `sender` (`human`/`assistant`), `text` (string) or `content`
(list of typed blocks), `created_at`.
- **Key finding — full-shape diffing is the wrong design.** Claude's
`settings` object is full of volatile internal codenames that churn
constantly: `enabled_bananagrams`, `enabled_sourdough`, `enabled_foccacia`,
`enabled_saffron`, `enabled_turmeric`, `enabled_monkeys_in_a_barrel`,
`paprika_mode`, `enabled_megaminds`, … A "any new/removed key = drift"
canary would fire on every UI experiment. The canary MUST target the
normalizer's load-bearing fields only, not the whole response. (Aligns with
the drop-noise-don't-retain-it principle.)
- **Real Claude messages are flat `text`/`sender`** — in this archive every
message had a string `text` and NO `content` block list (0 rich blocks
observed). So `_extract_claude_blocks` / `_dispatch_claude_block` (tool_use,
thinking, image, …) is an **unexercised theoretical path**; drift there
can't be "caught" by a canary because it never runs on real data — it's a
safety net for if Claude ever switches to block content. The canary should
assert the flat shape and *warn if `content` ever appears as a list* (that
itself is the drift event that would activate the dormant code).
- **Possible silent-loss spot (separate from drift):** Claude messages carry
`attachments` and `files` arrays (empty in this sample) that the normalizer
ignores entirely. If a user ever attaches files in Claude, they'd be
dropped without a LossReport entry. Worth a follow-up check.
**Live recon — ChatGPT captured 2026-06-28:**
- **Dependency surface (assert ONLY these):**
- listing item: `id`, `title`, `update_time`/`create_time`.
- conversation detail: `conversation_id`/`id`, `title`, `create_time`,
`update_time`, `mapping` (non-empty).
- mapping node: `message`, `children` (the tree walk depends on both);
message: `author.role`, `author.name`, `content.content_type`,
`content.parts`, `metadata.is_visually_hidden_from_conversation`.
- **content_type vocabulary observed (all currently handled):** `text`,
`model_editable_context`, `multimodal_text`, `thoughts`, `code`,
`execution_output`, `reasoning_recap`, `user_editable_context`,
`tether_browsing_display`. A *new* content_type already degrades gracefully
(visible `unknown` block + WARNING + LossReport tally) — so content_type
drift is **already non-silent**. The canary just needs to confirm the known
set still parses to non-empty blocks.
- **The genuinely silent drift risk — `author.name` collapse keys.**
`_COLLAPSE_TOOL_AUTHORS = {"file_search", "myfiles_browser"}`. Recon
confirms `file_search` is live (and `web.run`/`python` are correctly left
un-collapsed). If OpenAI renames `file_search`, the collapse **silently
stops** and the archive re-bloats with no error or LossReport entry. This is
the top canary target: assert that retrieval-dump tool authors are still
recognized, or at least flag unfamiliar `(role="tool", author.name)` pairs.
- **Second silent risk — empty `content.parts`.** A `text` message whose
`parts` field is renamed/emptied yields zero blocks and is skipped with only
a debug log = silent loss. Canary should assert a sampled `text` message
produces a non-empty block.
**Recon complete for both providers. Canary design (ready to build):** a
`doctor` check (or dedicated `canary` command) that, per provider, fetches one
listing page + one conversation and asserts the dependency-surface fields
above by presence+type — NOT full shape (Claude `settings` codenames prove
full-shape diffing is pure noise). Specific tripwires: (ChatGPT) unfamiliar
`(tool, author.name)` pair and empty `parts` on a text message; (Claude)
`content` appearing as a list, and non-empty `attachments`/`files`. Failures
surface as a warning, never a hard error (a backup tool must still run).
| Item | Verdict |
|------|---------|
| Brave/Chromium cookie auto-extraction | **Not viable** (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
| In-app watch/polling loop | **Dropped** (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as `sync` plus a systemd timer and a Windows scheduled task. |
| Proactive token-expiry notification | **Blocked, and moot** (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and `/api/auth/session`'s `expires` is a rolling window that advances even for a dead token. `doctor`'s `error`-based health check covers the real need. |
| Official export-ZIP fallback | **Closed** (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. `BaseProvider` still admits a `FileProvider` if that changes. |
| Joplin `--force` flag | Closed as not needed (2026-06-28). |
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
| Search command | Closed as not needed (2026-06-28). `grep`/`ripgrep` over `EXPORT_DIR` covers it. |
| Claude `attachments` / `files` handling | Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
---
# Deprioritized — CLOSED as not needed (2026-06-28)
# Completed
These were considered and intentionally **not** built. Closed, not planned —
the tool is feature-complete for its purpose. Kept for reference in case a
real need ever revives one: Joplin `--force`, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reasoning reclassification, Obsidian
output, token-expiry notifications (also moot — see §9), and a search
command. (Also closed, outside this list: handling Claude `attachments`/
`files`, which the canary will flag if they ever appear in real data.)
Detail for each release is in `CHANGELOG.md`.
## Export `--force` Flag — SHIPPED v0.6.0
Implemented 2026-06-12: `export --force` passes `force=True` to
`cache.get_new_or_updated()`. Shipped alongside a `mark_exported` fix that
preserves Joplin links across re-exports, so a forced re-render + `joplin`
updates existing notes instead of duplicating them.
## Joplin `--force` Flag
Similarly, add `--force` to the `joplin` command to re-sync all cached
conversations to Joplin regardless of whether they've been synced before.
Useful after making formatting changes to the Markdown exporter.
Implementation: in `get_joplin_pending()`, return all entries that have a
`file_path` when `force=True`, ignoring `joplin_synced_at`.
## Per-Conversation Cache Reset
Add `cache --reset --conversation <id>` to force re-export or re-sync of a
single conversation without clearing the entire provider cache.
Current workaround: manually edit `~/.ai-chat-exporter/manifest.json` and
delete the entry, then re-run export.
## Official API Fallback
If the unofficial internal web API approach breaks, migrate to official export
file parsing as a fallback:
- ChatGPT: parse `conversations.json` from Settings → Export Data
- Claude: parse `conversations.json` from Settings → Privacy → Export Data
The `BaseProvider` abstract class is intentionally designed so that a
`FileProvider` subclass can implement the same interface
(`list_conversations`, `get_conversation`, `normalize_conversation`)
without any changes to cache, exporters, or CLI code.
To add this: implement `src/providers/file_chatgpt.py` and
`src/providers/file_claude.py`, then add `--input-file` flag to the
export command to accept a pre-downloaded export ZIP or JSON.
Deprioritized 2026-06-12: the official ChatGPT export does not cover what
this user needs (project data), so it isn't a real fallback here.
## Reclassify o1/o3 Reasoning Subparts
v0.4.0 leaves dict parts inside `text` content_type messages with shape
`{"summary": ..., "content": ...}` rendered as plain text (defensive — the
shape was inferred from a code comment, not captured live). Once a real
reasoning conversation is captured, reclassify these as `thinking` blocks.
## Obsidian Vault Output
Add an `obsidian` command (or `--target obsidian` flag) to sync exported
conversations into an Obsidian vault directory. The current Markdown format
is already largely compatible; the main differences are:
- Obsidian uses YAML frontmatter `properties` (same format, already supported)
- Tags should use `#tag` inline or `tags:` list in frontmatter (already done)
- Wikilinks (`[[Title]]`) instead of Markdown links — optional, Obsidian
supports both
Implementation: the existing `MarkdownExporter` output is already valid in
Obsidian. An `ObsidianSyncer` class (mirroring `JoplinClient`) would simply
copy files to the vault directory and maintain a flat or nested folder
structure matching the user's Obsidian setup. No API needed — just file I/O.
## Token Expiry Notifications
Moved to the active roadmap as §9 (Token Validity on `doctor`). The original
"proactively notify before expiry" idea is blocked by the same finding: a
reliable expiry time isn't available client-side (ChatGPT token is encrypted,
Claude's is opaque, and `/api/auth/session`'s `expires` looks like a rolling
window rather than the real refresh cadence). Any heads-up — a `doctor`
line, an `expiry` subcommand, or a `notify-send` nudge — depends on first
resolving the §9 open question of what signal is actually trustworthy.
## Search Command
Add a `search` command to full-text search across all exported Markdown files:
```bash
python -m src.main search "kubernetes ingress"
python -m src.main search "kubernetes ingress" --provider claude --project devops
```
Implementation: `grep`/`ripgrep` over `EXPORT_DIR`, display results with
conversation title, date, and a snippet. No index needed — Markdown files are
small enough to grep directly.
- **v0.1.0** — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- **v0.2.0** — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- **v0.4.0** — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via `LossReport` and visible `unknown` blocks
- **v0.5.0** — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- **v0.6.0** — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT`); Claude Code session provider; `prune` + manifest integrity; binary content downloads; session limiter and request pacing; `export --force`
- **v0.7.0** — `canary` drift detection; real ChatGPT token-health check on `doctor`; removal of the non-viable browser-cookie path
- **v0.8.0** — Claude Code coverage reopened: subagent capture as folded `<details>`, repo `[tags]` in titles, its own `AI-ClaudeCode` notebook with self-healing note moves, multi-root scanning (`CLAUDE_CODE_DIR`, `CLAUDE_CONFIG_DIR`)
- **v0.9.0** — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms; `sync` with a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution via `gizmo_id` and the `projects` command; a silent data-loss fix in both local providers (`splitlines` breaking on U+0085/U+2028/U+2029)
+324 -12
View File
@@ -6,7 +6,28 @@ Supports incremental sync — only new or updated conversations are exported on
---
## ⚠️ Terms of Service Warning
## Contents
- [Terms of Service Warning](#terms-of-service-warning)
- [Installation](#installation)
- [First Run: Run Doctor](#first-run-run-doctor)
- [Getting Your Session Tokens](#getting-your-session-tokens)
- [The `auth` Command](#the-auth-command)
- [`.env` Setup](#env-setup)
- [ChatGPT Projects](#chatgpt-projects)
- [Claude Code Sessions](#claude-code-sessions)
- [Codex Sessions](#codex-sessions)
- [Scheduling a Daily Run](#scheduling-a-daily-run)
- [Output Structure](#output-structure)
- [CLI Reference](#cli-reference)
- [How the Cache Works](#how-the-cache-works)
- [Troubleshooting](#troubleshooting)
- [Future Work](#future-work)
- [Security Notes](#security-notes)
---
## Terms of Service Warning
**Read this before using this tool.**
@@ -32,10 +53,31 @@ This tool is designed for a single user backing up their own conversations. Do n
```bash
git clone <repo-url>
cd ai-chat-exporter
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cd AIChatExporter
./ai-chat-exporter doctor
```
That's the whole install. The `ai-chat-exporter` wrapper creates `.venv` and
installs dependencies on first run, and reinstalls whenever `pyproject.toml`
changes — so there is no `python3 -m venv` / `source .venv/bin/activate` to
remember, on this machine or the next one you clone onto. Every command in this
README works the same way:
```bash
./ai-chat-exporter export --provider all
./ai-chat-exporter sync
```
The wrapper does **not** change directory: `.env`, `cache/` and `exports/` all
resolve against your current directory, which is what lets one checkout archive
different machines into different places. Run it from the repo. (It warns if you
don't, because a different working directory means a different `cache/manifest.json`
— which would re-export everything and orphan your existing Joplin notes.)
Prefer the traditional route, or want the test dependencies? That still works:
```bash
python3 -m venv .venv && source .venv/bin/activate && pip install -e ".[dev]"
```
### Windows
@@ -44,12 +86,28 @@ No admin access required. Run these in **Command Prompt** (`cmd.exe`) — it's t
```bat
git clone <repo-url>
cd ai-chat-exporter
python -m venv .venv
.venv\Scripts\activate
pip install -e ".[dev]"
cd AIChatExporter
ai-chat-exporter doctor
```
`ai-chat-exporter.cmd` does the same bootstrap as the POSIX wrapper — it creates
`.venv` and installs dependencies on first run. No `python -m venv`, no
`.venv\Scripts\activate`.
**How you invoke it differs between the two Windows shells:**
| Shell | Command |
|---|---|
| Command Prompt (`cmd.exe`) | `ai-chat-exporter export --provider all` |
| PowerShell | `.\ai-chat-exporter.cmd export --provider all` |
`cmd.exe` searches the current directory before `PATH` and resolves the bare name
through `PATHEXT`, so it finds `ai-chat-exporter.cmd` with no prefix and no
extension. PowerShell deliberately does *not* search the current directory, and
`.\ai-chat-exporter` there would resolve to the extensionless POSIX script, which
PowerShell cannot execute — so name the `.cmd` explicitly. Command Prompt is the
simpler of the two here.
All `ai-chat-exporter` commands work identically in Command Prompt.
**Using PowerShell instead?** If you prefer PowerShell, you may need to allow script execution first (one-time, current user only):
@@ -186,12 +244,26 @@ cp .env.example .env
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
### Notifications
| Variable | Default | Description |
|----------|---------|-------------|
| `NTFY_TOPIC` | — | [ntfy](https://ntfy.sh) topic to push run results to. Unset disables notifications entirely. |
| `NTFY_SERVER` | `https://ntfy.sh` | Point at your own host if self-hosting. |
| `NTFY_TOKEN` | — | Bearer token, for access-controlled topics. |
| `NTFY_NOTIFY` | `always` | `always` notifies on every run, `failure` only when something failed, `off` never. |
A topic on public ntfy.sh is readable by anyone who knows its name, so
notifications carry per-provider counts and a machine name only — never
conversation titles. See [Getting notified](#getting-notified).
### Cache & logging
| Variable | Default | Description |
|----------|---------|-------------|
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
| `AI_CHAT_EXPORTER_QUIET_CWD` | — | Set to `1` to silence the launcher's warning when run from outside the repo. Read by the `ai-chat-exporter` wrapper scripts, not by Python; the scheduler installers set it, since they always set the correct working directory. |
---
@@ -201,6 +273,28 @@ ChatGPT project conversations are stored separately from your main conversation
### Finding your project IDs
The quickest way is to let the exporter find them:
```
ai-chat-exporter projects
```
It lists every project your conversations actually belong to, marks which are
missing from `.env`, and prints a paste-ready `CHATGPT_PROJECT_IDS=` line
(`--write` updates `.env` for you). If your ChatGPT account does not include
the project on conversation summaries, add `--deep` and it reads each
conversation's detail instead — slower, one request per conversation, but
complete.
Why it matters: project attribution is resolved from each conversation's own
`gizmo_id`, so exports file correctly whether or not a project is configured.
But the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that
live only inside a project never appear in the default conversation list, so an
unlisted project's chats are never fetched at all. Every export run also names
any unconfigured project it encounters.
To find them by hand instead:
1. Open ChatGPT and click a Project in the left sidebar
2. Look at the browser URL — it will look like:
`https://chatgpt.com/g/g-p-68c2b2b3037c8191890036fb4ae3ed9f-my-project/project`
@@ -241,6 +335,125 @@ Sessions from all roots are merged by folder (no per-machine label); if the same
---
## Codex Sessions
The `codex` provider archives your local [Codex CLI](https://chatgpt.com/codex) agent transcripts — same deal as Claude Code: no tokens, no API, no ToS exposure. Rollout files are read from `~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl`.
```bash
ai-chat-exporter export --provider codex
ai-chat-exporter joplin --provider codex
```
Exports are prose-only by default, with the same placeholder format as Claude Code, into their own top-level **`AI-Codex`** Joplin notebook. Repo tags work the same way (`CODEX_REPO_TAG_IGNORE` to suppress), and additional roots can be scanned with `CODEX_DIR` (`:`-separated); `CODEX_HOME`'s `sessions/` is picked up automatically when that variable is set.
Three things differ from Claude Code, all forced by how Codex stores its data:
**Reasoning cannot be exported.** Codex encrypts it at rest — every reasoning record carries `encrypted_content` with no plaintext summary in any layer of the file. It is always dropped and counted; `EXPORTER_HIDDEN_CONTENT=full` cannot bring it back.
**Incomplete tool calls are reported.** Codex writes each session twice in one file: its own typed items (what actually ran) and the raw model-facing wire format (everything attempted). The provider reads the typed layer — it is already decoded, and it omits harness plumbing that would otherwise need stripping — but cross-checks the raw layer for calls that never produced a result, so placeholders read `3 calls: exec_command ×3 (+2 did not complete)`. Those are commands that failed to launch, that you aborted, or that were still running when the turn ended.
**Cloud tasks are out of scope.** `codex cloud` tasks run server-side and are reachable at `chatgpt.com/backend-api/api/codex/tasks`, but local CLI sessions are never uploaded there, so the cloud API is not an alternative source for these transcripts and this provider stays entirely offline. If you start using `codex cloud exec`, those transcripts would be cloud-only and would need separate work.
---
## Scheduling a Daily Run
`sync` chains `export` then `joplin` in one invocation, which is what a scheduler
wants — one command, and a meaningful exit code so a failed run is visible
instead of silent.
```bash
./ai-chat-exporter sync # every configured provider
./ai-chat-exporter sync --provider codex # just one
```
### Linux (systemd user timer)
```bash
./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
./scheduling/install-systemd-timer.sh --uninstall
```
Defaults to 09:00 daily (`--time 21:30` to change). `Persistent=true` means a
machine that was off at the scheduled time runs the archive at next boot rather
than skipping the day. To archive while logged out, `loginctl enable-linger $USER`.
Check on it with `systemctl --user list-timers aichat-sync.timer` and
`journalctl --user -u aichat-sync.service -n 50`.
### Windows (Task Scheduler)
```powershell
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
```
Per-user task, no admin rights needed. `-StartWhenAvailable` is the counterpart
of systemd's `Persistent=true`.
### What to know before you rely on it
**List only the providers that work unattended on that machine.** The local
providers (`claude-code`, `codex`) need no credentials and always work. The web
providers depend on a session token that expires and can only be refreshed by
hand via DevTools — so on a machine where that token is stale, scheduling them
means a failed run every single day, which is a good way to learn to ignore
failures you actually want to see. That is why `--provider` is repeatable in both
installers: schedule the coding machine for `claude-code` + `codex`, the browser
machine for `chatgpt` + `claude`.
**Both installers pass `--joplin-optional`.** If Joplin desktop isn't running,
the sync warns instead of failing: the export has already captured the local
transcripts (the part that can disappear), and the notes are rebuilt from the
cache on the next run that finds Joplin up.
**One provider failing does not skip the rest.** Both installers loop over the
providers in a single action rather than one action each, because systemd
`oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only
the last action's result. Every provider is attempted; the run still exits
non-zero if any failed.
### Getting notified
A scheduled run is silent by default. Output goes to three pull-only places: the
exporter's own log (`cache/logs/exporter.log`), the systemd journal on Linux
(`journalctl --user -u aichat-sync.service`), and Task Scheduler's
`LastTaskResult` on Windows.
To have runs report back, set an [ntfy](https://ntfy.sh) topic in `.env`:
```bash
NTFY_TOPIC=my-archive-topic
```
Then check it works before waiting on a scheduled run:
```bash
ai-chat-exporter notify # show settings
ai-chat-exporter notify --test # send a test push
```
`sync` pushes a result whenever a topic is configured — a success carries the
per-provider counts at low/default priority, a failure carries the reason at
high priority with an alert tag, so a failed archive is distinguishable from a
quiet one on your phone. `NTFY_NOTIFY=failure` notifies only on failure; `off`
disables it; `--notify` / `--no-notify` override per run.
The message includes the **machine name**, which matters because both machines
archive into one topic. It contains counts only — never conversation titles. A
topic on public ntfy.sh is readable by anyone who knows its name, so if you want
it private, self-host (`NTFY_SERVER`) or use an access-controlled topic with
`NTFY_TOKEN`.
A notification is never fatal: if ntfy is unreachable, the run logs a warning and
still reports its real exit code.
**Acknowledge the ToS notice once, interactively.** It's stored in the cache
manifest per machine. Until then a scheduled run exits 1 with an explanation
rather than hanging on a prompt no one can answer.
---
## Output Structure
All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks.
@@ -349,7 +562,7 @@ ai-chat-exporter export --output /path/to/my/notes
ai-chat-exporter export --dry-run
```
Options: `--provider [chatgpt|claude|claude-code|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--hidden-content [full|placeholder|omit]`, `--download-media [images|all|off]`, `--max-conversations N`, `--force`, `--dry-run`
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--hidden-content [full|placeholder|omit]`, `--download-media [images|all|off]`, `--max-conversations N`, `--force`, `--dry-run`
**Re-rendering the whole archive after an upgrade.** New formatting or features (collapse policy, media downloads) only change conversations as they're re-exported. To re-render everything you already have, use `--force` — it re-exports every conversation even if unchanged, **without** `cache --clear`, so your Joplin note links are preserved (a later `joplin` run updates the existing notes instead of duplicating them).
@@ -411,7 +624,59 @@ Reads the local export cache and pushes each exported Markdown file to Joplin as
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
4. Joplin desktop must be open when you run this command
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--project NAME`, `--dry-run`
### `sync` — Export and sync in one run
```bash
# The whole archive run: export, then push to Joplin
ai-chat-exporter sync
# One provider
ai-chat-exporter sync --provider codex
# Export only; don't touch Joplin
ai-chat-exporter sync --skip-joplin
# Joplin being closed is a warning, not a failure (used by the schedulers)
ai-chat-exporter sync --joplin-optional
```
Equivalent to `export` followed by `joplin` with the same `--provider`. Intended
for scheduled runs — see [Scheduling a Daily Run](#scheduling-a-daily-run).
Unlike the individual commands, `sync` sets a **meaningful exit code**: non-zero
if any conversation failed to export or any note failed to sync. A provider whose
listing call fails outright (an expired web session token being the usual cause)
counts its whole batch as failed. A provider that is simply unconfigured, or that
had nothing new, is ordinary success. `export` on its own always exits 0, which
is fine when you're reading the summary table and useless to a scheduler.
`--joplin-optional` downgrades an unreachable Joplin to a warning: the export has
already captured the local transcripts, and the notes are rebuilt from the cache
by the next run that finds Joplin open.
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--since YYYY-MM-DD`, `--hidden-content [full|placeholder|omit]`, `--max-conversations N`, `--skip-joplin`, `--joplin-optional`, `--notify/--no-notify`, `--dry-run`
Note this is a deliberate subset of `export`'s options — `--format`, `--output`,
`--project`, `--download-media` and `--force` are not passed through. Use
`export` directly for those. (`--download-media` still applies from `.env`; the
flag is only a per-run override.)
### `notify` — Push-notification settings and test
```bash
# Show the current settings
ai-chat-exporter notify
# Send a test push to confirm the topic works
ai-chat-exporter notify --test
```
Shows the resolved ntfy configuration and which machine name will appear in the
title. See [Getting notified](#getting-notified) for what a scheduled run sends.
Options: `--test`
### `prune` — Delete stale export files
@@ -429,6 +694,50 @@ Joplin. Refuses to run when the manifest is empty (e.g. right after
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
command separately verifies that every manifest entry's file exists on disk.
### `projects` — Discover ChatGPT project IDs
```bash
# List the projects your conversations belong to
ai-chat-exporter projects
# Also inspect conversations whose listing entry doesn't name a project
ai-chat-exporter projects --deep
# Write the discovered IDs straight into .env
ai-chat-exporter projects --write
```
`CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is
invisible to the listing pass — conversations that live *only* inside that
project are never fetched at all. This reports every project your conversations
belong to, marks the ones absent from `.env`, and prints a paste-ready line.
`--deep` fetches each conversation's detail when the listing doesn't name its
project: complete, but one request per conversation, so it's slow.
Options: `--deep`, `--write`
### `canary` — Check for provider API drift
```bash
ai-chat-exporter canary
ai-chat-exporter canary --provider chatgpt
```
The web providers are undocumented internal APIs that can change shape without
notice, and the failure mode is silent — a renamed field means content is
quietly dropped rather than an error being raised. The canary fetches one
listing page and one conversation per provider and asserts only the fields the
normalizer actually depends on.
Findings are `ERROR` (a load-bearing field is missing or mistyped — the parser
will break or silently lose data) or `WARN` (something unfamiliar appeared;
worth investigating, not necessarily broken). **Exits non-zero on any ERROR**,
so it can be scheduled or run in CI. Local providers have no remote schema and
are not probed.
Options: `--provider [chatgpt|claude|all]`
### `cache` — Manage the sync manifest
```bash
@@ -528,7 +837,10 @@ No new or updated conversations since your last run. To verify: `ai-chat-exporte
See `FUTURE.md` for the full roadmap. Current priorities:
- **Watch/scheduled mode** on the way to a headless StartOS service
- **A StartOS service** that centralises every machine's conversations into one
corpus and owns the Joplin connection, so each machine only has to upload
(`FUTURE.md` §8)
- **Splitting this README** into a short overview plus separate documents
---
+67
View File
@@ -0,0 +1,67 @@
#!/usr/bin/env bash
# ai-chat-exporter — run the exporter without activating a virtualenv.
#
# cd /path/to/AIChatExporter
# ./ai-chat-exporter export --provider all
#
# Creates .venv and installs dependencies on first run, so a fresh clone on a
# new machine needs no `python3 -m venv` / `source .venv/bin/activate` ceremony.
# Reinstalls automatically when pyproject.toml changes.
#
# Working directory is deliberately NOT changed: EXPORT_DIR, CACHE_DIR and .env
# discovery are all relative to your current directory, which is what lets the
# same checkout archive different machines into different places. Run it from
# the repo (see the warning below) unless you mean otherwise.
#
# Windows equivalent: ai-chat-exporter.cmd (same directory).
set -euo pipefail
DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
VENV="$DIR/.venv"
PY="$VENV/bin/python"
STAMP="$VENV/.deps-stamp"
# The cache manifest is read from ./cache by default. Running from somewhere
# else silently starts a *new, empty* archive rather than failing, which would
# re-export everything and orphan the existing Joplin notes. Warn, don't block —
# a deliberate second archive is a legitimate thing to want.
if [ "$PWD" != "$DIR" ] && [ -z "${AI_CHAT_EXPORTER_QUIET_CWD:-}" ]; then
echo "warning: running from $PWD, not $DIR" >&2
echo " cache/ and exports/ resolve against the current directory," >&2
echo " so this may start a separate archive. Set" >&2
echo " AI_CHAT_EXPORTER_QUIET_CWD=1 to silence this." >&2
fi
find_python() {
for candidate in python3 python; do
if command -v "$candidate" >/dev/null 2>&1; then
# Needs >=3.11 (pyproject requires-python).
if "$candidate" -c 'import sys; sys.exit(0 if sys.version_info >= (3, 11) else 1)' 2>/dev/null; then
command -v "$candidate"
return 0
fi
fi
done
return 1
}
if [ ! -x "$PY" ]; then
BOOTSTRAP_PY="$(find_python)" || {
echo "error: no python3 >= 3.11 found on PATH — install it and re-run." >&2
exit 1
}
echo "Creating virtualenv in $VENV …" >&2
"$BOOTSTRAP_PY" -m venv "$VENV"
rm -f "$STAMP"
fi
# Install (or refresh) dependencies when the venv is new or pyproject changed.
if [ ! -f "$STAMP" ] || [ "$DIR/pyproject.toml" -nt "$STAMP" ]; then
echo "Installing dependencies …" >&2
"$PY" -m pip install --quiet --upgrade pip
"$PY" -m pip install --quiet -e "$DIR"
touch "$STAMP"
fi
exec "$PY" -m src.main "$@"
+83
View File
@@ -0,0 +1,83 @@
@echo off
rem ai-chat-exporter.cmd - run the exporter without activating a virtualenv.
rem
rem Command Prompt (cmd.exe) - the documented way:
rem cd C:\path\to\AIChatExporter
rem ai-chat-exporter export --provider all
rem
rem cmd.exe searches the current directory before PATH and resolves the bare
rem name through PATHEXT (which includes .CMD), so no ".\" and no extension are
rem needed. The extensionless POSIX sibling is ignored: it is not in PATHEXT.
rem
rem PowerShell does NOT search the current directory, and ".\ai-chat-exporter"
rem there would resolve to the extensionless POSIX script, which PowerShell
rem cannot run. In PowerShell, name this file explicitly:
rem .\ai-chat-exporter.cmd export --provider all
rem
rem Creates .venv and installs dependencies on first run, so a fresh clone needs
rem no "python -m venv" / ".venv\Scripts\activate" ceremony. Reinstalls
rem automatically when pyproject.toml changes.
rem
rem The working directory is deliberately NOT changed - EXPORT_DIR, CACHE_DIR
rem and .env discovery are all relative to it. POSIX equivalent: ai-chat-exporter
setlocal enabledelayedexpansion
set "DIR=%~dp0"
if "%DIR:~-1%"=="\" set "DIR=%DIR:~0,-1%"
set "VENV=%DIR%\.venv"
set "PY=%VENV%\Scripts\python.exe"
set "STAMP=%VENV%\.deps-stamp"
rem See the POSIX script for why this warns rather than blocks: cache\ resolves
rem against the current directory, so the wrong one silently starts a second
rem archive instead of failing.
if /i not "%CD%"=="%DIR%" if "%AI_CHAT_EXPORTER_QUIET_CWD%"=="" (
echo warning: running from %CD%, not %DIR% 1>&2
echo cache\ and exports\ resolve against the current directory, 1>&2
echo so this may start a separate archive. Set 1>&2
echo AI_CHAT_EXPORTER_QUIET_CWD=1 to silence this. 1>&2
)
if not exist "%PY%" (
echo Creating virtualenv in %VENV% ... 1>&2
rem The py launcher is the reliable way to get a specific version; fall back
rem to whatever "python" is if it is not installed.
where py >nul 2>&1
if !errorlevel! equ 0 (
py -3 -m venv "%VENV%"
) else (
python -m venv "%VENV%"
)
if not exist "%PY%" (
echo error: could not create a virtualenv - install Python 3.11+ from 1>&2
echo python.org or the Microsoft Store, then re-run. 1>&2
exit /b 1
)
if exist "%STAMP%" del "%STAMP%"
)
rem Staleness check, in pure batch: %%~tF is the file's last-modified stamp, so
rem storing it and comparing strings needs no external process. The obvious
rem alternative - asking PowerShell to compare timestamps - costs ~1s of
rem interpreter startup on *every* command, which is a lot to pay to almost
rem always learn that nothing changed.
set "PYPROJ_TIME="
for %%F in ("%DIR%\pyproject.toml") do set "PYPROJ_TIME=%%~tF"
set "SAVED_TIME="
if exist "%STAMP%" set /p SAVED_TIME=<"%STAMP%"
if not "%SAVED_TIME%"=="%PYPROJ_TIME%" (
echo Installing dependencies ... 1>&2
"%PY%" -m pip install --quiet --upgrade pip
"%PY%" -m pip install --quiet -e "%DIR%"
if !errorlevel! neq 0 (
echo error: dependency installation failed. 1>&2
exit /b !errorlevel!
)
> "%STAMP%" echo !PYPROJ_TIME!
)
"%PY%" -m src.main %*
exit /b %errorlevel%
+2 -2
View File
@@ -4,8 +4,8 @@ build-backend = "setuptools.build_meta"
[project]
name = "ai-chat-exporter"
version = "0.8.0"
description = "Export ChatGPT and Claude conversation history to Markdown for personal archival in Joplin"
version = "0.9.0"
description = "Archive ChatGPT, Claude, Claude Code and Codex conversation history to Markdown for personal backup in Joplin"
requires-python = ">=3.11"
dependencies = [
"requests==2.31.0",
+93
View File
@@ -0,0 +1,93 @@
<#
.SYNOPSIS
Register a Windows scheduled task that runs `ai-chat-exporter sync` daily.
.DESCRIPTION
The Windows counterpart to install-systemd-timer.sh. Creates a per-user task
(no admin rights needed) that runs the exporter from this repository, with
the working directory set to the repo so .env, cache\ and exports\ resolve
exactly as they do for an interactive run.
-Provider is repeatable. The CLI takes one provider per run, so each becomes
its own action, executed in order. List only the providers that work
unattended on this machine: a web provider whose session token has expired
fails the task every day, which trains you to ignore the failures you
actually want to notice.
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Provider all -Time 21:30
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
#>
[CmdletBinding()]
param(
[string[]]$Provider = @('all'),
[string]$Time = '09:00',
[string]$TaskName = 'AiChatExporterSync',
[switch]$Unregister
)
$ErrorActionPreference = 'Stop'
$repo = Split-Path -Parent $PSScriptRoot
$launcher = Join-Path $repo 'ai-chat-exporter.cmd'
if ($Unregister) {
Unregister-ScheduledTask -TaskName $TaskName -Confirm:$false -ErrorAction SilentlyContinue
Write-Host "Removed scheduled task '$TaskName'."
return
}
if (-not (Test-Path $launcher)) {
throw "Launcher not found at $launcher"
}
# A single action looping over the providers, rather than one action each.
# Task Scheduler runs multiple actions in order but reports only the last one's
# result, so a failure in an earlier provider would be invisible. The loop keeps
# going after a failure and propagates a non-zero exit code.
#
# /v:on and !RC! are required, not stylistic: cmd expands every %VAR% on a
# command line *before* running any of it, so "exit /b %RC%" would report the
# value RC had before the loop ever ran - i.e. always success. Delayed expansion
# reads it at the point of use.
$loop = ($Provider | ForEach-Object { "`"$launcher`" sync --provider $_ --joplin-optional || set RC=1" }) -join ' & '
$taskArgs = "/v:on /c set RC=0 & $loop & exit /b !RC!"
$actions = New-ScheduledTaskAction -Execute 'cmd.exe' `
-Argument $taskArgs `
-WorkingDirectory $repo
$trigger = New-ScheduledTaskTrigger -Daily -At $Time
# StartWhenAvailable is the counterpart of systemd's Persistent=true: a machine
# that was asleep at the scheduled time runs the archive when it wakes, rather
# than skipping the day entirely.
$settings = New-ScheduledTaskSettingsSet `
-StartWhenAvailable `
-DontStopIfGoingOnBatteries `
-AllowStartIfOnBatteries `
-ExecutionTimeLimit (New-TimeSpan -Hours 2)
Register-ScheduledTask -TaskName $TaskName `
-Action $actions `
-Trigger $trigger `
-Settings $settings `
-Description 'Export AI chat history and sync it to Joplin' `
-Force | Out-Null
Write-Host "Registered '$TaskName' - daily at $Time for: $($Provider -join ', ')"
Write-Host ''
Write-Host 'Command the task will run:'
Write-Host " cmd.exe $taskArgs"
Write-Host " (working directory: $repo)"
Write-Host ''
Write-Host 'Next steps:'
Write-Host " * Run it once now: Start-ScheduledTask -TaskName $TaskName"
Write-Host " * Check the result: Get-ScheduledTaskInfo -TaskName $TaskName"
Write-Host " * Read the log: Get-Content '$repo\cache\logs\exporter.log' -Tail 50"
Write-Host ' * The terms-of-service notice must have been acknowledged'
Write-Host ' interactively once on this machine, or the task exits 1.'
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Install a systemd *user* timer that runs `ai-chat-exporter sync` on a schedule.
#
# ./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
# ./scheduling/install-systemd-timer.sh --provider all --time 21:30
# ./scheduling/install-systemd-timer.sh --uninstall
#
# User units (not system units) are the right scope: the archive is per-user,
# the .env holds that user's session tokens, and Joplin runs in their session.
#
# --provider is repeatable. The CLI takes one provider per run, so each one
# becomes its own ExecStart line, executed in order. Prefer listing only the
# providers that actually work unattended on this machine — a web provider with
# an expired session token will fail the unit every single day, which trains you
# to ignore the failure you actually want to see.
set -euo pipefail
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
UNIT_DIR="$HOME/.config/systemd/user"
NAME="aichat-sync"
TIME="09:00"
PROVIDERS=()
UNINSTALL=0
while [ $# -gt 0 ]; do
case "$1" in
--provider) PROVIDERS+=("$2"); shift 2 ;;
--time) TIME="$2"; shift 2 ;;
--name) NAME="$2"; shift 2 ;;
--uninstall) UNINSTALL=1; shift ;;
-h|--help) sed -n '2,20p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
if [ "$UNINSTALL" -eq 1 ]; then
systemctl --user disable --now "$NAME.timer" 2>/dev/null || true
rm -f "$UNIT_DIR/$NAME.timer" "$UNIT_DIR/$NAME.service"
systemctl --user daemon-reload
echo "Removed $NAME.timer and $NAME.service."
exit 0
fi
[ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all")
if [ ! -x "$REPO/ai-chat-exporter" ]; then
echo "error: $REPO/ai-chat-exporter is missing or not executable." >&2
exit 1
fi
mkdir -p "$UNIT_DIR"
# WorkingDirectory is the point of the whole unit: .env, cache/ and exports/ all
# resolve against it, so the scheduled run writes to the same archive an
# interactive run from this directory would.
{
echo "[Unit]"
echo "Description=AI chat archive sync"
echo "Documentation=file://$REPO/README.md"
echo "After=network-online.target"
echo "Wants=network-online.target"
echo
echo "[Service]"
echo "Type=oneshot"
echo "WorkingDirectory=$REPO"
echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1"
# One ExecStart per provider would stop at the first failure, silently
# skipping the rest — an expired ChatGPT token would mean codex never runs.
# Loop instead, so every provider is attempted and the unit still reports
# failure if any of them failed.
printf 'ExecStart=/bin/sh -c '\''rc=0; for p in %s; do "$0" sync --provider "$p" --joplin-optional || rc=1; done; exit $rc'\'' %s\n' \
"${PROVIDERS[*]}" "$REPO/ai-chat-exporter"
} > "$UNIT_DIR/$NAME.service"
# Persistent=true runs a missed schedule at the next boot — the machine being
# off at 09:00 should delay the archive, not skip it. RandomizedDelaySec keeps
# the web providers from being hit at exactly the same second every day.
cat > "$UNIT_DIR/$NAME.timer" <<EOF
[Unit]
Description=Run the AI chat archive sync daily
[Timer]
OnCalendar=*-*-* $TIME:00
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now "$NAME.timer"
echo "Installed $NAME.timer — daily at $TIME for: ${PROVIDERS[*]}"
echo
systemctl --user list-timers "$NAME.timer" --no-pager || true
cat <<EOF
Next steps:
• Timers only run while you have a session. To archive when logged out:
loginctl enable-linger $USER
• Run it once by hand to confirm:
systemctl --user start $NAME.service
• Read the log:
journalctl --user -u $NAME.service -n 50
• The terms-of-service notice must have been acknowledged interactively at
least once on this machine, or the unit exits 1 with an explanation.
EOF
+4
View File
@@ -409,6 +409,10 @@ _PROVIDER_DISPLAY = {
# intermixed with Claude web projects and named by dev folder). Existing
# notes self-heal into here on the next sync — update_note moves them.
"claude-code": "AI-ClaudeCode",
# Codex CLI coding sessions get their own top-level notebook for the same
# reason Claude Code does — they are a distinct surface, not a ChatGPT
# project, and burying them under AI-ChatGPT makes both harder to browse.
"codex": "AI-Codex",
}
+387 -5
View File
@@ -99,6 +99,19 @@ def cli(ctx: click.Context, verbose: bool, quiet: bool, debug: bool, no_log_file
# ToS gate: must happen before any command executes
if not cache.is_tos_acknowledged():
# Non-interactive (cron, systemd timer, Task Scheduler): click.prompt
# raises Abort on a closed stdin, which used to exit 0 — a scheduled
# run would report success having archived nothing. Fail loudly instead
# and tell the operator how to clear the gate once, by hand.
if not sys.stdin.isatty():
err_console.print(
"[red]Terms-of-service notice has not been acknowledged, and there is "
"no terminal to ask on.[/red]\n"
"Run any command once interactively (e.g. 'ai-chat-exporter doctor') "
"and type 'yes' to acknowledge. The acknowledgement is stored in the "
"cache manifest and applies to every later run on this machine."
)
sys.exit(1)
try:
answer = click.prompt(TOS_NOTICE, default="", show_default=False).strip().lower()
except (click.Abort, KeyboardInterrupt):
@@ -332,6 +345,143 @@ def doctor(ctx: click.Context) -> None:
sys.exit(1)
@cli.command()
@click.option(
"--deep",
is_flag=True,
help=(
"Fetch each conversation's detail when the listing does not name its "
"project. Slow (one request per conversation) but complete."
),
)
@click.option("--write", is_flag=True, help="Write the discovered IDs to .env.")
@click.pass_context
def projects(ctx: click.Context, deep: bool, write: bool) -> None:
"""Discover ChatGPT project IDs, including ones missing from your config.
CHATGPT_PROJECT_IDS has to be maintained by hand, and a project missing
from it is invisible to the listing pass: conversations that live only
inside that project are never fetched at all. This finds the projects your
account actually uses and prints a paste-ready line.
"""
cfg = _load_config_or_exit(ctx.obj["debug"])
if not cfg.chatgpt_session_token:
console.print("[red]CHATGPT_SESSION_TOKEN is not set — run 'ai-chat-exporter auth'.[/red]")
sys.exit(1)
from src.providers.chatgpt import ChatGPTProvider
try:
prov = ChatGPTProvider(
session_token=cfg.chatgpt_session_token,
session_token_1=cfg.chatgpt_session_token_1,
project_ids=cfg.chatgpt_project_ids,
)
except ProviderError as e:
_handle_provider_error(e, ctx.obj["debug"])
sys.exit(1)
configured = list(cfg.chatgpt_project_ids)
console.print(f"\n[bold cyan][CHATGPT][/bold cyan] {len(configured)} project ID(s) configured")
console.print("Listing conversations…")
try:
summaries = prov.fetch_all_conversations(since=None)
except ProviderError as e:
_handle_provider_error(e, ctx.obj["debug"])
sys.exit(1)
found: dict[str, str] = {}
missing_gizmo: list[dict] = []
for conv in summaries:
gizmo_id = conv.get("gizmo_id")
if gizmo_id and str(gizmo_id).startswith("g-p-"):
found.setdefault(gizmo_id, "")
elif gizmo_id is None:
missing_gizmo.append(conv)
if found:
console.print(
f" [dim]The listing names a project on {len(summaries) - len(missing_gizmo)} "
f"conversation(s) — no detail fetches needed.[/dim]"
)
else:
console.print(
" [yellow]The listing does not carry gizmo_id, so projects can only be "
"read from conversation details.[/yellow]"
)
if deep and missing_gizmo:
console.print(
f" Fetching detail for {len(missing_gizmo)} conversation(s) "
f"(~{len(missing_gizmo) * cfg.request_delay / 60:.0f} min at "
f"{cfg.request_delay}s pacing)…"
)
from rich.progress import (
BarColumn, Progress, SpinnerColumn, TaskProgressColumn, TextColumn,
)
with Progress(
SpinnerColumn(),
TextColumn("[progress.description]{task.description}"),
BarColumn(),
TaskProgressColumn(),
console=console,
) as progress:
task = progress.add_task("Scanning…", total=len(missing_gizmo))
for conv in missing_gizmo:
conv_id = conv.get("id") or conv.get("conversation_id")
if conv_id:
try:
raw = prov.get_conversation(conv_id)
gizmo_id = raw.get("gizmo_id")
if gizmo_id and str(gizmo_id).startswith("g-p-"):
found.setdefault(gizmo_id, "")
except ProviderError:
pass
progress.advance(task)
elif missing_gizmo and not found:
console.print(
f" [dim]Re-run with --deep to read {len(missing_gizmo)} conversation "
"detail(s). Nothing to report without it.[/dim]"
)
for gizmo_id in list(found) + [p for p in configured if p not in found]:
found[gizmo_id] = prov._fetch_project_name(gizmo_id)
table = Table(title="ChatGPT Projects", show_header=True)
table.add_column("Project", style="cyan")
table.add_column("ID", style="dim")
table.add_column("In .env")
for gizmo_id, name in sorted(found.items(), key=lambda kv: kv[1].lower()):
in_env = gizmo_id in configured
table.add_row(
name,
gizmo_id,
"[green]yes[/green]" if in_env else "[yellow]NO — add it[/yellow]",
)
console.print(table)
new_ids = [g for g in found if g not in configured]
if not new_ids:
console.print("[green]Every project found is already configured.[/green]")
return
combined = configured + new_ids
line = "CHATGPT_PROJECT_IDS=" + ",".join(combined)
console.print(
f"\n[bold]{len(new_ids)} project(s) missing from your config.[/bold] "
"Conversations that live only inside them are not being exported."
)
console.print("\n[dim]Paste into .env:[/dim]")
console.print(line)
if write:
_set_env_key("CHATGPT_PROJECT_IDS", ",".join(combined))
else:
console.print("\n[dim]Or re-run with --write to update .env directly.[/dim]")
@cli.command()
@click.option(
"--provider",
@@ -547,7 +697,7 @@ def _print_doctor_table(checks: list[dict]) -> None:
@cli.command()
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False),
type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all",
show_default=True,
help="Which provider to export.",
@@ -853,6 +1003,10 @@ def export(
if cache.exported_at(prov_name, c.get("id") or c.get("uuid", "")) < campaign_at
)
# `sync` reads this to decide the process exit code — a scheduled run that
# exported nothing because a token expired must not look like success.
ctx.obj["last_export_summary"] = summary
if not dry_run:
_print_export_summary(summary)
if force:
@@ -939,6 +1093,21 @@ def _resolve_providers(provider: str, cfg) -> list[tuple[str, object]]:
", ".join(str(r) for r in cc_roots),
)
if provider in ("codex", "all"):
from src.providers.codex import CodexProvider, resolve_roots as codex_roots_fn
cx_roots = codex_roots_fn()
if any(r.is_dir() for r in cx_roots):
result.append((
"codex",
CodexProvider(hidden_content=cfg.hidden_content),
))
elif provider == "codex":
logging.getLogger(__name__).warning(
"[codex] Skipping — none of these roots found (set "
"CODEX_DIR, ':'-separated for multiple): %s",
", ".join(str(r) for r in cx_roots),
)
return result
@@ -1027,6 +1196,217 @@ def _print_export_summary(summary: dict[str, dict[str, int]]) -> None:
console.print(table)
# ──────────────────────────────────────────────────────────────────────────────
# sync command
# ──────────────────────────────────────────────────────────────────────────────
@cli.command()
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all",
show_default=True,
help="Which provider(s) to export and sync.",
)
@click.option("--since", default=None, help="Only export conversations updated after this date (YYYY-MM-DD).")
@click.option(
"--hidden-content",
type=click.Choice(["full", "placeholder", "omit"], case_sensitive=False),
default=None,
help="Overrides EXPORTER_HIDDEN_CONTENT for this run.",
)
@click.option(
"--max-conversations",
type=click.IntRange(min=1),
default=None,
help="Cap how many conversations are downloaded this run (per provider).",
)
@click.option("--skip-joplin", is_flag=True, help="Export only; do not sync to Joplin.")
@click.option(
"--joplin-optional",
is_flag=True,
help=(
"Treat an unreachable Joplin as a warning rather than a failure. For "
"scheduled runs, where Joplin desktop may simply not be open yet."
),
)
@click.option(
"--notify/--no-notify",
"notify_flag",
default=None,
help=(
"Send an ntfy push with the result. Defaults to on whenever NTFY_TOPIC "
"is configured (see NTFY_NOTIFY to notify only on failure)."
),
)
@click.option("--dry-run", is_flag=True, help="Show what would happen without writing or sending anything.")
@click.pass_context
def sync(
ctx: click.Context,
provider: str,
since: str | None,
hidden_content: str | None,
max_conversations: int | None,
skip_joplin: bool,
joplin_optional: bool,
notify_flag: bool | None,
dry_run: bool,
) -> None:
"""Export, then sync to Joplin — the whole archive run in one command.
Equivalent to `export` followed by `joplin` with the same --provider, which
is what a scheduled run wants: one line in a systemd timer, cron entry or
Windows scheduled task.
Unlike the individual commands, this one sets a meaningful exit code: it
exits non-zero if any conversation failed to export or any note failed to
sync. A provider whose listing call fails outright — an expired web session
token being the usual cause — counts its whole batch as failed, so a
scheduler sees a real failure rather than a silent no-op.
Note this does *not* flag a provider that is simply unconfigured, or one
that legitimately had nothing new to export; both are ordinary success.
"""
failures: list[str] = []
ctx.invoke(
export,
provider=provider,
since=since,
hidden_content=hidden_content,
max_conversations=max_conversations,
dry_run=dry_run,
)
export_summary = ctx.obj.get("last_export_summary") or {}
for prov_name, counts in export_summary.items():
if counts.get("failed"):
failures.append(f"{prov_name}: {counts['failed']} conversation(s) failed to export")
if skip_joplin:
console.print("[dim]Skipping Joplin sync (--skip-joplin).[/dim]")
else:
try:
ctx.invoke(joplin, provider=provider, dry_run=dry_run)
except SystemExit as e:
# `joplin` exits non-zero when the desktop app isn't reachable. On a
# timer that is routine — the machine may not be unlocked yet — and
# the export (the part that captures data which can disappear) has
# already succeeded. The notes are rebuilt from the cache on the next
# run that finds Joplin up, so nothing is lost by carrying on.
if not joplin_optional or e.code in (0, None):
raise
console.print(
"[yellow]Joplin sync skipped — Joplin is not reachable "
"(--joplin-optional). Exported files are on disk; the next run "
"with Joplin open will sync them.[/yellow]"
)
else:
joplin_summary = ctx.obj.get("last_joplin_summary") or {}
for prov_name, counts in joplin_summary.items():
if counts.get("failed"):
failures.append(f"{prov_name}: {counts['failed']} note(s) failed to sync")
_maybe_notify(ctx, failures, notify_flag, dry_run)
if failures:
err_console.print("\n[red]Sync completed with failures:[/red]")
for line in failures:
err_console.print(f" [red]•[/red] {line}")
sys.exit(1)
console.print("\n[green]Sync complete.[/green]")
def _maybe_notify(
ctx: click.Context, failures: list[str], notify_flag: bool | None, dry_run: bool
) -> None:
"""Push the run result to ntfy, if configured. Never raises."""
from src import notify as notify_mod
if dry_run:
return
# --notify/--no-notify overrides; otherwise notify whenever a topic is set.
if notify_flag is False:
return
if notify_flag is None and not notify_mod.is_configured():
return
if notify_flag is None and (
notify_mod.resolve_notify_policy() == notify_mod.NOTIFY_FAILURE and not failures
):
return
title, message, tags, priority = notify_mod.format_summary(
ctx.obj.get("last_export_summary"),
ctx.obj.get("last_joplin_summary"),
failures,
)
notify_mod.send(title, message, tags=tags, priority=priority)
# ──────────────────────────────────────────────────────────────────────────────
# notify command
# ──────────────────────────────────────────────────────────────────────────────
@cli.command()
@click.option("--test", "send_test", is_flag=True, help="Send a test notification now.")
@click.pass_context
def notify(ctx: click.Context, send_test: bool) -> None:
"""Show ntfy notification settings, and optionally send a test push.
Notifications are how an unattended run reports back: the log file, the
systemd journal and Task Scheduler's exit code are all pull-only.
"""
from src import notify as notify_mod
topic = os.getenv("NTFY_TOPIC", "").strip()
server = (os.getenv("NTFY_SERVER", "").strip() or notify_mod.DEFAULT_SERVER).rstrip("/")
policy = notify_mod.resolve_notify_policy()
has_token = bool(os.getenv("NTFY_TOKEN", "").strip())
table = Table(title="Notification settings")
table.add_column("Setting", style="bold")
table.add_column("Value")
table.add_row("NTFY_TOPIC", topic or "[dim]not set — notifications disabled[/dim]")
table.add_row("NTFY_SERVER", server)
table.add_row("NTFY_NOTIFY", policy)
table.add_row("NTFY_TOKEN", "set" if has_token else "[dim]not set[/dim]")
table.add_row("This machine", notify_mod.machine_name())
console.print(table)
if not topic:
console.print(
"\nSet [bold]NTFY_TOPIC[/bold] in .env to enable push notifications, "
"then re-run with --test."
)
return
if not has_token and server == notify_mod.DEFAULT_SERVER:
console.print(
"\n[yellow]Note:[/yellow] a topic on public ntfy.sh is readable by "
"anyone who knows its name. Notifications carry counts and a machine "
"name only — never conversation titles."
)
if send_test:
ok = notify_mod.send(
f"AI archive test - {notify_mod.machine_name()}",
"Test notification from ai-chat-exporter. If you can read this, "
"scheduled runs will reach you here.",
tags="test_tube",
priority="default",
)
if ok:
console.print(f"\n[green]Sent to {server}/{topic}[/green]")
else:
err_console.print(
f"\n[red]Could not send to {server}/{topic}[/red] — see the log "
"for the reason (run with --debug for detail)."
)
sys.exit(1)
# ──────────────────────────────────────────────────────────────────────────────
# list command
# ──────────────────────────────────────────────────────────────────────────────
@@ -1035,7 +1415,7 @@ def _print_export_summary(summary: dict[str, dict[str, int]]) -> None:
@cli.command(name="list")
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False),
type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all",
show_default=True,
)
@@ -1101,7 +1481,7 @@ def list_conversations(ctx: click.Context, provider: str, project_filter: str |
@click.option("--clear", is_flag=True, help="Clear cached entries.")
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False),
type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all",
help="Provider to target (used with --clear).",
)
@@ -1239,7 +1619,7 @@ def prune(ctx: click.Context, dry_run: bool, yes: bool) -> None:
@cli.command()
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False),
type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all",
show_default=True,
help="Which provider's conversations to sync to Joplin.",
@@ -1312,7 +1692,7 @@ def joplin(ctx: click.Context, provider: str, project_filter: str | None, dry_ru
# Determine which providers to process
providers_to_sync: list[str] = []
for prov in ("chatgpt", "claude", "claude-code"):
for prov in ("chatgpt", "claude", "claude-code", "codex"):
if provider in (prov, "all"):
providers_to_sync.append(prov)
@@ -1434,6 +1814,8 @@ def joplin(ctx: click.Context, provider: str, project_filter: str | None, dry_ru
finally:
progress.advance(task)
ctx.obj["last_joplin_summary"] = summary
if not dry_run:
_print_joplin_summary(summary)
+33 -4
View File
@@ -131,10 +131,7 @@ def resolve_media(
logger.warning(
"[media] Could not download %s: %s", ref[:60], e.original
)
reason = "expired-or-missing" if "404" in str(e.original) or "not found" in str(
e.original
).lower() else "download-error"
report.record_media_failed(reason)
report.record_media_failed(_classify_failure(e))
continue
ext = _pick_extension(mime, file_name)
@@ -153,6 +150,38 @@ def resolve_media(
return downloaded
def _classify_failure(error: ProviderError) -> str:
"""Bucket a download failure for the run summary.
The buckets separate "the asset is gone" from "the asset is there but we
were refused" — different causes, different fixes, so lumping both into
download-error hides which one you have.
Note that a bare 403 from ChatGPT usually means *gone*, not *refused*:
its download endpoint reports a deleted upload as 403 Forbidden. The
provider confirms that against /files/{id} before raising, so a failure
that reaches here still saying "forbidden" is one where the file record
survives.
``forbidden`` does NOT mean recoverable. Investigated exhaustively
2026-08-17 (19 assets): those records report ``state: "ready"`` and carry
a ``library_file_id``, yet /files/{id}/download — the endpoint that mints
the signed estuary/content URL the web UI itself fetches — refuses them,
and the Library id is rejected as ``file_not_found``. Decisively, the same
images render blank in ChatGPT's own UI. Nothing was withheld from us; the
bytes are gone and only the metadata survived, apparently from a Library
migration dated 2026-07-14. Do not spend another afternoon on it: no
header, scope, namespace or id form reaches these. The bucket stays
separate only because the two failures look different on the wire.
"""
detail = str(error.original).lower()
if "404" in detail or "not found" in detail or "no longer exists" in detail:
return "expired-or-missing"
if "403" in detail or "forbidden" in detail:
return "forbidden"
return "download-error"
def _safe_asset_name(provider, ref: str) -> str | None:
"""A stable, filesystem-safe name for the asset (the provider file ID)."""
parser = getattr(provider, "parse_asset_file_id", None)
+179
View File
@@ -0,0 +1,179 @@
"""Push notifications for unattended runs, via ntfy (https://ntfy.sh).
A scheduled archive is invisible by design: the exporter logs to
``cache/logs/exporter.log``, systemd logs to the journal and Task Scheduler
records an exit code, but all three are *pull* — you only learn a run failed by
going to look. This module pushes a one-line result instead.
Configured entirely from the environment (see ``.env.example``):
NTFY_TOPIC topic name; unset disables notifications entirely
NTFY_SERVER default https://ntfy.sh — set to your own host if self-hosting
NTFY_TOKEN optional bearer token, for access-controlled topics
NTFY_NOTIFY always (default) | failure | off
**Payload is deliberately counts-only.** A public ntfy topic is readable by
anyone who guesses the name — there is no per-topic secret unless you add
NTFY_TOKEN against a server that enforces it. Conversation titles are sensitive
(they are the first line of what you asked), so nothing but provider names and
numbers ever goes into a notification. The machine name is included because two
machines archive into one topic and "2 exported" is meaningless without knowing
which box it came from.
Never fatal: notification is a courtesy, not part of the archive. Every failure
path here logs a warning and returns False, so a down ntfy server can't fail a
run that actually captured your data.
"""
import logging
import os
import socket
import requests
logger = logging.getLogger(__name__)
DEFAULT_SERVER = "https://ntfy.sh"
NOTIFY_ALWAYS = "always"
NOTIFY_FAILURE = "failure"
NOTIFY_OFF = "off"
VALID_NOTIFY_POLICIES = {NOTIFY_ALWAYS, NOTIFY_FAILURE, NOTIFY_OFF}
# ntfy request timeout (connect, read). Short: a notification is never worth
# holding a scheduled run open for.
_TIMEOUT = (5, 10)
def resolve_notify_policy() -> str:
"""Read NTFY_NOTIFY from the environment, defaulting to ``always``."""
value = os.getenv("NTFY_NOTIFY", "").strip().lower()
if not value:
return NOTIFY_ALWAYS
if value not in VALID_NOTIFY_POLICIES:
logger.warning(
"NTFY_NOTIFY=%r is invalid (expected always|failure|off) — using 'always'.",
value,
)
return NOTIFY_ALWAYS
return value
def is_configured() -> bool:
"""True when a topic is set and the policy isn't ``off``."""
return bool(os.getenv("NTFY_TOPIC", "").strip()) and resolve_notify_policy() != NOTIFY_OFF
def machine_name() -> str:
"""Short hostname — two machines share one topic, so this is load-bearing."""
try:
return socket.gethostname().split(".")[0] or "unknown-host"
except Exception:
return "unknown-host"
def _ascii_header(value: str) -> str:
"""Reduce a header value to plain ASCII.
HTTP headers are latin-1 at best, and requests raises on anything outside
it — an em dash in a title is enough to lose the notification entirely
(observed 2026-08-18). The body has no such limit: it is sent as UTF-8
bytes, so only header values need flattening.
"""
replacements = {"\u2014": "-", "\u2013": "-", "\u2018": "'", "\u2019": "'",
"\u201c": '"', "\u201d": '"', "\u2026": "..."}
for bad, good in replacements.items():
value = value.replace(bad, good)
return value.encode("ascii", "replace").decode("ascii")
def send(
title: str,
message: str,
tags: str = "",
priority: str = "",
) -> bool:
"""POST a notification to ntfy. Returns True on success, never raises.
``tags`` is a comma-separated list of ntfy tag names (emoji shortcodes are
rendered as emoji); ``priority`` is one of min/low/default/high/urgent.
"""
topic = os.getenv("NTFY_TOPIC", "").strip()
if not topic:
return False
server = (os.getenv("NTFY_SERVER", "").strip() or DEFAULT_SERVER).rstrip("/")
url = f"{server}/{topic}"
headers = {"Title": _ascii_header(title)}
if tags:
headers["Tags"] = _ascii_header(tags)
if priority:
headers["Priority"] = _ascii_header(priority)
token = os.getenv("NTFY_TOKEN", "").strip()
if token:
headers["Authorization"] = f"Bearer {token}"
try:
resp = requests.post(
url,
data=message.encode("utf-8"),
headers=headers,
timeout=_TIMEOUT,
)
resp.raise_for_status()
logger.debug("[notify] Sent to %s", url)
return True
except Exception as e:
# Deliberately broad: a notification must never fail an archive run.
logger.warning("[notify] Could not send to %s: %s", url, e)
return False
def format_summary(
export_summary: dict | None,
joplin_summary: dict | None,
failures: list[str] | None,
) -> tuple[str, str, str, str]:
"""Build ``(title, message, tags, priority)`` for a completed sync run.
Counts only — see the module docstring on why no titles appear here.
"""
failures = failures or []
host = machine_name()
lines: list[str] = []
total_exported = 0
for prov, counts in (export_summary or {}).items():
exported = counts.get("exported", 0)
skipped = counts.get("skipped", 0)
failed = counts.get("failed", 0)
total_exported += exported
line = f"{prov}: {exported} exported, {skipped} up to date"
if failed:
line += f", {failed} FAILED"
lines.append(line)
synced = sum(
counts.get("created", 0) + counts.get("updated", 0)
for counts in (joplin_summary or {}).values()
)
if joplin_summary:
lines.append(f"joplin: {synced} note(s) created/updated")
if failures:
lines.append("")
lines.extend(failures)
title = f"AI archive FAILED - {host}"
tags = "rotating_light"
priority = "high"
else:
title = f"AI archive OK - {host}"
# A run that captured something is more interesting than a quiet one.
tags = "white_check_mark" if total_exported else "zzz"
priority = "default" if total_exported else "low"
if not lines:
lines = ["nothing to do"]
return title, "\n".join(lines), tags, priority
+42 -1
View File
@@ -33,6 +33,9 @@ logger = logging.getLogger(__name__)
# Request timeouts (connect, read) in seconds
REQUEST_TIMEOUT = (10, 30)
# Longest error-body excerpt to carry into a ProviderError message.
_ERROR_BODY_CHARS = 300
# Retry configuration
MAX_RETRIES = 3
BACKOFF_BASE = 2.0
@@ -115,6 +118,32 @@ def resolve_request_delay() -> float:
return DEFAULT_REQUEST_DELAY
return value
def _describe_error_body(response: Any) -> str:
"""Summarise an error response body for a log line.
Providers explain 4xx in the body (ChatGPT uses ``detail``), so a bare
status code is not a diagnosis. Prefers ``detail``/``error``/``message``,
falls back to a truncated raw excerpt, and redacts before it is logged.
"""
try:
body = response.json()
except Exception:
try:
text = (response.text or "").strip()
except Exception:
return "no response body"
if not text:
return "empty response body"
return f"body: {text[:_ERROR_BODY_CHARS]}"
if isinstance(body, dict):
for key in ("detail", "error", "message"):
if key in body:
value = redact_secrets(body[key])
return f"{key}: {str(value)[:_ERROR_BODY_CHARS]}"
return f"body: {str(redact_secrets(body))[:_ERROR_BODY_CHARS]}"
# Realistic Chrome User-Agent
USER_AGENT = (
"Mozilla/5.0 (X11; Linux x86_64) "
@@ -384,7 +413,19 @@ class BaseProvider(ABC):
continue
# ── Other HTTP errors ──────────────────────────────────────
response.raise_for_status()
# Raise with the response body attached. curl_cffi formats
# raise_for_status() as "HTTP Error {code}: {reason}", and
# HTTP/2 carries no reason phrase — so the bare exception
# reads "HTTP Error 403:" and says nothing about the cause.
# The provider's JSON `detail` is the only explanation there is.
if not response.ok:
raise ProviderError(
self.provider_name,
f"{method} {url}",
RuntimeError(
f"HTTP {response.status_code} — {_describe_error_body(response)}"
),
)
# ── Success ────────────────────────────────────────────────
body = response.json()
+98 -5
View File
@@ -401,6 +401,34 @@ class ChatGPTProvider(BaseProvider):
self._project_name_cache[project_id] = name
return name
def _note_unconfigured_project(self, gizmo_id: str, name: str) -> None:
"""Report a project discovered from a conversation but absent from config.
Attribution now works without CHATGPT_PROJECT_IDS, but the listing pass
still uses it: project-only conversations never appear in the default
listing, so an unconfigured project's chats are exported only if they
surface some other way. Naming the gap once per run is the difference
between "some chats are missing" and knowing which line to add.
"""
configured = getattr(self, "_project_ids", None) or []
if gizmo_id in configured:
return
seen = getattr(self, "_unconfigured_projects", None)
if seen is None:
seen = set()
self._unconfigured_projects = seen
if gizmo_id in seen:
return
seen.add(gizmo_id)
logger.info(
"[chatgpt] Project '%s' (%s) is not in CHATGPT_PROJECT_IDS — "
"attribution resolved from the conversation, but conversations "
"that live only in this project are not being listed. Add it to "
"CHATGPT_PROJECT_IDS to fetch them.",
name,
gizmo_id,
)
def list_project_conversations(
self, project_id: str, cursor: str = "0"
) -> tuple[list[dict], str | None]:
@@ -687,6 +715,15 @@ class ChatGPTProvider(BaseProvider):
signed ``download_url``; fetching that yields the bytes. Expired
assets (e.g. old generated images) 404 on the first hop.
A *deleted* asset does not 404 on the first hop — it answers
``403 {"detail":"Forbidden"}``, which reads like a permissions
problem and isn't one. Measured 2026-08-17 across 18 such assets:
every one returned ``404 {"detail":"File not found"}`` on
``/files/{id}`` while assets that downloaded fine returned 200 on
both, in the same session. So a 403 here is confirmed against the
metadata endpoint before it is reported, and a missing record is
called what it is rather than a refusal.
Returns:
(content_bytes, mime_type_or_None, file_name_or_None)
@@ -702,7 +739,21 @@ class ChatGPTProvider(BaseProvider):
ValueError(f"Unrecognised asset reference: {ref[:80]}"),
)
try:
meta = self._make_request("GET", f"{BASE_URL}/files/{file_id}/download")
except ProviderError as e:
if "HTTP 403" in str(e.original) and self._asset_record_missing(file_id):
raise ProviderError(
self.provider_name,
f"download_asset({file_id})",
RuntimeError(
"Asset no longer exists — HTTP 404 'File not found' on "
"/files/{id}. The upload was deleted or expired "
"server-side; it is not recoverable from ChatGPT."
),
) from e
raise
download_url = meta.get("download_url")
if not download_url:
raise ProviderError(
@@ -729,6 +780,26 @@ class ChatGPTProvider(BaseProvider):
)
return resp.content, mime, file_name
def _asset_record_missing(self, file_id: str) -> bool:
"""True if ``/files/{id}`` reports the asset gone.
Runs only on the 403 path, so it costs one extra request per failed
asset and none per successful one. Any other outcome — 200, a network
error, an unparseable body — returns False, leaving the original 403
to be reported as-is rather than guessing that it means deletion.
"""
try:
self._pace()
resp = self._session.request(
"GET", f"{BASE_URL}/files/{file_id}", timeout=REQUEST_TIMEOUT
)
except Exception as e: # noqa: BLE001 - a failed probe must not mask the 403
logger.debug(
"[chatgpt] Existence probe for %s failed: %s", file_id, e
)
return False
return resp.status_code == 404
# ------------------------------------------------------------------
# Normalization
# ------------------------------------------------------------------
@@ -754,15 +825,37 @@ class ChatGPTProvider(BaseProvider):
updated_at = _ts_to_iso(raw.get("update_time"))
# Prefer _project_name annotation injected from the listing summary
# (propagated by the export loop). Fall back to _project_map lookup.
project = raw.get("_project_name") or (
self._project_map.get(conv_id) if conv_id else None
)
# (propagated by the export loop), then the _project_map built from
# CHATGPT_PROJECT_IDS, then the conversation's own gizmo_id.
#
# That last fallback matters: both earlier sources are limited to the
# projects the user configured, so a conversation in an unconfigured
# project still exports (it appears in the default listing) but landed
# under no-project. The detail payload names its own project, so read
# it from there and the attribution stays correct without the user
# having to maintain a list. Observed 2026-08-17: a conversation whose
# gizmo_id was absent from CHATGPT_PROJECT_IDS filed under no-project.
source = "_project_name"
project = raw.get("_project_name")
if not project and conv_id:
project = self._project_map.get(conv_id)
source = "_project_map"
if not project:
gizmo_id = raw.get("gizmo_id")
# gizmo_type distinguishes a project from a custom GPT; only a
# project should become a folder. Absent on older payloads, so
# treat "unknown but has a project-shaped id" as a project.
if gizmo_id and str(gizmo_id).startswith("g-p-"):
project = self._fetch_project_name(gizmo_id)
source = "gizmo_id"
if conv_id:
self._project_map[conv_id] = project
self._note_unconfigured_project(gizmo_id, project)
logger.debug(
"[chatgpt] normalize_conversation[%s]: project=%r (source=%s)",
conv_id[:8] if conv_id else "?",
project,
"_project_name" if raw.get("_project_name") else "_project_map",
source,
)
mapping: dict = raw.get("mapping", {})
+7 -33
View File
@@ -54,6 +54,7 @@ from src.blocks import (
make_unknown_block,
)
from src.loss_report import LossReport
from src.utils import git_root_name
from src.providers.base import (
BaseProvider,
HIDDEN_CONTENT_FULL,
@@ -395,7 +396,11 @@ def _parse_jsonl(path: Path) -> list[dict]:
except OSError as e:
logger.warning("[claude-code] Could not read %s: %s", path, e)
return records
for line in text.splitlines():
# split("\n"), not splitlines(): splitlines() also breaks on U+0085,
# U+2028/9 and friends, which are legal *inside* a JSON string. A NEL in
# captured command output shreds one record into unparseable fragments and
# loses it silently (observed 2026-08-18 in a real Codex rollout).
for line in text.split("\n"):
line = line.strip()
if not line:
continue
@@ -458,37 +463,6 @@ def _ignored_repos() -> set[str]:
return {s.strip() for s in env.split(",") if s.strip()}
# dir Path → git-repo name it belongs to (or None). Process-wide; the working
# tree doesn't change under us mid-run, so caching walked dirs is safe.
_GIT_ROOT_CACHE: dict[Path, str | None] = {}
_CACHE_MISS = object()
def _git_root_name(path: Path, max_steps: int = 25) -> str | None:
"""Name of the git repo ``path`` lives in — nearest ancestor with ``.git``.
Walks up from ``path`` until a ``.git`` entry is found (returns that dir's
basename) or the filesystem root is reached (returns ``None``). Disk-based:
a path in no git repo, or a repo no longer on disk, yields ``None``.
"""
cur = path
for _ in range(max_steps):
cached = _GIT_ROOT_CACHE.get(cur, _CACHE_MISS)
if cached is not _CACHE_MISS:
return cached
try:
if (cur / ".git").exists():
_GIT_ROOT_CACHE[cur] = cur.name
return cur.name
except OSError:
break
if cur.parent == cur: # filesystem root
break
cur = cur.parent
_GIT_ROOT_CACHE[path] = None
return None
# Absolute path-like tokens inside Bash command strings (file_path/path keys are
# matched directly). git-root resolution short-circuits on non-repo paths.
_ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+")
@@ -521,7 +495,7 @@ def _repos_touched(
if p in seen:
return
seen.add(p)
name = _git_root_name(Path(p))
name = git_root_name(Path(p))
if name and not name.startswith(".") and name not in ignore:
counts[name] += 1
+881
View File
@@ -0,0 +1,881 @@
"""Codex CLI session provider — archives local agent transcripts.
Reads JSONL rollout files from ``~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl``
(override with ``CODEX_DIR``). Like Claude Code: no tokens, no rate limits, no
ToS risk — the data is local and Codex may prune it.
Local-only, deliberately
------------------------
Codex Cloud tasks (``codex cloud``) live server-side at
``https://chatgpt.com/backend-api/api/codex/tasks{,/list}`` — the same host and
``/backend-api`` root the ChatGPT provider already speaks. **CLI sessions are
never uploaded there**, so the cloud API is not an alternative source for these
transcripts and this provider does not talk to the network. Verified 2026-08-18
against Codex 0.147.0. If ``codex cloud exec`` ever enters regular use, those
transcripts *would* be cloud-only and would need a separate provider.
Two representations, one file (measured 2026-08-18 over 7 sessions / 0.147.0)
----------------------------------------------------------------------------
Every rollout line is ``{timestamp, ordinal, type, payload}``. Dialogue appears
twice, in two different shapes, and we parse the **typed** one:
* ``response_item`` — the model-facing wire format (mirrors the OpenAI Responses
API). Tool calls arrive as *JavaScript source* because Codex's ``exec`` tool is
code-mode::
const r = await tools.exec_command({"cmd":"git status","workdir":"/x", …});
text(r.output);
Exactly one ``tools.*`` call per invocation; three functions observed:
``exec_command`` (180), ``web__run`` (15), ``apply_patch`` (13).
* ``event_msg`` / ``item_completed`` — Codex's own typed items, already decoded:
``UserMessage``, ``AgentMessage``, ``Reasoning``, ``CommandExecution``,
``FileChange``, ``Extension``, ``ContextCompaction``.
The typed layer wins on every axis that matters here. It is 1:1 with the raw
layer for prose (91 ``AgentMessage`` ↔ 91 assistant messages, same ids; 345
``Reasoning`` ↔ 345), it hands us structured command/exit-code/output fields
instead of JS we would have to regex, and it pre-filters harness plumbing for
free: all 51 ``developer``-role messages (skills manifests, ``<multi_agent_mode>``,
"Approved command prefix saved") plus the 7 ``# AGENTS.md instructions…``
injections and 1 ``<environment_context>`` have no typed item. That is the same
noise ``claude_code._HARNESS_TAG_RE`` strips by hand.
Its one weakness: it records what *ran*, not what was *attempted*. 26 of 180
``exec_command`` calls produced no ``CommandExecution`` item — 14 sandbox launch
failures (``bwrap: loopback: Failed RTM_NEWADDR``), 6 user aborts ("aborted by
user after 504.6s"), ~5 still running at turn end, 1 "Script failed". So we read
the raw layer *only* to count attempts, and the collapsed placeholder reports the
shortfall ("3 did not complete") rather than silently under-reporting. ``wait``
function calls (50) are process polls, not attempts, and are not counted.
Reasoning is unrecoverable
--------------------------
All 345 reasoning items carry ``encrypted_content``; ``summary`` is ``[]`` in the
raw layer and ``summary_text``/``raw_content`` are empty in the typed layer, in
every session. Codex does not persist readable reasoning locally. Thinking is
therefore always dropped and counted — the same end state as the Claude Code
policy (decision 2026-06-12), but by necessity rather than by choice, so even
``full`` cannot surface it.
Why the sidecar SQLite is not read
----------------------------------
``~/.codex/state_5.sqlite`` carries a ``threads`` table (title, cwd, model,
tokens_used, rollout_path) and ``thread_history_1.sqlite`` a projection of the
items — but ``thread_history_projection_state`` tracks a byte offset *into the
rollout file*, i.e. the JSONL is canonical and SQLite is derived. Its ``title``
is just the first user message truncated (identical to ``first_user_message`` and
``preview``), so it offers nothing the JSONL lacks, and its filename carries a
schema version that will churn. We read the files.
Rendering follows the EXPORTER_HIDDEN_CONTENT policy: prose-only by default
(dialogue kept, tool traffic collapsed to one grouped placeholder per activity
run, reasoning dropped); ``full`` keeps the decoded tool calls and their output.
Subagents: Codex 0.147.0's ``thread_spawn_edges`` table exists but is empty and no
sub-transcripts were observed, so there is no subagent folding here (contrast
``claude_code._load_subagents``). If spawned agents start appearing they will
arrive as new item types and land in the loss report as unknowns.
"""
import json
import logging
import os
import re
from collections import Counter
from datetime import datetime, timezone
from pathlib import Path
from src.blocks import (
COLLAPSED_KIND_HIDDEN_CONTEXT,
COLLAPSED_KIND_TOOL_DUMP,
UNKNOWN_REASON_UNKNOWN_TYPE,
make_collapsed_block,
make_text_block,
make_tool_result_block,
make_tool_use_block,
make_unknown_block,
)
from src.loss_report import LossReport
from src.providers.base import (
BaseProvider,
HIDDEN_CONTENT_FULL,
ProviderError,
VALID_HIDDEN_CONTENT_POLICIES,
resolve_hidden_content_policy,
)
from src.utils import git_root_name
logger = logging.getLogger(__name__)
DEFAULT_SESSIONS_DIR = "~/.codex/sessions"
# rollout-<ISO-ish timestamp>-<uuid>.jsonl — the trailing UUID is the thread id.
_ROLLOUT_RE = re.compile(
r"^rollout-\d{4}-\d{2}-\d{2}T[\d-]+-"
r"([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12})$"
)
# The single tools.* call inside an `exec` custom_tool_call's JavaScript body.
_TOOLS_CALL_RE = re.compile(r"tools\.([A-Za-z_][\w.]*)\s*\(")
# Typed items that represent tool traffic, mapped to the label used in the
# collapsed placeholder. These labels must match the ones derived from the raw
# layer in _attempt_label, or the attempted-vs-completed delta is meaningless.
_TOOL_ITEM_LABELS = {
"CommandExecution": "exec_command",
"FileChange": "apply_patch",
# Extension is labelled by its `kind` (e.g. "web.search"); see _tool_label.
}
# function_call names that poll an already-running process rather than starting
# new work. Counting them as attempts would inflate the shortfall.
_POLLING_FUNCTIONS = {"wait"}
# How many distinct tool names to list in a collapsed-activity placeholder.
_TOOL_NAMES_SHOWN = 4
# Harness-injected user text. The typed layer already omits these (they have no
# UserMessage item), so this is a belt-and-braces guard for other Codex versions.
_HARNESS_USER_RE = re.compile(
r"^\s*(?:#\s*AGENTS\.md instructions for\b|<environment_context>)",
)
def resolve_roots(sessions_dir=None) -> list[Path]:
"""Resolve the ordered list of Codex ``sessions/`` roots to scan.
Precedence:
1. Explicit ``sessions_dir`` (a single path or a list) — used by tests.
2. ``CODEX_DIR`` env, split on ``os.pathsep`` (``:``) so multiple roots can
be given; a single path (no separator) stays backward compatible. Falls
back to the default root when unset.
3. Additionally, if ``CODEX_HOME`` is set (Codex's own name for its state
directory), its ``sessions/`` subdir is appended.
Roots are expanded and de-duplicated (order preserved); existence is checked
by the caller / ``_scan``.
"""
raw: list[str]
if sessions_dir is not None:
raw = (
[str(p) for p in sessions_dir]
if isinstance(sessions_dir, (list, tuple))
else [str(sessions_dir)]
)
else:
env = os.getenv("CODEX_DIR")
raw = env.split(os.pathsep) if env else [DEFAULT_SESSIONS_DIR]
codex_home = os.getenv("CODEX_HOME")
if codex_home:
raw.append(str(Path(codex_home) / "sessions"))
roots: list[Path] = []
for p in raw:
p = p.strip()
if not p:
continue
path = Path(p).expanduser()
if path not in roots:
roots.append(path)
return roots
class CodexProvider(BaseProvider):
"""Local-file provider over Codex CLI rollout transcripts."""
provider_name = "codex"
def __init__(
self,
sessions_dir: str | Path | None = None,
hidden_content: str | None = None,
) -> None:
super().__init__()
self._sessions_dirs = resolve_roots(sessions_dir)
self._hidden_content = (
hidden_content
if hidden_content in VALID_HIDDEN_CONTENT_POLICIES
else resolve_hidden_content_policy()
)
# conv_id → rollout file path, populated by _scan()
self._path_map: dict[str, Path] = {}
# ------------------------------------------------------------------
# BaseProvider interface
# ------------------------------------------------------------------
def list_conversations(self, offset: int = 0, limit: int = 100) -> list[dict]:
full = self._scan()
return full[offset : offset + limit]
def fetch_all_conversations(self, since: datetime | None = None) -> list[dict]:
convs = self._scan()
if since is not None:
since_aware = since if since.tzinfo else since.replace(tzinfo=timezone.utc)
convs = [
c for c in convs
if datetime.fromisoformat(c["updated_at"]) >= since_aware
]
logger.info(
"[codex] Found %d session(s) under %s",
len(convs),
", ".join(str(d) for d in self._sessions_dirs),
)
return convs
def get_conversation(self, conv_id: str) -> dict:
path = self._path_map.get(conv_id)
if path is None:
# Direct call without a prior listing (e.g. tests) — scan first.
self._scan()
path = self._path_map.get(conv_id)
if path is None or not path.exists():
raise ProviderError(
self.provider_name,
f"get_conversation({conv_id[:8]})",
FileNotFoundError(f"No rollout file for id {conv_id}"),
)
records = _parse_jsonl(path)
mtime = datetime.fromtimestamp(path.stat().st_mtime, tz=timezone.utc)
return {
"id": conv_id,
"_path": str(path),
"_records": records,
# Listing and normalized updated_at must match, or the cache
# staleness comparison would re-export every session every run.
"_mtime_iso": mtime.isoformat(),
}
def normalize_conversation(self, raw: dict, loss_report: LossReport | None = None) -> dict:
report = loss_report if loss_report is not None else LossReport()
policy = getattr(self, "_hidden_content", None) or resolve_hidden_content_policy()
conv_id = raw.get("id") or ""
records: list[dict] = raw.get("_records") or []
title = _extract_title(records)
# Append the repos this session touched, e.g. "… [repo-a, repo-b]".
# Codex sessions are commonly all launched from one workspace root, so
# without this every session lands in the same notebook with no way to
# tell them apart.
launch_cwd = _extract_launch_cwd(records)
if launch_cwd:
repos = _repos_touched(records, launch_cwd)
if repos:
title = f"{title} [{', '.join(repos)}]"
project = _extract_project(records)
created_at = _extract_created_at(records)
updated_at = raw.get("_mtime_iso") or next(
(r.get("timestamp") for r in reversed(records) if r.get("timestamp")), ""
)
messages = _extract_messages(records, conv_id, report, policy)
for _ in messages:
report.record_message()
report.record_conversation()
return {
"id": conv_id,
"title": title,
"provider": self.provider_name,
"project": project,
"created_at": created_at or "",
"updated_at": updated_at or "",
"message_count": len(messages),
"messages": messages,
}
# ------------------------------------------------------------------
# Scanning
# ------------------------------------------------------------------
def _scan(self) -> list[dict]:
existing = [d for d in self._sessions_dirs if d.is_dir()]
if not existing:
logger.warning(
"[codex] No sessions directory exists: %s",
", ".join(str(d) for d in self._sessions_dirs),
)
return []
# conv_id → (mtime, conv dict). If the same thread id appears under two
# roots (e.g. a live dir and a backup), the newer-mtime copy wins.
by_id: dict[str, tuple[float, dict]] = {}
for root in existing:
# Rollouts are filed under YYYY/MM/DD; rglob keeps us agnostic to
# that layout in case Codex reorganises it.
for session_file in sorted(root.rglob("rollout-*.jsonl")):
match = _ROLLOUT_RE.match(session_file.stem)
if not match:
logger.debug("[codex] Skipping unrecognised filename %s", session_file.name)
continue
try:
stat = session_file.stat()
except OSError:
continue
if stat.st_size == 0:
continue
conv_id = match.group(1)
prev = by_id.get(conv_id)
if prev is not None and prev[0] >= stat.st_mtime:
logger.debug(
"[codex] Duplicate session %s in %s; keeping newer copy",
conv_id[:8], root,
)
continue
self._path_map[conv_id] = session_file
title, project, created = _read_session_meta(session_file)
by_id[conv_id] = (
stat.st_mtime,
{
"id": conv_id,
"title": title,
"project": project,
# The --project filter and dry-run table read the
# listing dict, not the normalized conversation.
"_project_name": project,
"created_at": created,
"updated_at": datetime.fromtimestamp(
stat.st_mtime, tz=timezone.utc
).isoformat(),
"_path": str(session_file),
},
)
# Deterministic order: by creation date bucket, then thread id.
return sorted(
(conv for _, conv in by_id.values()),
key=lambda c: (c["created_at"], c["id"]),
)
# ---------------------------------------------------------------------------
# Internal helpers
# ---------------------------------------------------------------------------
def _parse_jsonl(path: Path) -> list[dict]:
"""Read a JSONL file into records, tolerating (and logging) bad lines."""
records: list[dict] = []
bad_lines = 0
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
logger.warning("[codex] Could not read %s: %s", path, e)
return records
# split("\n"), not splitlines(): splitlines() also breaks on U+0085,
# U+2028/9 and friends, which are legal *inside* a JSON string. A NEL in
# captured command output shreds one record into unparseable fragments and
# loses it silently (observed 2026-08-18 in a real Codex rollout).
for line in text.split("\n"):
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError:
bad_lines += 1
if bad_lines:
logger.warning("[codex] %s: skipped %d unparseable line(s)", path.name, bad_lines)
return records
def _item(rec: dict) -> dict | None:
"""Return the typed item from an ``event_msg``/``item_completed`` record."""
if rec.get("type") != "event_msg":
return None
payload = rec.get("payload") or {}
if payload.get("type") != "item_completed":
return None
item = payload.get("item")
return item if isinstance(item, dict) else None
def _item_kind(item: dict) -> str:
"""Typed item discriminator. 0.147.0 uses ``type``; ``item_type`` is a hedge."""
return str(item.get("item_type") or item.get("type") or "")
def _item_text(item: dict) -> str:
"""Concatenate the text of a UserMessage / AgentMessage item.
The two disagree on case — ``UserMessage`` blocks are ``"text"`` and
``AgentMessage`` blocks are ``"Text"`` — so the comparison is case-folded.
"""
content = item.get("content")
if isinstance(content, str):
return content.strip()
if not isinstance(content, list):
return ""
parts = []
for block in content:
if isinstance(block, dict) and str(block.get("type", "")).lower() == "text":
text = block.get("text")
if isinstance(text, str):
parts.append(text)
return "".join(parts).strip()
def _read_session_meta(path: Path) -> tuple[str, str | None, str]:
"""Light single-pass scan for listing metadata: (title, project, created_at).
Substring guards keep this cheap — only candidate lines are JSON-parsed, and
the scan stops as soon as the title is found (the first UserMessage is
usually within the first few dozen lines of a multi-megabyte file).
"""
title = ""
project: str | None = None
created = ""
try:
with path.open(encoding="utf-8") as fh:
for line in fh:
if not created and '"session_meta"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
payload = rec.get("payload") or {}
created = str(payload.get("timestamp") or rec.get("timestamp") or "")
cwd = payload.get("cwd")
if cwd:
project = Path(cwd).name or None
continue
if not title and '"UserMessage"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
item = _item(rec)
if item is None or _item_kind(item) != "UserMessage":
continue
text = _item_text(item)
if text and not _HARNESS_USER_RE.match(text):
title = text[:80]
break
except OSError as e:
logger.warning("[codex] Could not read %s: %s", path, e)
return title or "Untitled session", project, created
def _extract_title(records: list[dict]) -> str:
"""Codex has no AI-generated title — the first real user prompt is the title.
(``state_5.sqlite``'s ``title`` column is this same string truncated; see the
module docstring on why the sidecar DB is not consulted.)
"""
for rec in records:
item = _item(rec)
if item is None or _item_kind(item) != "UserMessage":
continue
text = _item_text(item)
if text and not _HARNESS_USER_RE.match(text):
return text[:80]
return "Untitled session"
def _session_meta_payload(records: list[dict]) -> dict:
for rec in records:
if rec.get("type") == "session_meta":
payload = rec.get("payload")
if isinstance(payload, dict):
return payload
return {}
def _extract_launch_cwd(records: list[dict]) -> str | None:
"""The session's working directory (constant per session)."""
cwd = _session_meta_payload(records).get("cwd")
return cwd if isinstance(cwd, str) and cwd else None
def _extract_project(records: list[dict]) -> str | None:
"""Project = basename of the session's working directory."""
cwd = _extract_launch_cwd(records)
if cwd:
return Path(cwd).name or None
return None
def _extract_created_at(records: list[dict]) -> str:
"""Session start time.
Prefers ``session_meta.payload.timestamp`` — the outer line ``timestamp`` is
when the record was *flushed*, which can trail the true start by minutes.
"""
payload = _session_meta_payload(records)
ts = payload.get("timestamp")
if isinstance(ts, str) and ts:
return ts
return next((r.get("timestamp") for r in records if r.get("timestamp")), "") or ""
def _ignored_repos() -> set[str]:
"""Optional ignore-list: git repos to never tag (comma-separated names)."""
env = os.getenv("CODEX_REPO_TAG_IGNORE", "")
return {s.strip() for s in env.split(",") if s.strip()}
def _strip_file_uri(path: str) -> str:
"""``CommandExecution.cwd`` is a ``file://`` URI; every other path is plain."""
return path[7:] if path.startswith("file://") else path
# Absolute path-like tokens inside command strings.
_ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+")
def _repos_touched(
records: list[dict], launch_cwd: str, cap: int = 3, ignore: set[str] | None = None
) -> list[str]:
"""Repos a session touched, for the title's ``[repo-a, repo-b]`` tag.
A file's repo is the git repository it lives in (nearest ancestor with a
``.git``). Paths come from the typed items: ``FileChange.changes`` keys
(absolute), and ``CommandExecution``'s argv plus its ``cwd``. Ordered by
touch frequency, capped with a trailing ``…``.
"""
ignore = _ignored_repos() if ignore is None else ignore
counts: Counter = Counter()
seen: set[str] = set()
def note(p) -> None:
if not isinstance(p, str) or not p:
return
p = _strip_file_uri(p)
if not p.startswith("/"): # resolve relative paths against the launch cwd
if not launch_cwd:
return
p = str(Path(launch_cwd) / p)
if p in seen:
return
seen.add(p)
name = git_root_name(Path(p))
if name and not name.startswith(".") and name not in ignore:
counts[name] += 1
for rec in records:
item = _item(rec)
if item is None:
continue
kind = _item_kind(item)
if kind == "FileChange":
changes = item.get("changes")
if isinstance(changes, dict):
for file_path in changes:
note(file_path)
elif kind == "CommandExecution":
note(item.get("cwd"))
command = item.get("command")
if isinstance(command, list) and command:
tail = command[-1]
if isinstance(tail, str):
for m in _ABS_PATH_RE.finditer(tail):
note(m.group(0))
ordered = [name for name, _ in counts.most_common()]
if len(ordered) > cap:
return ordered[:cap] + ["…"]
return ordered
def _tool_label(item: dict, kind: str) -> str:
"""Placeholder label for a completed tool item.
``Extension`` covers everything routed through the model's own extensions
(``web.search`` so far), so its ``kind`` field is the useful name.
"""
if kind == "Extension":
return str(item.get("kind") or "extension")
return _TOOL_ITEM_LABELS.get(kind, kind)
def _attempt_label(payload: dict) -> str | None:
"""Label for a raw tool call, or None if it should not count as an attempt.
``exec`` custom_tool_calls carry JavaScript; the inner ``tools.<fn>`` name is
what lines up with the typed items' labels. ``wait`` polls an already-running
process and starts no new work.
"""
ptype = payload.get("type")
name = payload.get("name") or ""
if ptype == "custom_tool_call":
if name == "exec":
match = _TOOLS_CALL_RE.search(payload.get("input") or "")
if match:
# tools.web__run → the Extension item calls itself "web.search";
# both are "one web call", which is all the count claims.
return match.group(1)
return "exec"
return name or "tool"
if ptype == "function_call":
if name in _POLLING_FUNCTIONS:
return None
return name or "function"
return None
def _command_string(item: dict) -> str:
"""Human-readable command from a ``CommandExecution`` argv list.
The argv is ``["/bin/bash", "-lc", "<script>"]``; the script is the part
worth showing.
"""
command = item.get("command")
if isinstance(command, list):
if len(command) >= 3 and command[0].endswith("sh") and command[1] in ("-lc", "-c"):
return str(command[-1])
return " ".join(str(c) for c in command)
return str(command or "")
def _extract_messages(
records: list[dict],
conv_id: str,
report: LossReport,
policy: str,
) -> list[dict]:
"""Normalize Codex rollout records into messages.
Reads the typed ``item_completed`` layer for content and the raw
``response_item`` layer only to count attempted tool calls, so a collapsed
placeholder can report calls that never produced an item (sandbox failures,
user aborts, still-running processes). See the module docstring.
"""
messages: list[dict] = []
# Pending collapsed tool activity for the current run.
pending_tools: Counter = Counter()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
def flush_pending() -> None:
nonlocal pending_bytes, pending_attempted, pending_completed
# An activity run with only failed attempts still deserves a placeholder
# — silence would imply nothing happened.
shortfall = max(pending_attempted - pending_completed, 0)
if not pending_tools and not shortfall:
return
if policy == HIDDEN_CONTENT_FULL:
# Under `full` the completed calls are already rendered as tool
# blocks, so only the shortfall is left to report — and a *collapsed*
# block would render the "omitted (set full to keep)" suffix, which
# contradicts the policy in force. Say it plainly instead.
report.record_collapsed("tool_call_incomplete", 0)
note = make_tool_result_block(
f"{shortfall} tool call(s) produced no result — the process failed to "
"launch, was aborted, or was still running when the turn ended.",
tool_name="incomplete",
is_error=True,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(note)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None,
"blocks": [note]}
)
pending_tools.clear()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
return
if pending_tools:
shown = ", ".join(
f"{name} ×{count}"
for name, count in pending_tools.most_common(_TOOL_NAMES_SHOWN)
)
if len(pending_tools) > _TOOL_NAMES_SHOWN:
shown += ", …"
origin = f"{sum(pending_tools.values())} calls: {shown}"
else:
origin = "0 calls"
if shortfall:
origin += f" (+{shortfall} did not complete)"
report.record_collapsed("tool_call_incomplete", 0)
block = make_collapsed_block(
origin=origin,
content_type="tool_activity",
size_bytes=pending_bytes,
kind=COLLAPSED_KIND_TOOL_DUMP,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(block)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None, "blocks": [block]}
)
pending_tools.clear()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
def append_message(role: str, blocks: list[dict], timestamp) -> None:
# Attach the preceding run's tool activity to the previous message
# before starting a new one, so the placeholder lands between the
# dialogue turns it actually occurred between.
flush_pending()
messages.append(
{
"role": role,
"content_type": "text",
"timestamp": timestamp,
"blocks": blocks,
}
)
for rec in records:
rec_type = rec.get("type")
payload = rec.get("payload") or {}
timestamp = rec.get("timestamp")
# ── Raw layer: attempt counting only ──────────────────────────────
if rec_type == "response_item":
if payload.get("type") in ("custom_tool_call", "function_call"):
if _attempt_label(payload) is not None:
pending_attempted += 1
continue
item = _item(rec)
if item is None:
continue
kind = _item_kind(item)
# ── Dialogue ──────────────────────────────────────────────────────
if kind in ("UserMessage", "AgentMessage"):
text = _item_text(item)
if kind == "UserMessage" and _HARNESS_USER_RE.match(text):
# Defensive: 0.147.0 gives these no typed item at all.
report.record_filtered_role("codex.harness_injection")
continue
block = make_text_block(text)
if block:
append_message(
"user" if kind == "UserMessage" else "assistant", [block], timestamp
)
continue
# ── Reasoning: encrypted at rest, nothing to keep under any policy ─
if kind == "Reasoning":
report.record_collapsed("reasoning", len(json.dumps(item, default=str)))
continue
# ── Context compaction: a visible marker, not a silent drop ────────
if kind == "ContextCompaction":
flush_pending()
marker = make_collapsed_block(
origin="context compacted — earlier turns dropped from the model's context",
content_type="context_compaction",
size_bytes=0,
kind=COLLAPSED_KIND_HIDDEN_CONTEXT,
)
report.record_collapsed("context_compaction", 0)
messages.append(
{"role": "tool", "content_type": "text", "timestamp": timestamp,
"blocks": [marker]}
)
continue
# ── Tool traffic ──────────────────────────────────────────────────
if kind in ("CommandExecution", "FileChange", "Extension"):
pending_completed += 1
label = _tool_label(item, kind)
size = len(json.dumps(item, default=str))
if policy == HIDDEN_CONTENT_FULL:
use_block, result_block = _full_tool_blocks(item, kind, label)
target = messages[-1] if messages and messages[-1]["role"] == "assistant" else None
if target is None:
messages.append(
{"role": "tool", "content_type": "text",
"timestamp": timestamp, "blocks": []}
)
target = messages[-1]
target["blocks"].append(use_block)
if result_block:
target["blocks"].append(result_block)
else:
pending_tools[label] += 1
pending_bytes += size
report.record_collapsed(label, size)
continue
# ── Anything new in a future Codex version ────────────────────────
logger.warning("[codex] Unknown item type %r in session %s", kind, conv_id[:8])
report.record_unknown(f"codex.{kind or '?'}")
append_message(
"tool",
[
make_unknown_block(
raw_type=f"codex.{kind or '?'}",
observed_keys=list(item.keys()),
reason=UNKNOWN_REASON_UNKNOWN_TYPE,
)
],
timestamp,
)
flush_pending()
return messages
def _full_tool_blocks(item: dict, kind: str, label: str) -> tuple[dict, dict | None]:
"""Decoded tool_use / tool_result blocks for EXPORTER_HIDDEN_CONTENT=full.
The typed item already carries structured fields, so this reads the command,
exit code and output directly rather than parsing the raw layer's JavaScript.
"""
if kind == "CommandExecution":
exit_code = item.get("exit_code")
use = make_tool_use_block(
label,
{
"command": _command_string(item),
"cwd": _strip_file_uri(str(item.get("cwd") or "")),
"status": item.get("status"),
"exit_code": exit_code,
},
item.get("id"),
)
output = item.get("aggregated_output") or item.get("formatted_output") or ""
result = make_tool_result_block(
str(output),
tool_name=label,
is_error=bool(exit_code not in (0, None)),
)
return use, result
if kind == "FileChange":
changes = item.get("changes") if isinstance(item.get("changes"), dict) else {}
use = make_tool_use_block(
label,
{
"files": {
path: (change.get("type") if isinstance(change, dict) else "?")
for path, change in changes.items()
},
"status": item.get("status"),
},
item.get("id"),
)
output = "\n".join(
part for part in (item.get("stdout"), item.get("stderr")) if part
)
result = make_tool_result_block(
output or f"{len(changes)} file(s) changed",
tool_name=label,
is_error=item.get("status") not in (None, "completed", "success"),
)
return use, result
# Extension (web.search and anything else routed through an extension).
use = make_tool_use_block(
label,
{"query": item.get("query"), "action": item.get("action")},
item.get("id"),
)
results = item.get("results")
result = make_tool_result_block(
json.dumps(results, default=str, indent=1) if results else "",
tool_name=label,
)
return use, result
+58 -3
View File
@@ -76,11 +76,27 @@ def build_export_path(
return base_dir.joinpath(*parts) / filename
def _is_sensitive_key(key: object) -> bool:
"""True if a mapping key names a secret.
Matches the whole key and each of its underscore/dash-separated words, so
compound names carry too: exact-match alone let ``access_token`` and
``api_key`` through into logged response bodies. Word-level matching keeps
innocent keys ("keywords", "monkey") intact.
"""
if not isinstance(key, str):
return False
lowered = key.lower()
if lowered in _SENSITIVE_KEYS:
return True
return any(part in _SENSITIVE_KEYS for part in re.split(r"[^a-z0-9]+", lowered))
def redact_secrets(data: object) -> object:
"""Recursively redact sensitive values from a dict/list for safe logging.
Keys matching _SENSITIVE_KEYS (case-insensitive) have their values
replaced with "[REDACTED]".
Keys naming a secret (see _is_sensitive_key) have their values replaced
with "[REDACTED]".
Args:
data: Any JSON-serializable object.
@@ -90,7 +106,7 @@ def redact_secrets(data: object) -> object:
"""
if isinstance(data, dict):
return {
k: "[REDACTED]" if k.lower() in _SENSITIVE_KEYS else redact_secrets(v)
k: "[REDACTED]" if _is_sensitive_key(k) else redact_secrets(v)
for k, v in data.items()
}
if isinstance(data, list):
@@ -155,3 +171,42 @@ def _parse_dt(ts: str) -> datetime:
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
# ---------------------------------------------------------------------------
# Git-repo resolution (shared by the local agent-transcript providers)
# ---------------------------------------------------------------------------
# dir Path → git-repo name it belongs to (or None). Process-wide; the working
# tree doesn't change under us mid-run, so caching walked dirs is safe.
_GIT_ROOT_CACHE: dict[Path, str | None] = {}
_CACHE_MISS = object()
def git_root_name(path: Path, max_steps: int = 25) -> str | None:
"""Name of the git repo ``path`` lives in — nearest ancestor with ``.git``.
Walks up from ``path`` until a ``.git`` entry is found (returns that dir's
basename) or the filesystem root is reached (returns ``None``). Disk-based:
a path in no git repo, or a repo no longer on disk, yields ``None``.
Used by the Claude Code and Codex providers to tag a session title with the
repos it touched, so sessions launched from a shared workspace root stay
distinguishable.
"""
cur = path
for _ in range(max_steps):
cached = _GIT_ROOT_CACHE.get(cur, _CACHE_MISS)
if cached is not _CACHE_MISS:
return cached
try:
if (cur / ".git").exists():
_GIT_ROOT_CACHE[cur] = cur.name
return cur.name
except OSError:
break
if cur.parent == cur: # filesystem root
break
cur = cur.parent
_GIT_ROOT_CACHE[path] = None
return None
+39
View File
@@ -0,0 +1,39 @@
"""Shared test fixtures.
The autouse fixture here exists because of a real incident (2026-08-18): the
`sync` CLI tests invoke the actual command, which calls `load_config()`, which
calls `load_dotenv()` — so the developer's real `.env` was loaded and its
`NTFY_TOPIC` used. Every `pytest` run fired real push notifications at the
developer's phone, including a fabricated "3 conversations failed to export"
from a fixture. Nothing appeared in the log to explain it, because the tests
pass `--no-log-file`.
The lesson generalises past ntfy: any test that exercises a command end to end
inherits whatever is in `.env` unless the environment is neutralised first.
"""
import pytest
@pytest.fixture(autouse=True)
def _no_outbound_side_effects(monkeypatch):
"""Neutralise every environment variable that could reach a real service.
Set to empty/unroutable values rather than deleted: `load_dotenv` is called
with ``override=False``, which only skips keys **already present** in the
environment. Deleting a key would let the developer's `.env` put it back.
Individual tests may still `monkeypatch.setenv` these — that is how the
notification tests point at a dead local port on purpose.
"""
# Notifications: an empty topic disables sending outright (`is_configured`
# strips and checks truthiness), and the server is pointed at a closed port
# so even a test that sets its own topic cannot reach the internet.
monkeypatch.setenv("NTFY_TOPIC", "")
monkeypatch.setenv("NTFY_SERVER", "http://127.0.0.1:9")
monkeypatch.setenv("NTFY_TOKEN", "")
# Joplin: a test that reached the developer's running desktop instance would
# create or overwrite real notes in their archive.
monkeypatch.setenv("JOPLIN_API_URL", "http://127.0.0.1:9")
monkeypatch.setenv("JOPLIN_API_TOKEN", "")
+31 -1
View File
@@ -20,7 +20,11 @@ def _write_session(tmp_path, records, project_dir="-home-jesse-myproj", name="ab
proj = tmp_path / project_dir
proj.mkdir(parents=True, exist_ok=True)
f = proj / f"{name}.jsonl"
f.write_text("\n".join(json.dumps(r) for r in records), encoding="utf-8")
# ensure_ascii=False so raw U+0085/U+2028 reach the parser — see
# TestExoticLineBreaks.
f.write_text(
"\n".join(json.dumps(r, ensure_ascii=False) for r in records), encoding="utf-8"
)
return f
@@ -501,3 +505,29 @@ class TestMultiRoot:
monkeypatch.setenv("CLAUDE_CONFIG_DIR", str(tmp_path / "cfg"))
roots = resolve_roots()
assert (tmp_path / "cfg" / "projects") in roots
class TestExoticLineBreaks:
"""Same U+0085 hazard as the Codex provider — see its TestExoticLineBreaks."""
# C0 controls (\x0b, \x0c) are excluded: JSON requires them escaped, so they
# never reach the splitter literally. These three do.
@pytest.mark.parametrize("sep", ["\x85", "
", "
"])
def test_record_with_exotic_break_survives(self, tmp_path, sep):
records = [
{
"type": "user",
"cwd": "/home/jesse/myproj",
"timestamp": "2026-05-01T10:00:00.000Z",
"message": {"role": "user", "content": f"before{sep}after"},
},
]
_write_session(tmp_path, records)
prov = ClaudeCodeProvider(projects_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation("abc-123"))
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert f"before{sep}after" in texts
+364
View File
@@ -1,5 +1,7 @@
"""CLI-level tests using Click's CliRunner — no live API calls required."""
from pathlib import Path
import pytest
from click.testing import CliRunner
@@ -297,3 +299,365 @@ class TestCanaryCommand:
)
assert result.exit_code == 1
assert "No web-API provider tokens" in result.output
class TestProjectsCommand:
"""`projects` discovers project IDs missing from CHATGPT_PROJECT_IDS."""
def _patch_provider(self, monkeypatch, summaries, details=None, names=None):
import src.providers.chatgpt as chatgpt_mod
names = names or {}
details = details or {}
class FakeProvider:
def __init__(self, **kwargs):
self._project_ids = kwargs.get("project_ids") or []
def fetch_all_conversations(self, since=None):
return summaries
def get_conversation(self, conv_id):
return details.get(conv_id, {})
def _fetch_project_name(self, gizmo_id):
return names.get(gizmo_id, gizmo_id)
monkeypatch.setattr(chatgpt_mod, "ChatGPTProvider", FakeProvider)
return FakeProvider
def _env(self, tmp_path, **extra):
# Clears the ToS gate and the first-run doctor check.
cache = Cache(tmp_path)
cache.acknowledge_tos()
cache.mark_exported("chatgpt", "dummy", {"updated_at": "2024-01-01T00:00:00Z"})
env = {
"CHATGPT_SESSION_TOKEN": "eyJtesttoken",
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
}
env.update(extra)
return env
def test_reports_project_absent_from_config(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-p-missing"}],
names={"g-p-missing": "Tech Questions"},
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects"], env=self._env(tmp_path)
)
assert result.exit_code == 0
assert "Tech Questions" in result.output
assert "g-p-missing" in result.output
assert "CHATGPT_PROJECT_IDS=g-p-missing" in result.output
def test_quiet_when_everything_is_configured(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-p-known"}],
names={"g-p-known": "Known"},
)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "projects"],
env=self._env(tmp_path, CHATGPT_PROJECT_IDS="g-p-known"),
)
assert result.exit_code == 0
assert "already configured" in result.output
def test_custom_gpt_ids_are_ignored(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-notaproject"}],
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects"], env=self._env(tmp_path)
)
assert "g-notaproject" not in result.output
def test_deep_reads_conversation_details(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1"}], # listing carries no gizmo_id
details={"c1": {"gizmo_id": "g-p-fromdetail"}},
names={"g-p-fromdetail": "Found Deep"},
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects", "--deep"], env=self._env(tmp_path)
)
assert "Found Deep" in result.output
assert "CHATGPT_PROJECT_IDS=g-p-fromdetail" in result.output
def test_without_deep_says_so_rather_than_reporting_nothing(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1"}],
details={"c1": {"gizmo_id": "g-p-fromdetail"}},
)
result = CliRunner(mix_stderr=True).invoke(
cli, ["--no-log-file", "projects"], env=self._env(tmp_path)
)
assert "--deep" in result.output
def test_write_updates_env(self, tmp_path, monkeypatch):
self._patch_provider(
monkeypatch,
summaries=[{"id": "c1", "gizmo_id": "g-p-new"}],
names={"g-p-new": "New Project"},
)
runner = CliRunner(mix_stderr=True)
with runner.isolated_filesystem(temp_dir=tmp_path) as fs:
result = runner.invoke(
cli, ["--no-log-file", "projects", "--write"], env=self._env(tmp_path)
)
assert result.exit_code == 0
env_text = (Path(fs) / ".env").read_text(encoding="utf-8")
assert "CHATGPT_PROJECT_IDS=g-p-new" in env_text
# ---------------------------------------------------------------------------
# sync command + non-interactive ToS gate
# ---------------------------------------------------------------------------
class TestSyncCommand:
"""`sync` chains export → joplin for schedulers, with a real exit code."""
def _cache(self, tmp_path) -> Cache:
cache = Cache(tmp_path)
cache.acknowledge_tos()
# Non-empty last_run so the first-run doctor gate stays out of the way.
cache.mark_exported("codex", "dummy", {"updated_at": "2024-01-01T00:00:00Z"})
return cache
def _env(self, tmp_path) -> dict:
"""A real (minimal) codex session — `export` exits 1 on no providers at
all, so an empty directory would test the wrong failure."""
import json
day = tmp_path / "sessions" / "2026" / "08" / "17"
day.mkdir(parents=True, exist_ok=True)
sid = "01a00e3f-a309-74a3-bf32-06c2cd87faa3"
records = [
{
"timestamp": "2026-08-17T05:44:06.666Z",
"type": "session_meta",
"payload": {"session_id": sid, "timestamp": "2026-08-17T05:44:06.666Z",
"cwd": str(tmp_path / "ws")},
},
{
"timestamp": "2026-08-17T05:44:09.000Z",
"type": "event_msg",
"payload": {"type": "item_completed", "item": {
"type": "UserMessage", "id": "u1",
"content": [{"type": "text", "text": "hello"}]}},
},
]
(day / f"rollout-2026-08-17T01-44-06-{sid}.jsonl").write_text(
"\n".join(json.dumps(r) for r in records), encoding="utf-8"
)
return {
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
"CODEX_DIR": str(tmp_path / "sessions"),
}
def test_skip_joplin_exits_zero(self, tmp_path):
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--skip-joplin"],
env=self._env(tmp_path),
)
assert result.exit_code == 0
assert "Skipping Joplin sync" in result.output
assert "Sync complete" in result.output
def test_joplin_optional_survives_unreachable_joplin(self, tmp_path):
"""Joplin being closed must not fail a scheduled run — the export is done."""
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--joplin-optional"],
env={**self._env(tmp_path), "JOPLIN_API_URL": "http://127.0.0.1:9"},
)
assert result.exit_code == 0
assert "Joplin sync skipped" in result.output
def test_unreachable_joplin_fails_without_the_flag(self, tmp_path):
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex"],
env={**self._env(tmp_path), "JOPLIN_API_URL": "http://127.0.0.1:9"},
)
assert result.exit_code == 1
def test_export_failures_set_nonzero_exit(self, tmp_path, monkeypatch):
"""A scheduler must be able to tell a real run from a silent no-op."""
self._cache(tmp_path)
import src.main as main_mod
real_export = main_mod.export.callback
def fake_export(*args, **kwargs):
import click
ctx = click.get_current_context()
ctx.obj["last_export_summary"] = {
"codex": {"exported": 0, "skipped": 0, "failed": 3}
}
monkeypatch.setattr(main_mod.export, "callback", fake_export)
try:
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--skip-joplin"],
env=self._env(tmp_path),
)
finally:
monkeypatch.setattr(main_mod.export, "callback", real_export)
assert result.exit_code == 1
assert "3 conversation(s) failed to export" in result.output
class TestNonInteractiveTosGate:
"""Without a TTY the gate must fail loudly, not exit 0 having done nothing."""
def test_no_tty_exits_one_with_explanation(self, tmp_path, monkeypatch):
Cache(tmp_path) # fresh cache: ToS not acknowledged
monkeypatch.setattr("sys.stdin.isatty", lambda: False)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "doctor"],
env={"CACHE_DIR": str(tmp_path), "EXPORT_DIR": str(tmp_path / "exports")},
)
assert result.exit_code == 1
assert "no terminal to" in result.output
# ---------------------------------------------------------------------------
# ntfy notifications
# ---------------------------------------------------------------------------
class TestNotifyFormatting:
"""The payload must stay counts-only: a public ntfy topic is world-readable."""
def test_success_summary(self):
from src.notify import format_summary
title, body, tags, priority = format_summary(
{"codex": {"exported": 2, "skipped": 5, "failed": 0}}, None, []
)
assert title.startswith("AI archive OK")
assert "codex: 2 exported, 5 up to date" in body
assert tags == "white_check_mark"
assert priority == "default"
def test_quiet_run_is_low_priority(self):
from src.notify import format_summary
_, _, tags, priority = format_summary(
{"codex": {"exported": 0, "skipped": 7, "failed": 0}}, None, []
)
assert tags == "zzz"
assert priority == "low"
def test_failure_summary_is_high_priority(self):
from src.notify import format_summary
title, body, tags, priority = format_summary(
{"chatgpt": {"exported": 0, "skipped": 0, "failed": 12}},
None,
["chatgpt: 12 conversation(s) failed to export"],
)
assert "FAILED" in title
assert "12 FAILED" in body
assert tags == "rotating_light"
assert priority == "high"
def test_hostname_present(self):
"""Two machines share one topic — counts are meaningless without it."""
from src.notify import format_summary, machine_name
title, _, _, _ = format_summary({}, None, [])
assert machine_name() in title
class TestNotifyHeaderEncoding:
"""HTTP headers are latin-1; an em dash in a title loses the notification."""
def test_smart_punctuation_flattened(self):
from src.notify import _ascii_header
out = _ascii_header("AI archive — don’t “fail”…")
assert out == 'AI archive - don\'t "fail"...'
out.encode("ascii") # must not raise
def test_arbitrary_unicode_survives_as_ascii(self):
from src.notify import _ascii_header
_ascii_header("héllo — 世界").encode("ascii")
class TestNotifySend:
def test_no_topic_is_a_no_op(self, monkeypatch):
from src import notify as notify_mod
monkeypatch.delenv("NTFY_TOPIC", raising=False)
assert notify_mod.send("t", "m") is False
assert notify_mod.is_configured() is False
def test_network_failure_never_raises(self, monkeypatch):
"""A down ntfy server must not fail a run that captured data."""
from src import notify as notify_mod
monkeypatch.setenv("NTFY_TOPIC", "unit-test-topic")
monkeypatch.setenv("NTFY_SERVER", "http://127.0.0.1:9")
assert notify_mod.send("t", "m") is False
def test_off_policy_disables(self, monkeypatch):
from src import notify as notify_mod
monkeypatch.setenv("NTFY_TOPIC", "unit-test-topic")
monkeypatch.setenv("NTFY_NOTIFY", "off")
assert notify_mod.is_configured() is False
class TestNoRealNotificationsDuringTests:
"""Regression guard for the 2026-08-18 incident: the suite pushed to the
developer's real ntfy topic because `sync` loads `.env` via `load_dotenv`.
Asserts the conftest neutralisation holds even though a real `.env` with a
live NTFY_TOPIC sits beside the tests.
"""
def test_notifications_are_disabled(self):
from src import notify as notify_mod
assert notify_mod.is_configured() is False
def test_dotenv_cannot_reintroduce_a_topic(self):
"""`load_dotenv(override=False)` skips keys already present — including
empty ones. Deleting the var instead of emptying it would reopen this."""
import os
from dotenv import load_dotenv
from src import notify as notify_mod
load_dotenv(override=False)
assert os.getenv("NTFY_TOPIC", "").strip() == ""
assert notify_mod.is_configured() is False
def test_send_cannot_reach_the_public_server(self, monkeypatch):
"""Even a test that sets its own topic is pinned to a dead local port."""
import os
from src import notify as notify_mod
monkeypatch.setenv("NTFY_TOPIC", "some-topic")
assert "127.0.0.1" in os.getenv("NTFY_SERVER", "")
assert notify_mod.send("t", "m") is False
+410
View File
@@ -0,0 +1,410 @@
"""Unit tests for the Codex CLI session provider.
Fixtures mirror the real 0.147.0 rollout shape observed on 2026-08-18: dialogue
carried twice (typed ``item_completed`` items plus raw ``response_item``s),
code-mode ``exec`` calls whose input is JavaScript, encrypted reasoning, and
harness-injected user messages that exist only in the raw layer.
"""
import json
import pytest
from src.blocks import (
BLOCK_TYPE_COLLAPSED,
BLOCK_TYPE_TEXT,
BLOCK_TYPE_TOOL_RESULT,
BLOCK_TYPE_TOOL_USE,
COLLAPSED_KIND_HIDDEN_CONTEXT,
)
from src.loss_report import LossReport
from src.providers.codex import CodexProvider, resolve_roots
SESSION_ID = "01a00e3f-a309-74a3-bf32-06c2cd87faa3"
FILENAME = f"rollout-2026-08-17T01-44-06-{SESSION_ID}.jsonl"
def _write_session(tmp_path, records, name=FILENAME, day="2026/08/17"):
day_dir = tmp_path / day
day_dir.mkdir(parents=True, exist_ok=True)
f = day_dir / name
# ensure_ascii=False: Codex (Rust) writes non-ASCII literally, so real
# rollouts contain raw U+0085/U+2028 inside JSON strings. Escaping them here
# would hide exactly the hazard TestExoticLineBreaks exists to catch.
f.write_text(
"\n".join(json.dumps(r, ensure_ascii=False) for r in records), encoding="utf-8"
)
return f
def _item(item_type, ts="2026-08-17T05:44:10.000Z", **fields):
# `item_type`, not `kind` — Extension items carry their own `kind` field.
return {
"timestamp": ts,
"type": "event_msg",
"payload": {"type": "item_completed", "item": {"type": item_type, **fields}},
}
def _exec_call(cmd, fn="exec_command"):
"""A raw code-mode custom_tool_call — its input is JavaScript, not JSON."""
return {
"timestamp": "2026-08-17T05:44:11.000Z",
"type": "response_item",
"payload": {
"type": "custom_tool_call",
"name": "exec",
"call_id": "call_1",
"input": (
f'const r = await tools.{fn}({{"cmd":{json.dumps(cmd)},'
f'"workdir":"/home/jesse/ws","yield_time_ms":30000}});\ntext(r.output);'
),
},
}
def _records(cwd="/home/jesse/ws"):
"""A representative session: meta, harness noise, dialogue, tool traffic."""
return [
{
"timestamp": "2026-08-17T05:49:13.629Z",
"ordinal": 0,
"type": "session_meta",
"payload": {
"session_id": SESSION_ID,
"timestamp": "2026-08-17T05:44:06.666Z",
"cwd": cwd,
"cli_version": "0.147.0",
"source": "cli",
},
},
# Harness plumbing: present in the raw layer only, exactly as 0.147.0
# writes it. Must not become a message.
{
"timestamp": "2026-08-17T05:44:07.000Z",
"type": "response_item",
"payload": {
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "# AGENTS.md instructions for /x"}],
},
},
{
"timestamp": "2026-08-17T05:44:08.000Z",
"type": "response_item",
"payload": {
"type": "message",
"role": "developer",
"content": [{"type": "input_text", "text": "<skills_instructions>…"}],
},
},
_item(
"UserMessage",
ts="2026-08-17T05:44:09.000Z",
id="u1",
content=[{"type": "text", "text": "Write the backup guide.", "text_elements": []}],
),
# Encrypted reasoning — nothing recoverable in either layer.
{
"timestamp": "2026-08-17T05:44:09.500Z",
"type": "response_item",
"payload": {"type": "reasoning", "summary": [], "encrypted_content": "gAAAA…"},
},
_item("Reasoning", id="r1", summary_text=[], raw_content=[]),
# AgentMessage content blocks are "Text" (capital T), unlike UserMessage.
_item(
"AgentMessage",
id="a1",
phase="commentary",
content=[{"type": "Text", "text": "I'll inspect the repo first."}],
),
_exec_call("ls -la"),
_item(
"CommandExecution",
id="exec-1",
process_id="123",
command=["/bin/bash", "-lc", "ls -la"],
cwd="file:///home/jesse/ws",
source="unified_exec_startup",
status="completed",
exit_code=0,
stdout="total 4\n",
aggregated_output="total 4\n",
formatted_output="total 4\n",
),
_exec_call("cat missing"),
_item(
"CommandExecution",
id="exec-2",
command=["/bin/bash", "-lc", "cat missing"],
cwd="file:///home/jesse/ws",
status="completed",
exit_code=1,
stderr="No such file\n",
aggregated_output="No such file\n",
),
# An attempt that never produced an item (sandbox failure / abort).
_exec_call("npm test"),
# A `wait` poll: not an attempt, must not inflate the shortfall.
{
"timestamp": "2026-08-17T05:44:12.000Z",
"type": "response_item",
"payload": {"type": "function_call", "name": "wait", "call_id": "call_w"},
},
_item(
"AgentMessage",
ts="2026-08-17T05:45:00.000Z",
id="a2",
phase="final_answer",
content=[{"type": "Text", "text": "Done — the guide is written."}],
),
]
class TestCodexProvider:
def test_scan_lists_session_with_metadata(self, tmp_path):
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
convs = prov.list_conversations()
assert len(convs) == 1
assert convs[0]["id"] == SESSION_ID
assert convs[0]["title"] == "Write the backup guide."
assert convs[0]["project"] == "ws"
# session_meta payload timestamp, not the (later) flush timestamp.
assert convs[0]["created_at"] == "2026-08-17T05:44:06.666Z"
def test_ignores_non_rollout_files(self, tmp_path):
_write_session(tmp_path, _records())
(tmp_path / "2026/08/17/notes.jsonl").write_text("{}", encoding="utf-8")
(tmp_path / "2026/08/17/rollout-garbage.jsonl").write_text("{}", encoding="utf-8")
assert len(CodexProvider(sessions_dir=tmp_path).list_conversations()) == 1
def test_empty_file_skipped(self, tmp_path):
_write_session(tmp_path, _records())
(tmp_path / "2026/08/16").mkdir(parents=True)
(tmp_path / "2026/08/16" / FILENAME.replace("17T01", "16T01")).write_text("")
assert len(CodexProvider(sessions_dir=tmp_path).list_conversations()) == 1
def test_missing_root_returns_empty(self, tmp_path):
prov = CodexProvider(sessions_dir=tmp_path / "nope")
assert prov.list_conversations() == []
def test_get_conversation_unknown_id_raises(self, tmp_path):
from src.providers.base import ProviderError
prov = CodexProvider(sessions_dir=tmp_path)
with pytest.raises(ProviderError):
prov.get_conversation("does-not-exist")
class TestNormalize:
def _normalized(self, tmp_path, policy="placeholder", records=None):
_write_session(tmp_path, records if records is not None else _records())
prov = CodexProvider(sessions_dir=tmp_path, hidden_content=policy)
prov.list_conversations()
report = LossReport()
return prov.normalize_conversation(prov.get_conversation(SESSION_ID), report), report
def test_dialogue_only_by_default(self, tmp_path):
conv, _ = self._normalized(tmp_path)
roles = [m["role"] for m in conv["messages"]]
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert roles[0] == "user"
assert texts == [
"Write the backup guide.",
"I'll inspect the repo first.",
"Done — the guide is written.",
]
def test_harness_injections_never_become_messages(self, tmp_path):
conv, _ = self._normalized(tmp_path)
blob = json.dumps(conv)
assert "AGENTS.md instructions" not in blob
assert "skills_instructions" not in blob
def test_reasoning_is_dropped_and_counted(self, tmp_path):
conv, report = self._normalized(tmp_path)
assert "thinking" not in json.dumps(conv)
assert "reasoning" in report.format_summary()
def test_reasoning_stays_dropped_under_full(self, tmp_path):
# Unlike Claude Code, `full` cannot surface it — it is encrypted at rest.
conv, _ = self._normalized(tmp_path, policy="full")
assert "encrypted" not in json.dumps(conv)
def test_tool_traffic_collapses_with_shortfall(self, tmp_path):
conv, _ = self._normalized(tmp_path)
collapsed = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
]
assert len(collapsed) == 1
origin = collapsed[0]["origin"]
# 2 completed of 3 attempted; the `wait` poll is not an attempt.
assert "2 calls: exec_command ×2" in origin
assert "+1 did not complete" in origin
def test_full_policy_emits_decoded_tool_blocks(self, tmp_path):
conv, _ = self._normalized(tmp_path, policy="full")
uses = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TOOL_USE
]
results = [
b for m in conv["messages"] for b in m["blocks"]
# The "incomplete" note is asserted by TestFullPolicyShortfall.
if b["type"] == BLOCK_TYPE_TOOL_RESULT and b.get("tool_name") != "incomplete"
]
assert len(uses) == 2 and len(results) == 2
# The command is read from the typed item, not parsed out of the JS.
assert uses[0]["input"]["command"] == "ls -la"
assert uses[0]["input"]["cwd"] == "/home/jesse/ws" # file:// stripped
assert results[0]["is_error"] is False
assert results[1]["is_error"] is True # exit_code 1
def test_message_count_matches(self, tmp_path):
conv, _ = self._normalized(tmp_path)
assert conv["message_count"] == len(conv["messages"])
def test_updated_at_matches_listing(self, tmp_path):
"""Cache staleness compares these two; a mismatch re-exports every run."""
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
listed = prov.list_conversations()[0]
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["updated_at"] == listed["updated_at"]
def test_context_compaction_is_visible(self, tmp_path):
records = _records() + [_item("ContextCompaction", id="c1")]
conv, report = self._normalized(tmp_path, records=records)
markers = [
b for m in conv["messages"] for b in m["blocks"]
if b.get("kind") == COLLAPSED_KIND_HIDDEN_CONTEXT
]
assert len(markers) == 1
assert "compacted" in markers[0]["origin"]
assert "context_compaction" in report.format_summary()
def test_unknown_item_type_is_reported(self, tmp_path):
records = _records() + [_item("QuantumMessage", id="q1", mystery=True)]
conv, report = self._normalized(tmp_path, records=records)
unknowns = [
b for m in conv["messages"] for b in m["blocks"] if b["type"] == "unknown"
]
assert len(unknowns) == 1
assert unknowns[0]["raw_type"] == "codex.QuantumMessage"
assert "codex.QuantumMessage" in report.format_summary()
def test_extension_labelled_by_kind(self, tmp_path):
records = _records() + [
_item("Extension", id="e1", kind="web.search", query="hsts", results=[]),
]
conv, _ = self._normalized(tmp_path, records=records)
origins = " ".join(
b["origin"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
)
assert "web.search ×1" in origins
class TestRepoTags:
def test_title_tagged_with_repos_touched(self, tmp_path):
repo = tmp_path / "ws" / "myrepo"
(repo / ".git").mkdir(parents=True)
records = _records(cwd=str(tmp_path / "ws")) + [
_item(
"FileChange",
id="fc1",
status="completed",
changes={str(repo / "README.md"): {"type": "add", "content": "x"}},
)
]
_write_session(tmp_path, records)
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["title"].endswith("[myrepo]")
def test_no_tag_when_nothing_touched(self, tmp_path):
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["title"] == "Write the backup guide."
class TestResolveRoots:
def test_default_root(self, monkeypatch):
monkeypatch.delenv("CODEX_DIR", raising=False)
monkeypatch.delenv("CODEX_HOME", raising=False)
assert resolve_roots()[0].name == "sessions"
def test_codex_dir_splits_on_pathsep(self, monkeypatch):
monkeypatch.setenv("CODEX_DIR", "/a/sessions:/b/sessions")
monkeypatch.delenv("CODEX_HOME", raising=False)
assert [str(p) for p in resolve_roots()] == ["/a/sessions", "/b/sessions"]
def test_codex_home_appended_and_deduped(self, monkeypatch):
monkeypatch.setenv("CODEX_DIR", "/a/sessions")
monkeypatch.setenv("CODEX_HOME", "/a")
# /a/sessions is already listed — must not appear twice.
assert [str(p) for p in resolve_roots()] == ["/a/sessions"]
class TestExoticLineBreaks:
"""U+0085 (NEL) and friends are legal inside a JSON string.
``str.splitlines()`` breaks on them, shredding one record into unparseable
fragments and losing it silently. Observed 2026-08-18 in a real rollout,
where captured command output contained two NELs.
"""
# C0 controls (\x0b, \x0c) are excluded: JSON requires them escaped, so they
# never reach the splitter literally. These three do.
@pytest.mark.parametrize("sep", ["\x85", "
", "
"])
def test_record_with_exotic_break_survives(self, tmp_path, sep):
records = _records()
records.append(
_item(
"AgentMessage",
ts="2026-08-17T05:46:00.000Z",
id="a3",
phase="final_answer",
content=[{"type": "Text", "text": f"before{sep}after"}],
)
)
_write_session(tmp_path, records)
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert f"before{sep}after" in texts
class TestFullPolicyShortfall:
def test_shortfall_note_is_not_a_collapsed_block(self, tmp_path):
"""Under `full`, a collapsed block would advise setting the policy that
is already in force. The shortfall is stated plainly instead."""
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path, hidden_content="full")
prov.list_conversations()
report = LossReport()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID), report)
collapsed = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
]
assert collapsed == []
notes = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TOOL_RESULT and b.get("tool_name") == "incomplete"
]
assert len(notes) == 1
assert "1 tool call(s) produced no result" in notes[0]["output"]
assert "tool_call_incomplete" in report.format_summary()
+6
View File
@@ -60,7 +60,13 @@ class TestSessionLimiterConfig:
"""MAX_CONVERSATIONS_PER_RUN and REQUEST_DELAY parsing in load_config."""
def _load(self, monkeypatch, tmp_path, **env):
from src import config as config_module
from src.config import load_config
# load_config() calls load_dotenv(override=False), which re-populates
# any variable this test just deleted from the developer's real .env —
# so test_defaults only saw defaults on a machine without one. Stub it:
# these tests are about parsing the environment, not discovering .env.
monkeypatch.setattr(config_module, "load_dotenv", lambda *a, **k: False)
monkeypatch.setenv("EXPORT_DIR", str(tmp_path / "exports"))
monkeypatch.setenv("CACHE_DIR", str(tmp_path / "cache"))
for key in (
+16
View File
@@ -170,6 +170,22 @@ class TestResolveMedia:
# Still renders as a placeholder, not a broken image link
assert render_blocks_to_markdown([block]).startswith("> 🖼️")
def test_forbidden_counted_separately_from_generic_error(self, tmp_path):
"""403 is a distinct bucket: the asset exists, we were refused."""
ref = "sediment://file_denied"
provider = _FakeProvider(
fail_refs={ref: RuntimeError("HTTP 403 — detail: unauthorized")}
)
block = make_image_placeholder(ref=ref, source="model_generated")
report = LossReport()
resolve_media(
_conv_with([block]), provider, tmp_path, "provider/project/year",
"images", report,
)
assert report.media_failed["forbidden"] == 1
assert "download-error" not in report.media_failed
def test_provider_without_download_asset(self, tmp_path):
"""claude-code has no remote assets — resolve_media must no-op."""
class NoDownload:
+263
View File
@@ -1116,3 +1116,266 @@ class TestClaudeDriftCanary:
p = self._provider([{"uuid": "u1", "name": "N", "updated_at": "z"}],
self._detail([]))
assert DRIFT_ERROR in _sev(p.check_drift())
# ---------------------------------------------------------------------------
# 4xx diagnostics: curl_cffi renders raise_for_status() as
# "HTTP Error {code}: {reason}", and HTTP/2 has no reason phrase — so a bare
# 403 logged as "HTTP Error 403:" says nothing. The body carries the cause.
# ---------------------------------------------------------------------------
class TestErrorBodyDiagnostics:
class _Resp:
ok = False
status_code = 403
reason = ""
headers: dict = {}
def __init__(self, payload=None, text=""):
self._payload = payload
self.text = text
def json(self):
if self._payload is None:
raise ValueError("not json")
return self._payload
def _provider(self, response):
from src.providers.chatgpt import ChatGPTProvider
p = ChatGPTProvider.__new__(ChatGPTProvider)
p._request_delay = 0
p._last_request_at = None
p._session = type("S", (), {"request": lambda *a, **k: response})()
return p
def test_detail_field_surfaces_in_error(self):
from src.providers.base import ProviderError
resp = self._Resp(payload={"detail": "File not accessible to this account"})
with pytest.raises(ProviderError) as exc:
self._provider(resp)._make_request("GET", "https://x/files/f1/download")
message = str(exc.value.original)
assert "403" in message
assert "File not accessible to this account" in message
def test_non_json_body_excerpted(self):
from src.providers.base import ProviderError
resp = self._Resp(text="<html>Forbidden</html>")
with pytest.raises(ProviderError) as exc:
self._provider(resp)._make_request("GET", "https://x/files/f1/download")
assert "Forbidden" in str(exc.value.original)
def test_empty_body_says_so_rather_than_nothing(self):
from src.providers.base import ProviderError
resp = self._Resp(text="")
with pytest.raises(ProviderError) as exc:
self._provider(resp)._make_request("GET", "https://x/files/f1/download")
assert "empty response body" in str(exc.value.original)
def test_secrets_in_error_body_are_redacted(self):
from src.providers.base import _describe_error_body
resp = self._Resp(payload={"error": {"message": "no", "access_token": "sk-abc"}})
described = _describe_error_body(resp)
assert "sk-abc" not in described
assert "[REDACTED]" in described
# ---------------------------------------------------------------------------
# Deleted assets: ChatGPT's download endpoint answers a missing upload with
# 403 Forbidden, not 404. Measured 2026-08-17 over 18 such assets — every one
# returned 404 "File not found" on /files/{id}, while assets that downloaded
# fine returned 200 on both in the same session.
# ---------------------------------------------------------------------------
class TestDeletedAssetReporting:
class _Resp:
def __init__(self, status, payload=None):
self.status_code = status
self.ok = 200 <= status < 400
self.reason = ""
self.headers: dict = {}
self.text = ""
self._payload = payload or {}
def json(self):
return self._payload
def _provider(self, responses):
"""responses: dict of url-substring → _Resp, consumed by substring match."""
from src.providers.chatgpt import ChatGPTProvider
p = ChatGPTProvider.__new__(ChatGPTProvider)
p._request_delay = 0
p._last_request_at = None
calls: list[str] = []
def request(method, url, **kwargs):
calls.append(url)
for fragment, resp in responses.items():
if url.endswith(fragment):
return resp
raise AssertionError(f"unexpected URL: {url}")
p._session = type("S", (), {"request": staticmethod(request)})()
p._calls = calls
return p
def test_403_confirmed_missing_is_reported_as_gone(self):
from src.providers.base import ProviderError
p = self._provider({
"/download": self._Resp(403, {"detail": "Forbidden"}),
"/files/file_x": self._Resp(404, {"detail": "File not found"}),
})
with pytest.raises(ProviderError) as exc:
p.download_asset("sediment://file_x")
message = str(exc.value.original)
assert "no longer exists" in message
assert "not recoverable" in message
assert any(u.endswith("/files/file_x") for u in p._calls), "probe not sent"
def test_403_on_an_asset_that_still_exists_stays_a_403(self):
"""Don't call a live asset deleted — that would hide a real problem."""
from src.providers.base import ProviderError
p = self._provider({
"/download": self._Resp(403, {"detail": "Forbidden"}),
"/files/file_x": self._Resp(200, {"id": "file_x"}),
})
with pytest.raises(ProviderError) as exc:
p.download_asset("sediment://file_x")
message = str(exc.value.original)
assert "403" in message
assert "no longer exists" not in message
def test_probe_failure_does_not_mask_the_original_403(self):
"""A probe that errors must leave the 403 intact, not swallow it."""
from src.providers.base import ProviderError
p = self._provider({"/download": self._Resp(403, {"detail": "Forbidden"})})
# No entry for /files/file_x — the fake session raises, standing in
# for a network error on the probe.
with pytest.raises(ProviderError) as exc:
p.download_asset("sediment://file_x")
assert "403" in str(exc.value.original)
assert "no longer exists" not in str(exc.value.original)
def test_no_probe_on_the_happy_path(self):
"""The probe costs a request — it must not fire on a good download."""
class _Bytes:
status_code = 200
content = b"data"
headers = {"content-type": "image/png"}
p = self._provider({
"/download": self._Resp(200, {"download_url": "https://cdn/x", "file_name": "a.png"}),
"https://cdn/x": _Bytes(),
})
content, mime, name = p.download_asset("sediment://file_x")
assert (content, mime, name) == (b"data", "image/png", "a.png")
assert not any(u.endswith("/files/file_x") for u in p._calls), "probed needlessly"
def test_classified_as_expired_not_forbidden(self):
"""The run summary must not call a deleted upload a permissions error."""
from src.media import _classify_failure
from src.providers.base import ProviderError
err = ProviderError(
"chatgpt",
"download_asset(file_x)",
RuntimeError(
"Asset no longer exists — HTTP 404 'File not found' on /files/{id}. "
"The upload was deleted or expired server-side; it is not "
"recoverable from ChatGPT."
),
)
assert _classify_failure(err) == "expired-or-missing"
# ---------------------------------------------------------------------------
# Project attribution: a conversation names its own project via gizmo_id, so
# it should not depend on the user having listed that project in config.
# ---------------------------------------------------------------------------
class TestProjectAttribution:
def _provider(self, *, project_ids=None, gizmo_name="My Project"):
from src.providers.chatgpt import ChatGPTProvider
p = ChatGPTProvider.__new__(ChatGPTProvider)
p._project_map = {}
p._project_name_cache = {}
p._project_ids = project_ids or []
p._hidden_content = "placeholder"
p._fetch_project_name = lambda gid: gizmo_name
return p
def _raw(self, **extra):
raw = {
"conversation_id": "conv-1",
"title": "T",
"create_time": 1_700_000_000,
"update_time": 1_700_000_000,
"mapping": {},
}
raw.update(extra)
return raw
def test_gizmo_id_supplies_project_when_unconfigured(self):
p = self._provider(gizmo_name="Tech Questions")
out = p.normalize_conversation(self._raw(gizmo_id="g-p-abc123"))
assert out["project"] == "Tech Questions"
def test_explicit_annotation_still_wins(self):
p = self._provider(gizmo_name="From Gizmo")
out = p.normalize_conversation(
self._raw(gizmo_id="g-p-abc123", _project_name="From Listing")
)
assert out["project"] == "From Listing"
def test_project_map_beats_gizmo_lookup(self):
p = self._provider(gizmo_name="From Gizmo")
p._project_map["conv-1"] = "From Map"
out = p.normalize_conversation(self._raw(gizmo_id="g-p-abc123"))
assert out["project"] == "From Map"
def test_custom_gpt_is_not_treated_as_a_project(self):
"""Only g-p- ids are projects; a custom GPT must not become a folder."""
p = self._provider()
out = p.normalize_conversation(self._raw(gizmo_id="g-xyz789"))
assert out["project"] is None
def test_no_gizmo_id_stays_unprojected(self):
p = self._provider()
assert p.normalize_conversation(self._raw())["project"] is None
def test_resolved_project_is_cached_into_the_map(self):
p = self._provider(gizmo_name="Cached")
p.normalize_conversation(self._raw(gizmo_id="g-p-abc123"))
assert p._project_map["conv-1"] == "Cached"
def test_unconfigured_project_is_reported_once(self, caplog):
p = self._provider(project_ids=["g-p-known"], gizmo_name="Surprise")
with caplog.at_level(logging.INFO):
p.normalize_conversation(self._raw(gizmo_id="g-p-surprise"))
p.normalize_conversation(
self._raw(conversation_id="conv-2", gizmo_id="g-p-surprise")
)
hits = [r for r in caplog.records if "not in CHATGPT_PROJECT_IDS" in r.message]
assert len(hits) == 1
assert "Surprise" in hits[0].message
def test_configured_project_is_not_reported(self, caplog):
p = self._provider(project_ids=["g-p-known"], gizmo_name="Known")
with caplog.at_level(logging.INFO):
p.normalize_conversation(self._raw(gizmo_id="g-p-known"))
assert not [r for r in caplog.records if "not in CHATGPT_PROJECT_IDS" in r.message]
+18
View File
@@ -145,3 +145,21 @@ class TestFormatTokenStatus:
expiry = datetime.now(tz=timezone.utc) + timedelta(days=10, hours=12)
result = format_token_status("tok", expiry)
assert "10 days" in result
class TestRedactCompoundKeys:
"""Exact-match redaction let compound secret names through into logs."""
def test_compound_secret_keys_redacted(self):
result = redact_secrets(
{"access_token": "sk-abc", "api_key": "k1", "session-token": "s1"}
)
assert result == {
"access_token": "[REDACTED]",
"api_key": "[REDACTED]",
"session-token": "[REDACTED]",
}
def test_innocent_keys_containing_a_secret_word_kept(self):
result = redact_secrets({"keywords": ["a"], "monkey": "b", "tokenizer": "c"})
assert result == {"keywords": ["a"], "monkey": "b", "tokenizer": "c"}