Commit Graph
20 Commits
Author SHA1 Message Date
JesseMarkowitz 3f82b35a55 tools: try fetching refused images by their Library ID
Dumping the raw message metadata found the identity the asset pointer never
carried:

    "id":              "file_000000003454722f9481506b96aed510"   ← refused
    "library_file_id": "libfile_4eb82f478fe081919127e2eba9886e86"
    "source":          "local"

The exporter only knows the sediment id from the asset_pointer and asks
/files/{sediment}/download, which 403s for these. The Library is a separate
store with its own ids, so we have been asking for the conversation-scoped
copy of a file that now lives in the Library. That also fits the 2026-07-14
creation-time cluster: a Library migration would mint exactly these new
records.

Neither gizmo_id, conversation_id, a /gizmos path nor a project Referer
helped (403/404), so scope was never the missing piece — identity was.

library_download_probe.py pairs each refused sediment id with its
library_file_id from message.metadata.attachments and tries the endpoints
that could serve it.
2026-08-17 09:33:05 -04:00
JesseMarkowitz 03646009b9 tools: probe how to download a gizmo-scoped file
metadata_diff found the discriminator. Every refused image carries
use_case="gizmo" — ChatGPT's name for Projects and custom GPTs — while
everything that downloads is image_gen or multimodal. All are state=ready,
so nothing is damaged: the plain /files/{id}/download endpoint just will
not serve a project-scoped file.

Second clue: all 11 were created 2026-07-14 within minutes of each other,
yet appear in conversations dated 2025-11 through 2026-08, several
predating their own creation_time. Something on July 14 re-created them as
gizmo-scoped copies and repointed the conversations — which is also why
the losses start in July. Not a policy change, an event.

gizmo_download_probe.py dumps the raw conversation part carrying one of
these assets (it may simply name the scope the download wants) and then
tries the plausible calls — gizmo_id/conversation_id query params, a
/gizmos/{id}/files path, a project Referer. A 200 is the fix.
2026-08-17 09:22:48 -04:00
JesseMarkowitz a3ac279e39 tools: find what separates a refused image from a served one
branch_check killed the abandoned-branch hypothesis — every lost image is
on the live branch, and the 3 images that do sit on abandoned branches are
alive. But it turned up something the earlier probing missed by sampling
four IDs from one family and generalising: of the 19 failures, only 7 are
actually gone. The other 12 answer /files/{id} with 200 and full metadata
and refuse only /download. They exist, and may be recoverable.

Three states, then: gone (404), refused (200 meta + 403 download), working
(200 both). Since metadata comes back for the refused ones, the
discriminator can be read straight off — fetch it for every file in each
state and compare fields, flagging any field whose values never overlap
between states.

Also re-checks /download now: if a file refused during the export serves
today, those 403s were transient and a retry pass recovers them, which is
a completely different fix from anything permanent.
2026-08-17 09:18:32 -04:00
JesseMarkowitz 021a76c628 tools: test whether lost images sit on abandoned branches
Expiry is now ruled out: 36/36 sampled images from 2025-09 through
2026-08 are still live server-side, so nothing dies of age and export
cadence is not the variable. The loss is per-asset — 2026-07-09 kept 38
images and lost 12 in one conversation.

Next hypothesis: those images belong to messages that were edited or
regenerated. ChatGPT stores a conversation as a tree and editing forks it,
stranding the superseded messages on an abandoned branch. The exporter
walks every node (chatgpt.py:963), so it exports those branches too — and
an attachment unreachable from live history is a natural GC target.

branch_check.py fetches the raw conversation, derives the live branch by
walking parent links up from current_node, and cross-tabulates every image
against (on the live branch?) x (still downloadable?). If the lost images
are all off-branch and nothing on the live branch is missing, this is not
data loss at all — it is attachments to messages that were replaced.
2026-08-17 08:58:20 -04:00
JesseMarkowitz b9e8896b33 tools: distinguish "captured in time" from "still alive"
analyze_media_age.py claimed age was ruled out because an image from
2025-11 was saved while one from 2026-08 was not. That conclusion does not
follow. "Saved" means some earlier run downloaded it, not that ChatGPT
still holds it — an image captured in June looks saved forever after, even
if it died in July. A month with no losses shows the exports were timely,
not that the assets survived.

- probe_survival_by_month.py: sample images already on disk, grouped by
  their conversation's month, and ask /files/{id} whether each still
  exists today. Old months still 200 → no expiry, and cadence did not save
  them. Old months now 404 → uploads do expire and cadence is the whole
  ballgame.
- analyze_media_age.py: stop asserting the unsupported verdict; say what
  the number does and does not show, and point at the probe.

One signal there is immune to the confound and survives: 2026-07-09 kept
38 images and lost 12. Same conversation, same day, opposite outcomes —
no retention policy does that, so at least part of this is per-asset.
2026-08-17 08:53:54 -04:00
JesseMarkowitz 04191eed8c tools: analyze whether media loss is age-based
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
(![source](media/…) when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.

Offline, no token, no API calls.
2026-08-17 08:45:21 -04:00
JesseMarkowitz f40b25001a tools: add one-off probe for the media 403s
The improved error reporting landed, and the answer it produced is
{"detail":"Forbidden"} — generic, no reason. The cause has to be narrowed
by experiment instead, so collect the experiments in one script:

- classify the failed assets offline from the exported placeholders
  (user_upload vs model_generated) — no API call needed
- probe a known-good asset in the same session as a control, to rule the
  session in or out
- retry the 403 with ChatGPT-Account-Id, which the exporter never sends
  and which workspace-scoped resources can require
- compare /files/{id} against /files/{id}/download

Temporary: delete once the cause is known, or fold into `doctor` if the
check earns a permanent home.
2026-08-17 08:08:01 -04:00
JesseMarkowitz 395ea19ca8 fix: surface the response body on 4xx so 403s are diagnosable
Media downloads logged "HTTP Error 403:" with no reason. That string is
curl_cffi's raise_for_status() format, "HTTP Error {code}: {reason}", and
HTTP/2 carries no reason phrase — so the message said nothing, and
_make_request threw the response body away. The provider's JSON `detail`
is the only explanation available for a refused asset.

- base._make_request: end non-retryable statuses with a ProviderError
  carrying the body's detail/error/message (redacted, truncated to 300
  chars) instead of a bare raise_for_status().
- media: bucket 403 as `forbidden` in the run summary, separately from
  `download-error` — "the asset is gone" and "we were refused" are
  different problems.
- utils.redact_secrets: match secret key names per word. Exact matching
  let access_token, api_key, and session-token through into logged
  bodies; "keywords"/"monkey"/"tokenizer" stay intact.
- tests/test_config.py: test_defaults depended on the absence of a local
  .env — load_config() calls load_dotenv(override=False), which restored
  the variable the test had just deleted. Stub dotenv discovery.

305 tests pass.
2026-08-17 07:58:02 -04:00
JesseMarkowitzandClaude Opus 4.8 1f5a445ada feat: v0.8.0 — Claude Code subagent capture, own Joplin notebook, git-root repo tags, multi-root scanning
- Subagents: fold Task-tool transcripts (subagents/*.jsonl) inline as
  collapsible <details> blocks at the spawn point; their own tool traffic
  collapses under the same policy; nesting handled via toolUseId matching
- Joplin: Claude Code gets its own top-level AI-ClaudeCode notebook;
  update_note sets parent_id so notes self-heal/relocate on re-sync
- Repo [tags] in titles via git-root detection (nearest .git ancestor),
  home-wide and cross-workspace; CLAUDE_CODE_REPO_TAG_IGNORE escape hatch
- Multi-root scanning: CLAUDE_CODE_DIR as os.pathsep list + CLAUDE_CONFIG_DIR;
  merged by launch-folder, newer-mtime wins on UUID collision
- 298 tests passing

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 05:04:41 -04:00
JesseMarkowitzandClaude Opus 4.8 bbcb29c856 docs: close backlog — project feature-complete for now
Mark FUTURE.md as feature-complete as of v0.7.0: active roadmap is empty and
the deprioritized backlog (Joplin --force, per-conversation cache reset,
official export-ZIP fallback, o1/o3 reclassification, Obsidian output,
token-expiry notifications, search) is closed as not needed, kept for
reference only. No further work planned.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:51:05 -04:00
JesseMarkowitzandClaude Opus 4.8 1e016ea652 canary: detect provider API schema drift; release v0.7.0
The export reads ChatGPT's and Claude's undocumented internal web APIs,
which can change shape without notice; the worst failure for a backup tool
is a silent one (skipped/mis-parsed content with no error). Add a `canary`
command + BaseProvider.check_drift() (overridden by ChatGPT/Claude) that
fetches one listing page + one conversation per provider and asserts only
the normalizer's load-bearing fields — not the full response shape, which
churns harmlessly. The top silent risk it guards is a renamed retrieval-tool
author bypassing the hidden-content collapse. ERROR findings exit non-zero;
WARN findings are surfaced but non-fatal so a backup run is never blocked.

Also retires the in-app watch mode and headless StartOS direction from the
roadmap (tool stays a local, manually-run CLI) and updates docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:37:24 -04:00
JesseMarkowitzandClaude Opus 4.8 ef603cf659 doctor: report ChatGPT token health via /api/auth/session error field
ChatGPT session tokens are JWEs whose `exp` is encrypted and unreadable
client-side, so the old JWT-decode path could never yield an expiry, and the
`/api/auth/session` `expires` is a misleading rolling window (it advances on
every call even for a dead token). The honest signal is the `error` field:
`RefreshAccessTokenError` means the session token is dead while `expires` and
a stale `accessToken` are still echoed (verified live).

- Add ChatGPTProvider.session_health() and a "ChatGPT token active" doctor
  check based on it; drop the dead JWE/exp decode path in doctor and auth.
- Fix _fetch_access_token to fail fast on a set `error` instead of returning
  the stale accessToken (which produced a confusing downstream 401).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:37:01 -04:00
JesseMarkowitzandClaude Opus 4.8 4a7ee5773f Remove auth --from-browser browser cookie extraction
Browser cookie auto-extraction is not viable: modern Chromium App-Bound
Encryption (Chrome 127+/current Brave) keys cookies off a SYSTEM-level
layer that cannot be decrypted off disk without admin rights and AV-flagged
SYSTEM impersonation, and fails on Brave specifically. Drop the
`--from-browser` flag, `_auth_from_browser`, `src/browser_tokens.py`, its
test, and the browser-cookie3 dependency. Auth is manual (DevTools wizard)
only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:35:42 -04:00
JesseMarkowitz 3cf0b1eaa8 fix: force re-render campaign tracking, cache dir loading, Joplin link preservation. Co-Authored-By: Fable 5 2026-06-12 22:39:12 -04:00
JesseMarkowitz 456975ad50 fix: --force re-render progression, cache dir loading, and Joplin link preservation. Co-Authored-By: Fable 5 <noreply@anthropic.com> 2026-06-12 21:25:47 -04:00
JesseMarkowitz 9e1a8ab7cb feat: v0.6.0 — collapse policy, session limiter, Claude Code provider, prune, browser auth, media downloads 2026-06-12 18:26:14 -04:00
JesseMarkowitzandClaude Sonnet 4.6 557994f7d9 fix: persist created_at in cache so Joplin note titles get date prefix
mark_exported() was discarding created_at from the metadata dict because
it wasn't in the hardcoded stored-key list, so the joplin sync always
saw an empty date and omitted the prefix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 11:36:21 -04:00
JesseMarkowitzandClaude Sonnet 4.6 e9b2e42893 feat: v0.5.0 — nested Joplin notebooks, date-prefixed note titles, flat year folders
Joplin notebooks now use a two-level hierarchy: AI-ChatGPT / <project> and
AI-Claude / <project> instead of a single flat title. Note titles are prefixed
with the conversation created_at date (YYYY-MM-DD). Export folders collapse
provider/project/year into a single provider/project.year directory.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-05 11:05:39 -04:00
JesseMarkowitzandClaude Opus 4.7 68e8d532be feat: v0.4.1 — ChatGPT tool-output content types and conv_id fix
First real-data export against v0.4.0 surfaced 66 unknown blocks across
three content types — captured live and added.

Added:
- execution_output (Code Interpreter / container.exec / python tool
  output) → tool_result block. output=content.text,
  tool_name=author.name, is_error=metadata.aggregate_result.status,
  summary=metadata.reasoning_title
- system_error → error tool_result with tool_name=author.name
- tether_browsing_display: spinner placeholders (empty result+summary)
  skip silently with DEBUG log; defensive populated-case branch maps
  to tool_result (untested in real data)
- tool_result block schema: optional `summary` field rendered as
  italic line between header and fence
- tool_result rendering: tool_name appears in header when present
  (e.g. `📤 Result: container.exec`); existing tool_name=None calls
  unchanged
- _ROLE_LABELS["tool"] = ("🔧 Tool", "tool")

Fixed:
- chatgpt.normalize_conversation reads `conversation_id` as fallback
  for `id`. Live API uses conversation_id; fixtures use id.
  Pre-fix: empty id in YAML frontmatter and missing context in
  WARNING logs.

Tests: 11 new (192 total, 0 failures). Fixture extended with 4
tool-output cases (execution_output success, empty execution_output
that should skip, system_error, tether_browsing_display spinner).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-05 09:25:55 -04:00
JesseMarkowitzandClaude Opus 4.7 473d02f71a feat: v0.4.0 — rich content support with typed blocks and loss visibility
Extracts per-message content into a typed `blocks` list (text, code,
thinking, tool_use, tool_result, image_placeholder, file_placeholder,
unknown) and renders them at exporter write time. Voice transcripts,
Custom Instructions, and image references now appear in exports
instead of being silently dropped.

Foundation:
- src/blocks.py: pure block constructors, _safe_fence (fence-corruption
  defense, verified live in Joplin), _blockquote_prefix, render
- src/loss_report.py: per-run tally surfaced as INFO summary at end of
  export so silently-dropped data becomes visible

Providers:
- ChatGPT: dispatch on content_type produces typed blocks; voice shapes
  (audio_transcription, audio_asset_pointer, real_time_user_audio_video_
  asset_pointer) locked from live DevTools capture; Custom Instructions
  bug fix (parts-vs-direct-fields); role filter lifted; hidden-context
  marker driven by is_visually_hidden_from_conversation flag
- Claude: defensive dispatch for text/thinking/tool_use/tool_result/image
  with recursive nested-block flattening; untested against real rich-
  content data — fix-forward in v0.4.1

Exporter:
- Markdown renders from blocks at write time via render_blocks_to_markdown;
  backward-compat fallback to content for any pre-v0.4.0 cached data

Tests:
- 27 new tests across providers, exporters, CLI; fixtures rebuilt with
  real-shape ChatGPT voice + Custom Instructions cases
- 181/181 pass

Behavior changes (intentional):
- JSON output omits content; consumers should read blocks
- Per-conversation message counts increase (Custom Instructions, image-
  only, tool-only messages now appear)
- Existing exports not auto-re-rendered; users wanting fresh output run
  cache --clear then export

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-04 23:17:18 -04:00