FUTURE.md had become a 648-line archaeological record: six shipped roadmap items, two dropped, two implemented, an investigation trail for each, and a backlog closed as not needed — with the two genuinely planned items buried at the bottom. It is now 139 lines, planned work first. - Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon behind decisions now recorded in one line each — why Brave cookie extraction is not viable, why the ChatGPT token's expiry cannot be read client-side, how the drift canary was designed — which would be expensive to rediscover. - Roadmap is now two items: the StartOS service (including the 2026-08-18 decision that clients upload to StartOS storage and the service owns the Joplin connection), and the README split. - Everything shipped or dropped is off the roadmap. Decisions not to build survive as a one-line table so they are not re-proposed. - Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0. Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased] section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy notifications, gizmo_id project attribution and the `projects` command, and the splitlines data-loss fix in both local providers. Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose — exported conversations are Markdown and may contain private content — with each doc re-included by name, so a new doc is silently untracked rather than rejected. The archive would not have been committed at all. Recorded in the README-split roadmap item, since `docs/*.md` will hit exactly this. Also refreshes the package description, which still named only ChatGPT and Claude after two local providers were added. Carries the documentation and table-of-contents work from earlier today.
9.2 KiB
Planned Future Work
Status 2026-08-18 (v0.9.0). The tool archives four providers — two web (
chatgpt,claude) and two local agent-transcript (claude-code,codex) — on a schedule, on Linux and Windows, reporting results by push notification. Two items below are genuinely planned. Everything else has shipped or been decided against.
Completed work moves to the changelog; this file holds only what is not built
yet. Decisions not to build something are recorded at the bottom in one line
each, so they are not re-proposed — the full investigation trails behind them
are in FUTURE-ARCHIVE.md.
Roadmap
1. StartOS Service — one corpus, not per-machine islands
Each machine currently archives to its own exports/ and its own Joplin. Work
is split across boxes — coding sessions (claude-code, codex) on the Linux
machine, web chats (chatgpt, claude) on the Windows one — so there is no
single place where all conversations exist together.
This was dropped on 2026-06-28 and reopened 2026-08-18, because the
reasoning behind the drop has gone stale. It rested on "the source conversations
live in the providers' clouds and can be re-downloaded". That is no longer true
of half the providers: claude-code and codex transcripts exist only on the
machine that produced them, Codex prunes its rollout files, and neither is
recoverable from any cloud. The motivation is also different from the one
weighed then — consolidation, not durability.
Intended split of responsibility (decided 2026-08-18). Each machine keeps running the exporter locally and keeps doing what only it can do: read that machine's local transcripts, and hold the browser session for the web providers. What changes is where the output goes. Instead of syncing to Joplin itself, a local run uploads its conversations to the StartOS storage area, and the StartOS service owns the Joplin connection for the whole corpus.
That inverts today's arrangement and removes two problems we already have:
- The Joplin-availability race leaves the clients. A scheduled run currently
has to find Joplin desktop open on that same machine — the 09:02 timer run on
2026-08-18 exported fine and then skipped the sync because Joplin did not
start until 09:07. A server that is always up has no such window, and
--joplin-optionalstops being load-bearing. - One Joplin integration instead of N. Notebook naming, resource upload and note updates run once, server-side, against one manifest — rather than each machine independently deciding what a notebook is called and racing to update the same note.
What this needs, and what it does not:
- Not a headless web-provider login. The hard sub-problem the original drop retired stays retired: the web providers keep running interactively on the machine that has the browser, and push their output to the server. Only the local providers would run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's
joplincommand, pointed at the StartOS service instead of a local Joplin API. Probably apushalongsidesync, so a scheduled client run stays one line. - Per-machine identity in the corpus, which the exporter does not track today:
claude_code.resolve_rootsdeliberately merges multiple roots with "no per-machine label". Centralizing makes that label load-bearing. - Conflict handling for one conversation seen by two machines, and a decision about whether the server or the client owns the cache manifest. It is per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of
claude-code/codextranscripts once Codex prunes its rollouts, so the client should keep them rather than hand them off.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is where the archive is actually read. This is a "nice eventually", not a gap.
2. Split the README into separate documents
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61 headings. An H2-only table of contents was added 2026-08-18 and helps, but it treats the symptom: the file is doing at least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|---|---|
README.md |
What it is, install, first run, a pointer to the rest |
docs/providers.md |
ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
docs/cli.md |
CLI Reference — ~250 lines on its own, the single biggest section |
docs/scheduling.md |
Scheduling a Daily Run, notifications |
docs/troubleshooting.md |
Troubleshooting, How the Cache Works |
Not urgent, because it has a real cost the TOC did not: any existing link into a README section — a bookmark, a note, another repo, a commit message — breaks when that section moves to another file. Worth folding into the next substantial README edit rather than doing as a change of its own.
Two things to settle when it happens:
.gitignoreignores*.mdon purpose — exported conversations are Markdown and may contain private content — and re-includes each doc by name (!README.md,!FUTURE.md, …). New files underdocs/will be silently ignored, with no error, until!docs/*.mdis added. This already bit the creation ofFUTURE-ARCHIVE.mdon 2026-08-18.- Whether
docs/renders acceptably on the Gitea instance hosting this repo. Relative links between Markdown files do work there, but confirm before splitting rather than after. - Whether anchors in the split files are stable enough to link between documents, or whether cross-references should point at file tops only. GFM anchors derive from heading text and break silently on a reword — the same fragility the TOC already carries.
Decided against
Recorded so they are not re-proposed. Full reasoning and the recon behind each
is in FUTURE-ARCHIVE.md.
| Item | Verdict |
|---|---|
| Brave/Chromium cookie auto-extraction | Not viable (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
| In-app watch/polling loop | Dropped (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as sync plus a systemd timer and a Windows scheduled task. |
| Proactive token-expiry notification | Blocked, and moot (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and /api/auth/session's expires is a rolling window that advances even for a dead token. doctor's error-based health check covers the real need. |
| Official export-ZIP fallback | Closed (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. BaseProvider still admits a FileProvider if that changes. |
Joplin --force flag |
Closed as not needed (2026-06-28). |
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
| Search command | Closed as not needed (2026-06-28). grep/ripgrep over EXPORT_DIR covers it. |
Claude attachments / files handling |
Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
Completed
Detail for each release is in CHANGELOG.md.
- v0.1.0 — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- v0.2.0 — Joplin import automation (
joplincommand, create/update notes, notebook auto-creation) - v0.4.0 — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via
LossReportand visibleunknownblocks - v0.5.0 — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- v0.6.0 — Collapse tool retrieval dumps & hidden context (
EXPORTER_HIDDEN_CONTENT); Claude Code session provider;prune+ manifest integrity; binary content downloads; session limiter and request pacing;export --force - v0.7.0 —
canarydrift detection; real ChatGPT token-health check ondoctor; removal of the non-viable browser-cookie path - v0.8.0 — Claude Code coverage reopened: subagent capture as folded
<details>, repo[tags]in titles, its ownAI-ClaudeCodenotebook with self-healing note moves, multi-root scanning (CLAUDE_CODE_DIR,CLAUDE_CONFIG_DIR) - v0.9.0 — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms;
syncwith a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution viagizmo_idand theprojectscommand; a silent data-loss fix in both local providers (splitlinesbreaking on U+0085/U+2028/U+2029)