Files
AIChatExporter/FUTURE.md
T
JesseMarkowitz 2d5fcb26f5 release: v0.9.0, and rewrite FUTURE.md around what is actually planned
FUTURE.md had become a 648-line archaeological record: six shipped roadmap
items, two dropped, two implemented, an investigation trail for each, and a
backlog closed as not needed — with the two genuinely planned items buried
at the bottom. It is now 139 lines, planned work first.

- Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon
  behind decisions now recorded in one line each — why Brave cookie
  extraction is not viable, why the ChatGPT token's expiry cannot be read
  client-side, how the drift canary was designed — which would be expensive
  to rediscover.

- Roadmap is now two items: the StartOS service (including the 2026-08-18
  decision that clients upload to StartOS storage and the service owns the
  Joplin connection), and the README split.

- Everything shipped or dropped is off the roadmap. Decisions not to build
  survive as a one-line table so they are not re-proposed.

- Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0.

Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased]
section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the
launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy
notifications, gizmo_id project attribution and the `projects` command, and
the splitlines data-loss fix in both local providers.

Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose —
exported conversations are Markdown and may contain private content — with
each doc re-included by name, so a new doc is silently untracked rather than
rejected. The archive would not have been committed at all. Recorded in the
README-split roadmap item, since `docs/*.md` will hit exactly this.

Also refreshes the package description, which still named only ChatGPT and
Claude after two local providers were added.

Carries the documentation and table-of-contents work from earlier today.
2026-08-18 13:51:34 -04:00

145 lines
9.2 KiB
Markdown

# Planned Future Work
> **Status 2026-08-18 (v0.9.0).** The tool archives four providers — two web
> (`chatgpt`, `claude`) and two local agent-transcript (`claude-code`, `codex`)
> — on a schedule, on Linux and Windows, reporting results by push notification.
> Two items below are genuinely planned. Everything else has shipped or been
> decided against.
Completed work moves to the changelog; this file holds only what is *not* built
yet. Decisions not to build something are recorded at the bottom in one line
each, so they are not re-proposed — the full investigation trails behind them
are in `FUTURE-ARCHIVE.md`.
---
# Roadmap
## 1. StartOS Service — one corpus, not per-machine islands
Each machine currently archives to its own `exports/` and its own Joplin. Work
is split across boxes — coding sessions (`claude-code`, `codex`) on the Linux
machine, web chats (`chatgpt`, `claude`) on the Windows one — so there is no
single place where all conversations exist together.
This was dropped on 2026-06-28 and **reopened 2026-08-18**, because the
reasoning behind the drop has gone stale. It rested on "the source conversations
live in the providers' clouds and can be re-downloaded". That is no longer true
of half the providers: `claude-code` and `codex` transcripts exist *only* on the
machine that produced them, Codex prunes its rollout files, and neither is
recoverable from any cloud. The motivation is also different from the one
weighed then — consolidation, not durability.
**Intended split of responsibility (decided 2026-08-18).** Each machine keeps
running the exporter locally and keeps doing what only it can do: read that
machine's local transcripts, and hold the browser session for the web providers.
What changes is where the output goes. Instead of syncing to Joplin itself, a
local run **uploads its conversations to the StartOS storage area**, and the
StartOS service owns the Joplin connection for the whole corpus.
That inverts today's arrangement and removes two problems we already have:
- **The Joplin-availability race leaves the clients.** A scheduled run currently
has to find Joplin desktop open on that same machine — the 09:02 timer run on
2026-08-18 exported fine and then skipped the sync because Joplin did not
start until 09:07. A server that is always up has no such window, and
`--joplin-optional` stops being load-bearing.
- **One Joplin integration instead of N.** Notebook naming, resource upload and
note updates run once, server-side, against one manifest — rather than each
machine independently deciding what a notebook is called and racing to update
the same note.
What this needs, and what it does *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers keep running interactively on the
machine that has the browser, and push their output to the server. Only the
local providers would run server-side, and they need no tokens at all.
- An upload step in the client — the counterpart of today's `joplin` command,
pointed at the StartOS service instead of a local Joplin API. Probably a
`push` alongside `sync`, so a scheduled client run stays one line.
- Per-machine identity in the corpus, which the exporter does not track today:
`claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing makes that label load-bearing.
- Conflict handling for one conversation seen by two machines, and a decision
about whether the server or the client owns the cache manifest. It is
per-machine today, and that is what makes "already up to date" mean anything.
- A story for what the client keeps locally after a successful upload. Exports
are the only copy of `claude-code` / `codex` transcripts once Codex prunes its
rollouts, so the client should keep them rather than hand them off.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 2. Split the README into separate documents
The README is 830 lines / 5,562 words / 37 KB — about a 25-minute read, with 61
headings. An H2-only table of contents was added 2026-08-18 and helps, but it
treats the symptom: the file is doing at least four unrelated jobs at once.
Rough shape of a split:
| Document | Content today |
|----------|---------------|
| `README.md` | What it is, install, first run, a pointer to the rest |
| `docs/providers.md` | ChatGPT Projects, Claude Code Sessions, Codex Sessions, session tokens |
| `docs/cli.md` | CLI Reference — ~250 lines on its own, the single biggest section |
| `docs/scheduling.md` | Scheduling a Daily Run, notifications |
| `docs/troubleshooting.md` | Troubleshooting, How the Cache Works |
Not urgent, because it has a real cost the TOC did not: any existing link into a
README section — a bookmark, a note, another repo, a commit message — breaks
when that section moves to another file. Worth folding into the next substantial
README edit rather than doing as a change of its own.
Two things to settle when it happens:
- **`.gitignore` ignores `*.md` on purpose** — exported conversations are
Markdown and may contain private content — and re-includes each doc by name
(`!README.md`, `!FUTURE.md`, …). New files under `docs/` will be silently
ignored, with no error, until `!docs/*.md` is added. This already bit the
creation of `FUTURE-ARCHIVE.md` on 2026-08-18.
- Whether `docs/` renders acceptably on the Gitea instance hosting this repo.
Relative links between Markdown files do work there, but confirm before
splitting rather than after.
- Whether anchors in the split files are stable enough to link *between*
documents, or whether cross-references should point at file tops only. GFM
anchors derive from heading text and break silently on a reword — the same
fragility the TOC already carries.
---
# Decided against
Recorded so they are not re-proposed. Full reasoning and the recon behind each
is in `FUTURE-ARCHIVE.md`.
| Item | Verdict |
|------|---------|
| Brave/Chromium cookie auto-extraction | **Not viable** (2026-06-12). App-Bound Encryption in Chrome 127+/current Brave needs admin + SYSTEM impersonation that AV flags as credential theft, and fails on Brave specifically. Auth stays manual via DevTools. |
| In-app watch/polling loop | **Dropped** (2026-06-28). Scheduling belongs to the host. Delivered instead in v0.9.0 as `sync` plus a systemd timer and a Windows scheduled task. |
| Proactive token-expiry notification | **Blocked, and moot** (2026-06-28). No trustworthy client-side expiry exists: the ChatGPT token is a JWE, Claude's is opaque, and `/api/auth/session`'s `expires` is a rolling window that advances even for a dead token. `doctor`'s `error`-based health check covers the real need. |
| Official export-ZIP fallback | **Closed** (2026-06-12). ChatGPT's official export omits project data, so it is not a real fallback here. `BaseProvider` still admits a `FileProvider` if that changes. |
| Joplin `--force` flag | Closed as not needed (2026-06-28). |
| Per-conversation cache reset | Closed as not needed (2026-06-28). Workaround: edit the manifest. |
| o1/o3 reasoning subpart reclassification | Closed (2026-06-28) — never seen in real captured data. |
| Obsidian vault output | Closed as not needed (2026-06-28). The Markdown is already Obsidian-valid; only file copying would be needed. |
| Search command | Closed as not needed (2026-06-28). `grep`/`ripgrep` over `EXPORT_DIR` covers it. |
| Claude `attachments` / `files` handling | Closed (2026-06-28) — never seen in real data; the drift canary will flag it if it appears. |
| Additional web providers (Gemini, Grok, Perplexity) | Out of scope — no significant usage to archive. |
---
# Completed
Detail for each release is in `CHANGELOG.md`.
- **v0.1.0** — Core export: ChatGPT + Claude, incremental sync, Markdown + JSON output
- **v0.2.0** — Joplin import automation (`joplin` command, create/update notes, notebook auto-creation)
- **v0.4.0** — Rich content: typed message blocks, ChatGPT voice transcripts, Custom Instructions extraction, data-loss visibility via `LossReport` and visible `unknown` blocks
- **v0.5.0** — Nested Joplin notebooks, date-prefixed note titles, flat year folders
- **v0.6.0** — Collapse tool retrieval dumps & hidden context (`EXPORTER_HIDDEN_CONTENT`); Claude Code session provider; `prune` + manifest integrity; binary content downloads; session limiter and request pacing; `export --force`
- **v0.7.0** — `canary` drift detection; real ChatGPT token-health check on `doctor`; removal of the non-viable browser-cookie path
- **v0.8.0** — Claude Code coverage reopened: subagent capture as folded `<details>`, repo `[tags]` in titles, its own `AI-ClaudeCode` notebook with self-healing note moves, multi-root scanning (`CLAUDE_CODE_DIR`, `CLAUDE_CONFIG_DIR`)
- **v0.9.0** — Codex CLI provider; launcher scripts removing the virtualenv ceremony on both platforms; `sync` with a meaningful exit code; daily scheduling for Linux and Windows; ntfy push notifications; ChatGPT project attribution via `gizmo_id` and the `projects` command; a silent data-loss fix in both local providers (`splitlines` breaking on U+0085/U+2028/U+2029)