release: v0.9.0, and rewrite FUTURE.md around what is actually planned

FUTURE.md had become a 648-line archaeological record: six shipped roadmap
items, two dropped, two implemented, an investigation trail for each, and a
backlog closed as not needed — with the two genuinely planned items buried
at the bottom. It is now 139 lines, planned work first.

- Archives the old file verbatim as FUTURE-ARCHIVE.md. It carries the recon
  behind decisions now recorded in one line each — why Brave cookie
  extraction is not viable, why the ChatGPT token's expiry cannot be read
  client-side, how the drift canary was designed — which would be expensive
  to rediscover.

- Roadmap is now two items: the StartOS service (including the 2026-08-18
  decision that clients upload to StartOS storage and the service owns the
  Joplin connection), and the README split.

- Everything shipped or dropped is off the roadmap. Decisions not to build
  survive as a one-line table so they are not re-proposed.

- Status notes for v0.7.0 and v0.8.0 move into Completed, joined by v0.9.0.

Releases v0.9.0: pyproject 0.8.0 -> 0.9.0, and the changelog's [Unreleased]
section becomes [0.9.0] - 2026-08-18. Covers the Codex provider, the
launcher scripts, `sync`, daily scheduling on Linux and Windows, ntfy
notifications, gizmo_id project attribution and the `projects` command, and
the splitlines data-loss fix in both local providers.

Whitelists FUTURE-ARCHIVE.md in .gitignore. `*.md` is ignored on purpose —
exported conversations are Markdown and may contain private content — with
each doc re-included by name, so a new doc is silently untracked rather than
rejected. The archive would not have been committed at all. Recorded in the
README-split roadmap item, since `docs/*.md` will hit exactly this.

Also refreshes the package description, which still named only ChatGPT and
Claude after two local providers were added.

Carries the documentation and table-of-contents work from earlier today.
This commit is contained in:
JesseMarkowitz
2026-08-18 13:51:34 -04:00
parent 55b9ce12f6
commit 2d5fcb26f5
7 changed files with 909 additions and 579 deletions
+137 -3
View File
@@ -6,7 +6,28 @@ Supports incremental sync — only new or updated conversations are exported on
---
## ⚠️ Terms of Service Warning
## Contents
- [Terms of Service Warning](#terms-of-service-warning)
- [Installation](#installation)
- [First Run: Run Doctor](#first-run-run-doctor)
- [Getting Your Session Tokens](#getting-your-session-tokens)
- [The `auth` Command](#the-auth-command)
- [`.env` Setup](#env-setup)
- [ChatGPT Projects](#chatgpt-projects)
- [Claude Code Sessions](#claude-code-sessions)
- [Codex Sessions](#codex-sessions)
- [Scheduling a Daily Run](#scheduling-a-daily-run)
- [Output Structure](#output-structure)
- [CLI Reference](#cli-reference)
- [How the Cache Works](#how-the-cache-works)
- [Troubleshooting](#troubleshooting)
- [Future Work](#future-work)
- [Security Notes](#security-notes)
---
## Terms of Service Warning
**Read this before using this tool.**
@@ -223,12 +244,26 @@ cp .env.example .env
| `JOPLIN_API_URL` | `http://localhost:41184` | Joplin API URL (change only if you've customised the port) |
| `JOPLIN_REQUEST_TIMEOUT` | `30` | Seconds before an API call times out. Increase for very large conversations. |
### Notifications
| Variable | Default | Description |
|----------|---------|-------------|
| `NTFY_TOPIC` | — | [ntfy](https://ntfy.sh) topic to push run results to. Unset disables notifications entirely. |
| `NTFY_SERVER` | `https://ntfy.sh` | Point at your own host if self-hosting. |
| `NTFY_TOKEN` | — | Bearer token, for access-controlled topics. |
| `NTFY_NOTIFY` | `always` | `always` notifies on every run, `failure` only when something failed, `off` never. |
A topic on public ntfy.sh is readable by anyone who knows its name, so
notifications carry per-provider counts and a machine name only — never
conversation titles. See [Getting notified](#getting-notified).
### Cache & logging
| Variable | Default | Description |
|----------|---------|-------------|
| `CACHE_DIR` | `./cache` | Where to store the sync manifest |
| `LOG_FILE` | `./cache/logs/exporter.log` | Log file path (`none` to disable) |
| `AI_CHAT_EXPORTER_QUIET_CWD` | — | Set to `1` to silence the launcher's warning when run from outside the repo. Read by the `ai-chat-exporter` wrapper scripts, not by Python; the scheduler installers set it, since they always set the correct working directory. |
---
@@ -589,7 +624,59 @@ Reads the local export cache and pushes each exported Markdown file to Joplin as
3. Copy the Authorization token and add `JOPLIN_API_TOKEN=<token>` to your `.env`
4. Joplin desktop must be open when you run this command
Options: `--provider [chatgpt|claude|all]`, `--project NAME`, `--dry-run`
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--project NAME`, `--dry-run`
### `sync` — Export and sync in one run
```bash
# The whole archive run: export, then push to Joplin
ai-chat-exporter sync
# One provider
ai-chat-exporter sync --provider codex
# Export only; don't touch Joplin
ai-chat-exporter sync --skip-joplin
# Joplin being closed is a warning, not a failure (used by the schedulers)
ai-chat-exporter sync --joplin-optional
```
Equivalent to `export` followed by `joplin` with the same `--provider`. Intended
for scheduled runs — see [Scheduling a Daily Run](#scheduling-a-daily-run).
Unlike the individual commands, `sync` sets a **meaningful exit code**: non-zero
if any conversation failed to export or any note failed to sync. A provider whose
listing call fails outright (an expired web session token being the usual cause)
counts its whole batch as failed. A provider that is simply unconfigured, or that
had nothing new, is ordinary success. `export` on its own always exits 0, which
is fine when you're reading the summary table and useless to a scheduler.
`--joplin-optional` downgrades an unreachable Joplin to a warning: the export has
already captured the local transcripts, and the notes are rebuilt from the cache
by the next run that finds Joplin open.
Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--since YYYY-MM-DD`, `--hidden-content [full|placeholder|omit]`, `--max-conversations N`, `--skip-joplin`, `--joplin-optional`, `--notify/--no-notify`, `--dry-run`
Note this is a deliberate subset of `export`'s options — `--format`, `--output`,
`--project`, `--download-media` and `--force` are not passed through. Use
`export` directly for those. (`--download-media` still applies from `.env`; the
flag is only a per-run override.)
### `notify` — Push-notification settings and test
```bash
# Show the current settings
ai-chat-exporter notify
# Send a test push to confirm the topic works
ai-chat-exporter notify --test
```
Shows the resolved ntfy configuration and which machine name will appear in the
title. See [Getting notified](#getting-notified) for what a scheduled run sends.
Options: `--test`
### `prune` — Delete stale export files
@@ -607,6 +694,50 @@ Joplin. Refuses to run when the manifest is empty (e.g. right after
`cache --clear`) so it can never wipe a freshly cleared archive. The `doctor`
command separately verifies that every manifest entry's file exists on disk.
### `projects` — Discover ChatGPT project IDs
```bash
# List the projects your conversations belong to
ai-chat-exporter projects
# Also inspect conversations whose listing entry doesn't name a project
ai-chat-exporter projects --deep
# Write the discovered IDs straight into .env
ai-chat-exporter projects --write
```
`CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is
invisible to the listing pass — conversations that live *only* inside that
project are never fetched at all. This reports every project your conversations
belong to, marks the ones absent from `.env`, and prints a paste-ready line.
`--deep` fetches each conversation's detail when the listing doesn't name its
project: complete, but one request per conversation, so it's slow.
Options: `--deep`, `--write`
### `canary` — Check for provider API drift
```bash
ai-chat-exporter canary
ai-chat-exporter canary --provider chatgpt
```
The web providers are undocumented internal APIs that can change shape without
notice, and the failure mode is silent — a renamed field means content is
quietly dropped rather than an error being raised. The canary fetches one
listing page and one conversation per provider and asserts only the fields the
normalizer actually depends on.
Findings are `ERROR` (a load-bearing field is missing or mistyped — the parser
will break or silently lose data) or `WARN` (something unfamiliar appeared;
worth investigating, not necessarily broken). **Exits non-zero on any ERROR**,
so it can be scheduled or run in CI. Local providers have no remote schema and
are not probed.
Options: `--provider [chatgpt|claude|all]`
### `cache` — Manage the sync manifest
```bash
@@ -706,7 +837,10 @@ No new or updated conversations since your last run. To verify: `ai-chat-exporte
See `FUTURE.md` for the full roadmap. Current priorities:
- **Watch/scheduled mode** on the way to a headless StartOS service
- **A StartOS service** that centralises every machine's conversations into one
corpus and owns the Joplin connection, so each machine only has to upload
(`FUTURE.md` §8)
- **Splitting this README** into a short overview plus separate documents
---