add alias to avoid explicit python commands, scheduled runs and bug fixes

This commit is contained in:
JesseMarkowitz
2026-08-18 09:59:15 -04:00
parent fe5ed341ad
commit 999429f61e
9 changed files with 737 additions and 8 deletions
+4
View File
@@ -6,6 +6,7 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
## [Unreleased]
### Fixed
- **The terms-of-service gate exited 0 without a terminal.** `click.prompt` raises `Abort` on a closed stdin, which the handler treated as a user Ctrl-C and exited 0 — so a scheduled run on a machine that had never acknowledged the notice would report success having archived nothing. Non-interactive invocations now exit 1 with an explanation of how to clear the gate once by hand. Found by running the new systemd unit rather than by reading the code.
- **A single U+0085 in a transcript silently dropped a whole record.** Both local providers split session files with `str.splitlines()`, which breaks not just on `\n` but on U+0085 (NEL), U+2028 and U+2029 — all of which are legal *inside* a JSON string and are written literally by Codex (Rust does not escape non-ASCII). One NEL in captured command output shredded one record into unparseable fragments; the parser logged "skipped 3 unparseable line(s)" and lost the record. Found while exporting a real rollout. Both providers now split on `\n` only, and both have regression tests that write their fixtures with `ensure_ascii=False` — with `json.dumps`' default the hazardous characters are escaped and the bug cannot reproduce.
- **Deleted uploads are no longer reported as permission errors.** ChatGPT's `/backend-api/files/{id}/download` answers a *missing* asset with `403 {"detail":"Forbidden"}`, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned `404 {"detail":"File not found"}` on `/files/{id}`, while assets that downloaded fine returned 200 on both in the same session, and `ChatGPT-Account-Id` made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as `expired-or-missing`. A 403 on an asset that *does* still exist is left alone as `forbidden` — that one would be a real problem.
- **4xx errors now report why.** `_make_request` ended non-retryable statuses with `raise_for_status()`, whose curl_cffi message is `HTTP Error {code}: {reason}` — and HTTP/2 carries no reason phrase, so a refused request logged as bare `HTTP Error 403:` and the response body (the only explanation the provider gives) was discarded. The body's `detail`/`error`/`message` is now carried into the `ProviderError`, redacted and truncated. This is what made the media 403s on `GET /backend-api/files/{id}/download` undiagnosable.
@@ -13,6 +14,9 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
### Added
- **`ai-chat-exporter` / `ai-chat-exporter.cmd` launchers — no virtualenv ceremony.** `cd` into the repo and run; the wrapper creates `.venv`, installs dependencies on first run, and reinstalls when `pyproject.toml` changes. A fresh clone goes from nothing to a working command in one step (measured: ~10s), on Linux/macOS and Windows alike, which matters for a tool meant to run on several machines. `cmd.exe` searches the current directory before `PATH`, so Windows needs no `.\` prefix. The working directory is deliberately not changed — `.env`, `cache/` and `exports/` still resolve against it, which is what lets one checkout archive different machines into different places — but the wrapper now warns when you run it from elsewhere, because a different `cache/manifest.json` silently starts a *second* archive rather than failing.
- **`sync` command — `export` then `joplin` in one invocation, with a real exit code.** Intended for schedulers (and the "trivial add-on" FUTURE.md §7 anticipated): it exits non-zero if any conversation failed to export or any note failed to sync, so a scheduled run that achieved nothing is distinguishable from one that had nothing to do. `--skip-joplin` exports only; `--joplin-optional` downgrades an unreachable Joplin to a warning, since the export has already captured the local transcripts and the notes rebuild from the cache on the next run that finds Joplin up.
- **Daily scheduling for both platforms.** `scheduling/install-systemd-timer.sh` (systemd user timer, `Persistent=true` so a machine that was off catches up at boot) and `scheduling/Register-AiChatSyncTask.ps1` (per-user Task Scheduler entry, `-StartWhenAvailable`). `--provider` is repeatable in both, because the right set differs per machine: the local providers need no credentials and always work unattended, while a web provider whose session token has expired would fail the job every single day and train you to ignore it.
- **Codex CLI provider (`--provider codex`).** Archives local Codex agent transcripts from `~/.codex/sessions/**/rollout-*.jsonl` — local-only, like `claude-code`: no tokens, no rate limits, no ToS exposure. Sessions land in their own top-level `AI-Codex` Joplin notebook, with the same prose-only default, repo tags (`CODEX_REPO_TAG_IGNORE`) and multi-root scanning (`CODEX_DIR`, plus `$CODEX_HOME/sessions`).
Codex writes each session twice in one file and the choice between the two layers is the whole design. `response_item` records are the model-facing wire format, where a tool call arrives as *JavaScript* (`tools.exec_command({...})`) because Codex's `exec` tool is code-mode; `event_msg`/`item_completed` records are Codex's own typed items, already decoded into `CommandExecution`/`FileChange`/`Extension` with argv, cwd, exit code and output as fields. Measured over 7 sessions on 0.147.0 (2026-08-18), the typed layer is 1:1 with the raw layer for prose (91 `AgentMessage` ↔ 91 assistant messages, sharing ids) and additionally omits every piece of harness plumbing — all 51 `developer`-role messages plus the 7 `# AGENTS.md instructions…` and 1 `<environment_context>` injections — which the Claude Code provider has to strip by regex. So the typed layer is parsed for content.
+34
View File
@@ -312,6 +312,40 @@ session-token freshness without a browser. There is no headless context to
keep fresh; the weekly manual DevTools refresh is acceptable. (Local cookie
extraction remains a dead end regardless — see #2.)
### REOPENED as a TODO (2026-08-18) — centralization, not durability
Worth revisiting, for a reason the 2026-06-28 decision did not weigh. That
decision rested on "the source conversations live in the providers' clouds and
can be re-downloaded". **That is no longer true of half the providers.**
`claude-code` (shipped v0.6.0) and `codex` (shipped 2026-08-18) read transcripts
that exist *only* on the machine that produced them — Codex prunes its rollout
files, and neither is recoverable from any cloud. The re-download premise now
covers the web providers only.
The new motivation is consolidation rather than durability: work is split across
machines — coding sessions (`claude-code`, `codex`) on the Linux box, web chats
(`chatgpt`, `claude`) on the Windows box — and each archives to its own local
`exports/` + Joplin. A StartOS service would give one server-side corpus of all
conversations from everywhere, instead of per-machine islands that only meet
inside Joplin.
What this would need, and what it would *not*:
- **Not** a headless web-provider login. The hard sub-problem the original drop
retired stays retired: the web providers can keep running interactively on the
machine that has the browser, pushing their output to the server. Only the
local providers need to run server-side, and they need no tokens at all.
- Machines would need to reach the server — a sync/push step, or the service
reading transcript directories exported from each machine.
- Per-machine identity in the corpus, which the exporter currently does not
track: `claude_code.resolve_roots` deliberately merges multiple roots with "no
per-machine label". Centralizing would make that label load-bearing.
- Conflict handling for one conversation seen by two machines. The cache
manifest is per-machine today.
Meanwhile Joplin is sufficient — it already syncs (encrypted) offsite, and it is
where the archive is actually read. This is a "nice eventually", not a gap.
## 9. Token Validity on `doctor` — IMPLEMENTED (2026-06-28)
Shipped: `doctor` now adds a "ChatGPT token active" check via
+108 -8
View File
@@ -32,10 +32,31 @@ This tool is designed for a single user backing up their own conversations. Do n
```bash
git clone <repo-url>
cd ai-chat-exporter
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cd AIChatExporter
./ai-chat-exporter doctor
```
That's the whole install. The `ai-chat-exporter` wrapper creates `.venv` and
installs dependencies on first run, and reinstalls whenever `pyproject.toml`
changes — so there is no `python3 -m venv` / `source .venv/bin/activate` to
remember, on this machine or the next one you clone onto. Every command in this
README works the same way:
```bash
./ai-chat-exporter export --provider all
./ai-chat-exporter sync
```
The wrapper does **not** change directory: `.env`, `cache/` and `exports/` all
resolve against your current directory, which is what lets one checkout archive
different machines into different places. Run it from the repo. (It warns if you
don't, because a different working directory means a different `cache/manifest.json`
— which would re-export everything and orphan your existing Joplin notes.)
Prefer the traditional route, or want the test dependencies? That still works:
```bash
python3 -m venv .venv && source .venv/bin/activate && pip install -e ".[dev]"
```
### Windows
@@ -44,12 +65,28 @@ No admin access required. Run these in **Command Prompt** (`cmd.exe`) — it's t
```bat
git clone <repo-url>
cd ai-chat-exporter
python -m venv .venv
.venv\Scripts\activate
pip install -e ".[dev]"
cd AIChatExporter
ai-chat-exporter doctor
```
`ai-chat-exporter.cmd` does the same bootstrap as the POSIX wrapper — it creates
`.venv` and installs dependencies on first run. No `python -m venv`, no
`.venv\Scripts\activate`.
**How you invoke it differs between the two Windows shells:**
| Shell | Command |
|---|---|
| Command Prompt (`cmd.exe`) | `ai-chat-exporter export --provider all` |
| PowerShell | `.\ai-chat-exporter.cmd export --provider all` |
`cmd.exe` searches the current directory before `PATH` and resolves the bare name
through `PATHEXT`, so it finds `ai-chat-exporter.cmd` with no prefix and no
extension. PowerShell deliberately does *not* search the current directory, and
`.\ai-chat-exporter` there would resolve to the extensionless POSIX script, which
PowerShell cannot execute — so name the `.cmd` explicitly. Command Prompt is the
simpler of the two here.
All `ai-chat-exporter` commands work identically in Command Prompt.
**Using PowerShell instead?** If you prefer PowerShell, you may need to allow script execution first (one-time, current user only):
@@ -284,6 +321,69 @@ Three things differ from Claude Code, all forced by how Codex stores its data:
---
## Scheduling a Daily Run
`sync` chains `export` then `joplin` in one invocation, which is what a scheduler
wants — one command, and a meaningful exit code so a failed run is visible
instead of silent.
```bash
./ai-chat-exporter sync # every configured provider
./ai-chat-exporter sync --provider codex # just one
```
### Linux (systemd user timer)
```bash
./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
./scheduling/install-systemd-timer.sh --uninstall
```
Defaults to 09:00 daily (`--time 21:30` to change). `Persistent=true` means a
machine that was off at the scheduled time runs the archive at next boot rather
than skipping the day. To archive while logged out, `loginctl enable-linger $USER`.
Check on it with `systemctl --user list-timers aichat-sync.timer` and
`journalctl --user -u aichat-sync.service -n 50`.
### Windows (Task Scheduler)
```powershell
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
```
Per-user task, no admin rights needed. `-StartWhenAvailable` is the counterpart
of systemd's `Persistent=true`.
### What to know before you rely on it
**List only the providers that work unattended on that machine.** The local
providers (`claude-code`, `codex`) need no credentials and always work. The web
providers depend on a session token that expires and can only be refreshed by
hand via DevTools — so on a machine where that token is stale, scheduling them
means a failed run every single day, which is a good way to learn to ignore
failures you actually want to see. That is why `--provider` is repeatable in both
installers: schedule the coding machine for `claude-code` + `codex`, the browser
machine for `chatgpt` + `claude`.
**Both installers pass `--joplin-optional`.** If Joplin desktop isn't running,
the sync warns instead of failing: the export has already captured the local
transcripts (the part that can disappear), and the notes are rebuilt from the
cache on the next run that finds Joplin up.
**One provider failing does not skip the rest.** Both installers loop over the
providers in a single action rather than one action each, because systemd
`oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only
the last action's result. Every provider is attempted; the run still exits
non-zero if any failed.
**Acknowledge the ToS notice once, interactively.** It's stored in the cache
manifest per machine. Until then a scheduled run exits 1 with an explanation
rather than hanging on a prompt no one can answer.
---
## Output Structure
All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks.
+67
View File
@@ -0,0 +1,67 @@
#!/usr/bin/env bash
# ai-chat-exporter — run the exporter without activating a virtualenv.
#
# cd /path/to/AIChatExporter
# ./ai-chat-exporter export --provider all
#
# Creates .venv and installs dependencies on first run, so a fresh clone on a
# new machine needs no `python3 -m venv` / `source .venv/bin/activate` ceremony.
# Reinstalls automatically when pyproject.toml changes.
#
# Working directory is deliberately NOT changed: EXPORT_DIR, CACHE_DIR and .env
# discovery are all relative to your current directory, which is what lets the
# same checkout archive different machines into different places. Run it from
# the repo (see the warning below) unless you mean otherwise.
#
# Windows equivalent: ai-chat-exporter.cmd (same directory).
set -euo pipefail
DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
VENV="$DIR/.venv"
PY="$VENV/bin/python"
STAMP="$VENV/.deps-stamp"
# The cache manifest is read from ./cache by default. Running from somewhere
# else silently starts a *new, empty* archive rather than failing, which would
# re-export everything and orphan the existing Joplin notes. Warn, don't block —
# a deliberate second archive is a legitimate thing to want.
if [ "$PWD" != "$DIR" ] && [ -z "${AI_CHAT_EXPORTER_QUIET_CWD:-}" ]; then
echo "warning: running from $PWD, not $DIR" >&2
echo " cache/ and exports/ resolve against the current directory," >&2
echo " so this may start a separate archive. Set" >&2
echo " AI_CHAT_EXPORTER_QUIET_CWD=1 to silence this." >&2
fi
find_python() {
for candidate in python3 python; do
if command -v "$candidate" >/dev/null 2>&1; then
# Needs >=3.11 (pyproject requires-python).
if "$candidate" -c 'import sys; sys.exit(0 if sys.version_info >= (3, 11) else 1)' 2>/dev/null; then
command -v "$candidate"
return 0
fi
fi
done
return 1
}
if [ ! -x "$PY" ]; then
BOOTSTRAP_PY="$(find_python)" || {
echo "error: no python3 >= 3.11 found on PATH — install it and re-run." >&2
exit 1
}
echo "Creating virtualenv in $VENV …" >&2
"$BOOTSTRAP_PY" -m venv "$VENV"
rm -f "$STAMP"
fi
# Install (or refresh) dependencies when the venv is new or pyproject changed.
if [ ! -f "$STAMP" ] || [ "$DIR/pyproject.toml" -nt "$STAMP" ]; then
echo "Installing dependencies …" >&2
"$PY" -m pip install --quiet --upgrade pip
"$PY" -m pip install --quiet -e "$DIR"
touch "$STAMP"
fi
exec "$PY" -m src.main "$@"
+83
View File
@@ -0,0 +1,83 @@
@echo off
rem ai-chat-exporter.cmd - run the exporter without activating a virtualenv.
rem
rem Command Prompt (cmd.exe) - the documented way:
rem cd C:\path\to\AIChatExporter
rem ai-chat-exporter export --provider all
rem
rem cmd.exe searches the current directory before PATH and resolves the bare
rem name through PATHEXT (which includes .CMD), so no ".\" and no extension are
rem needed. The extensionless POSIX sibling is ignored: it is not in PATHEXT.
rem
rem PowerShell does NOT search the current directory, and ".\ai-chat-exporter"
rem there would resolve to the extensionless POSIX script, which PowerShell
rem cannot run. In PowerShell, name this file explicitly:
rem .\ai-chat-exporter.cmd export --provider all
rem
rem Creates .venv and installs dependencies on first run, so a fresh clone needs
rem no "python -m venv" / ".venv\Scripts\activate" ceremony. Reinstalls
rem automatically when pyproject.toml changes.
rem
rem The working directory is deliberately NOT changed - EXPORT_DIR, CACHE_DIR
rem and .env discovery are all relative to it. POSIX equivalent: ai-chat-exporter
setlocal enabledelayedexpansion
set "DIR=%~dp0"
if "%DIR:~-1%"=="\" set "DIR=%DIR:~0,-1%"
set "VENV=%DIR%\.venv"
set "PY=%VENV%\Scripts\python.exe"
set "STAMP=%VENV%\.deps-stamp"
rem See the POSIX script for why this warns rather than blocks: cache\ resolves
rem against the current directory, so the wrong one silently starts a second
rem archive instead of failing.
if /i not "%CD%"=="%DIR%" if "%AI_CHAT_EXPORTER_QUIET_CWD%"=="" (
echo warning: running from %CD%, not %DIR% 1>&2
echo cache\ and exports\ resolve against the current directory, 1>&2
echo so this may start a separate archive. Set 1>&2
echo AI_CHAT_EXPORTER_QUIET_CWD=1 to silence this. 1>&2
)
if not exist "%PY%" (
echo Creating virtualenv in %VENV% ... 1>&2
rem The py launcher is the reliable way to get a specific version; fall back
rem to whatever "python" is if it is not installed.
where py >nul 2>&1
if !errorlevel! equ 0 (
py -3 -m venv "%VENV%"
) else (
python -m venv "%VENV%"
)
if not exist "%PY%" (
echo error: could not create a virtualenv - install Python 3.11+ from 1>&2
echo python.org or the Microsoft Store, then re-run. 1>&2
exit /b 1
)
if exist "%STAMP%" del "%STAMP%"
)
rem Staleness check, in pure batch: %%~tF is the file's last-modified stamp, so
rem storing it and comparing strings needs no external process. The obvious
rem alternative - asking PowerShell to compare timestamps - costs ~1s of
rem interpreter startup on *every* command, which is a lot to pay to almost
rem always learn that nothing changed.
set "PYPROJ_TIME="
for %%F in ("%DIR%\pyproject.toml") do set "PYPROJ_TIME=%%~tF"
set "SAVED_TIME="
if exist "%STAMP%" set /p SAVED_TIME=<"%STAMP%"
if not "%SAVED_TIME%"=="%PYPROJ_TIME%" (
echo Installing dependencies ... 1>&2
"%PY%" -m pip install --quiet --upgrade pip
"%PY%" -m pip install --quiet -e "%DIR%"
if !errorlevel! neq 0 (
echo error: dependency installation failed. 1>&2
exit /b !errorlevel!
)
> "%STAMP%" echo !PYPROJ_TIME!
)
"%PY%" -m src.main %*
exit /b %errorlevel%
+83
View File
@@ -0,0 +1,83 @@
<#
.SYNOPSIS
Register a Windows scheduled task that runs `ai-chat-exporter sync` daily.
.DESCRIPTION
The Windows counterpart to install-systemd-timer.sh. Creates a per-user task
(no admin rights needed) that runs the exporter from this repository, with
the working directory set to the repo so .env, cache\ and exports\ resolve
exactly as they do for an interactive run.
-Provider is repeatable. The CLI takes one provider per run, so each becomes
its own action, executed in order. List only the providers that work
unattended on this machine: a web provider whose session token has expired
fails the task every day, which trains you to ignore the failures you
actually want to notice.
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Provider chatgpt,claude
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Provider all -Time 21:30
.EXAMPLE
.\scheduling\Register-AiChatSyncTask.ps1 -Unregister
#>
[CmdletBinding()]
param(
[string[]]$Provider = @('all'),
[string]$Time = '09:00',
[string]$TaskName = 'AiChatExporterSync',
[switch]$Unregister
)
$ErrorActionPreference = 'Stop'
$repo = Split-Path -Parent $PSScriptRoot
$launcher = Join-Path $repo 'ai-chat-exporter.cmd'
if ($Unregister) {
Unregister-ScheduledTask -TaskName $TaskName -Confirm:$false -ErrorAction SilentlyContinue
Write-Host "Removed scheduled task '$TaskName'."
return
}
if (-not (Test-Path $launcher)) {
throw "Launcher not found at $launcher"
}
# A single action looping over the providers, rather than one action each.
# Task Scheduler runs multiple actions in order but reports only the last one's
# result, so a failure in an earlier provider would be invisible. The loop keeps
# going after a failure and propagates a non-zero exit code.
$loop = ($Provider | ForEach-Object { "`"$launcher`" sync --provider $_ --joplin-optional || set RC=1" }) -join ' & '
$actions = New-ScheduledTaskAction -Execute 'cmd.exe' `
-Argument "/c set RC=0 & $loop & exit /b %RC%" `
-WorkingDirectory $repo
$trigger = New-ScheduledTaskTrigger -Daily -At $Time
# StartWhenAvailable is the counterpart of systemd's Persistent=true: a machine
# that was asleep at the scheduled time runs the archive when it wakes, rather
# than skipping the day entirely.
$settings = New-ScheduledTaskSettingsSet `
-StartWhenAvailable `
-DontStopIfGoingOnBatteries `
-AllowStartIfOnBatteries `
-ExecutionTimeLimit (New-TimeSpan -Hours 2)
Register-ScheduledTask -TaskName $TaskName `
-Action $actions `
-Trigger $trigger `
-Settings $settings `
-Description 'Export AI chat history and sync it to Joplin' `
-Force | Out-Null
Write-Host "Registered '$TaskName' - daily at $Time for: $($Provider -join ', ')"
Write-Host ''
Write-Host 'Next steps:'
Write-Host " * Run it once now: Start-ScheduledTask -TaskName $TaskName"
Write-Host " * Check the result: Get-ScheduledTaskInfo -TaskName $TaskName"
Write-Host " * Read the log: Get-Content '$repo\cache\logs\exporter.log' -Tail 50"
Write-Host ' * The terms-of-service notice must have been acknowledged'
Write-Host ' interactively once on this machine, or the task exits 1.'
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Install a systemd *user* timer that runs `ai-chat-exporter sync` on a schedule.
#
# ./scheduling/install-systemd-timer.sh --provider claude-code --provider codex
# ./scheduling/install-systemd-timer.sh --provider all --time 21:30
# ./scheduling/install-systemd-timer.sh --uninstall
#
# User units (not system units) are the right scope: the archive is per-user,
# the .env holds that user's session tokens, and Joplin runs in their session.
#
# --provider is repeatable. The CLI takes one provider per run, so each one
# becomes its own ExecStart line, executed in order. Prefer listing only the
# providers that actually work unattended on this machine — a web provider with
# an expired session token will fail the unit every single day, which trains you
# to ignore the failure you actually want to see.
set -euo pipefail
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
UNIT_DIR="$HOME/.config/systemd/user"
NAME="aichat-sync"
TIME="09:00"
PROVIDERS=()
UNINSTALL=0
while [ $# -gt 0 ]; do
case "$1" in
--provider) PROVIDERS+=("$2"); shift 2 ;;
--time) TIME="$2"; shift 2 ;;
--name) NAME="$2"; shift 2 ;;
--uninstall) UNINSTALL=1; shift ;;
-h|--help) sed -n '2,20p' "${BASH_SOURCE[0]}"; exit 0 ;;
*) echo "unknown argument: $1" >&2; exit 1 ;;
esac
done
if [ "$UNINSTALL" -eq 1 ]; then
systemctl --user disable --now "$NAME.timer" 2>/dev/null || true
rm -f "$UNIT_DIR/$NAME.timer" "$UNIT_DIR/$NAME.service"
systemctl --user daemon-reload
echo "Removed $NAME.timer and $NAME.service."
exit 0
fi
[ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all")
if [ ! -x "$REPO/ai-chat-exporter" ]; then
echo "error: $REPO/ai-chat-exporter is missing or not executable." >&2
exit 1
fi
mkdir -p "$UNIT_DIR"
# WorkingDirectory is the point of the whole unit: .env, cache/ and exports/ all
# resolve against it, so the scheduled run writes to the same archive an
# interactive run from this directory would.
{
echo "[Unit]"
echo "Description=AI chat archive sync"
echo "Documentation=file://$REPO/README.md"
echo "After=network-online.target"
echo "Wants=network-online.target"
echo
echo "[Service]"
echo "Type=oneshot"
echo "WorkingDirectory=$REPO"
echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1"
# One ExecStart per provider would stop at the first failure, silently
# skipping the rest — an expired ChatGPT token would mean codex never runs.
# Loop instead, so every provider is attempted and the unit still reports
# failure if any of them failed.
printf 'ExecStart=/bin/sh -c '\''rc=0; for p in %s; do "$0" sync --provider "$p" --joplin-optional || rc=1; done; exit $rc'\'' %s\n' \
"${PROVIDERS[*]}" "$REPO/ai-chat-exporter"
} > "$UNIT_DIR/$NAME.service"
# Persistent=true runs a missed schedule at the next boot — the machine being
# off at 09:00 should delay the archive, not skip it. RandomizedDelaySec keeps
# the web providers from being hit at exactly the same second every day.
cat > "$UNIT_DIR/$NAME.timer" <<EOF
[Unit]
Description=Run the AI chat archive sync daily
[Timer]
OnCalendar=*-*-* $TIME:00
Persistent=true
RandomizedDelaySec=300
[Install]
WantedBy=timers.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now "$NAME.timer"
echo "Installed $NAME.timer — daily at $TIME for: ${PROVIDERS[*]}"
echo
systemctl --user list-timers "$NAME.timer" --no-pager || true
cat <<EOF
Next steps:
• Timers only run while you have a session. To archive when logged out:
loginctl enable-linger $USER
• Run it once by hand to confirm:
systemctl --user start $NAME.service
• Read the log:
journalctl --user -u $NAME.service -n 50
• The terms-of-service notice must have been acknowledged interactively at
least once on this machine, or the unit exits 1 with an explanation.
EOF
+129
View File
@@ -99,6 +99,19 @@ def cli(ctx: click.Context, verbose: bool, quiet: bool, debug: bool, no_log_file
# ToS gate: must happen before any command executes
if not cache.is_tos_acknowledged():
# Non-interactive (cron, systemd timer, Task Scheduler): click.prompt
# raises Abort on a closed stdin, which used to exit 0 — a scheduled
# run would report success having archived nothing. Fail loudly instead
# and tell the operator how to clear the gate once, by hand.
if not sys.stdin.isatty():
err_console.print(
"[red]Terms-of-service notice has not been acknowledged, and there is "
"no terminal to ask on.[/red]\n"
"Run any command once interactively (e.g. 'ai-chat-exporter doctor') "
"and type 'yes' to acknowledge. The acknowledgement is stored in the "
"cache manifest and applies to every later run on this machine."
)
sys.exit(1)
try:
answer = click.prompt(TOS_NOTICE, default="", show_default=False).strip().lower()
except (click.Abort, KeyboardInterrupt):
@@ -990,6 +1003,10 @@ def export(
if cache.exported_at(prov_name, c.get("id") or c.get("uuid", "")) < campaign_at
)
# `sync` reads this to decide the process exit code — a scheduled run that
# exported nothing because a token expired must not look like success.
ctx.obj["last_export_summary"] = summary
if not dry_run:
_print_export_summary(summary)
if force:
@@ -1179,6 +1196,116 @@ def _print_export_summary(summary: dict[str, dict[str, int]]) -> None:
console.print(table)
# ──────────────────────────────────────────────────────────────────────────────
# sync command
# ──────────────────────────────────────────────────────────────────────────────
@cli.command()
@click.option(
"--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all",
show_default=True,
help="Which provider(s) to export and sync.",
)
@click.option("--since", default=None, help="Only export conversations updated after this date (YYYY-MM-DD).")
@click.option(
"--hidden-content",
type=click.Choice(["full", "placeholder", "omit"], case_sensitive=False),
default=None,
help="Overrides EXPORTER_HIDDEN_CONTENT for this run.",
)
@click.option(
"--max-conversations",
type=click.IntRange(min=1),
default=None,
help="Cap how many conversations are downloaded this run (per provider).",
)
@click.option("--skip-joplin", is_flag=True, help="Export only; do not sync to Joplin.")
@click.option(
"--joplin-optional",
is_flag=True,
help=(
"Treat an unreachable Joplin as a warning rather than a failure. For "
"scheduled runs, where Joplin desktop may simply not be open yet."
),
)
@click.option("--dry-run", is_flag=True, help="Show what would happen without writing or sending anything.")
@click.pass_context
def sync(
ctx: click.Context,
provider: str,
since: str | None,
hidden_content: str | None,
max_conversations: int | None,
skip_joplin: bool,
joplin_optional: bool,
dry_run: bool,
) -> None:
"""Export, then sync to Joplin — the whole archive run in one command.
Equivalent to `export` followed by `joplin` with the same --provider, which
is what a scheduled run wants: one line in a systemd timer, cron entry or
Windows scheduled task.
Unlike the individual commands, this one sets a meaningful exit code: it
exits non-zero if any conversation failed to export or any note failed to
sync. A provider whose listing call fails outright — an expired web session
token being the usual cause — counts its whole batch as failed, so a
scheduler sees a real failure rather than a silent no-op.
Note this does *not* flag a provider that is simply unconfigured, or one
that legitimately had nothing new to export; both are ordinary success.
"""
failures: list[str] = []
ctx.invoke(
export,
provider=provider,
since=since,
hidden_content=hidden_content,
max_conversations=max_conversations,
dry_run=dry_run,
)
export_summary = ctx.obj.get("last_export_summary") or {}
for prov_name, counts in export_summary.items():
if counts.get("failed"):
failures.append(f"{prov_name}: {counts['failed']} conversation(s) failed to export")
if skip_joplin:
console.print("[dim]Skipping Joplin sync (--skip-joplin).[/dim]")
else:
try:
ctx.invoke(joplin, provider=provider, dry_run=dry_run)
except SystemExit as e:
# `joplin` exits non-zero when the desktop app isn't reachable. On a
# timer that is routine — the machine may not be unlocked yet — and
# the export (the part that captures data which can disappear) has
# already succeeded. The notes are rebuilt from the cache on the next
# run that finds Joplin up, so nothing is lost by carrying on.
if not joplin_optional or e.code in (0, None):
raise
console.print(
"[yellow]Joplin sync skipped — Joplin is not reachable "
"(--joplin-optional). Exported files are on disk; the next run "
"with Joplin open will sync them.[/yellow]"
)
else:
joplin_summary = ctx.obj.get("last_joplin_summary") or {}
for prov_name, counts in joplin_summary.items():
if counts.get("failed"):
failures.append(f"{prov_name}: {counts['failed']} note(s) failed to sync")
if failures:
err_console.print("\n[red]Sync completed with failures:[/red]")
for line in failures:
err_console.print(f" [red]•[/red] {line}")
sys.exit(1)
console.print("\n[green]Sync complete.[/green]")
# ──────────────────────────────────────────────────────────────────────────────
# list command
# ──────────────────────────────────────────────────────────────────────────────
@@ -1586,6 +1713,8 @@ def joplin(ctx: click.Context, provider: str, project_filter: str | None, dry_ru
finally:
progress.advance(task)
ctx.obj["last_joplin_summary"] = summary
if not dry_run:
_print_joplin_summary(summary)
+120
View File
@@ -415,3 +415,123 @@ class TestProjectsCommand:
assert result.exit_code == 0
env_text = (Path(fs) / ".env").read_text(encoding="utf-8")
assert "CHATGPT_PROJECT_IDS=g-p-new" in env_text
# ---------------------------------------------------------------------------
# sync command + non-interactive ToS gate
# ---------------------------------------------------------------------------
class TestSyncCommand:
"""`sync` chains export → joplin for schedulers, with a real exit code."""
def _cache(self, tmp_path) -> Cache:
cache = Cache(tmp_path)
cache.acknowledge_tos()
# Non-empty last_run so the first-run doctor gate stays out of the way.
cache.mark_exported("codex", "dummy", {"updated_at": "2024-01-01T00:00:00Z"})
return cache
def _env(self, tmp_path) -> dict:
"""A real (minimal) codex session — `export` exits 1 on no providers at
all, so an empty directory would test the wrong failure."""
import json
day = tmp_path / "sessions" / "2026" / "08" / "17"
day.mkdir(parents=True, exist_ok=True)
sid = "01a00e3f-a309-74a3-bf32-06c2cd87faa3"
records = [
{
"timestamp": "2026-08-17T05:44:06.666Z",
"type": "session_meta",
"payload": {"session_id": sid, "timestamp": "2026-08-17T05:44:06.666Z",
"cwd": str(tmp_path / "ws")},
},
{
"timestamp": "2026-08-17T05:44:09.000Z",
"type": "event_msg",
"payload": {"type": "item_completed", "item": {
"type": "UserMessage", "id": "u1",
"content": [{"type": "text", "text": "hello"}]}},
},
]
(day / f"rollout-2026-08-17T01-44-06-{sid}.jsonl").write_text(
"\n".join(json.dumps(r) for r in records), encoding="utf-8"
)
return {
"CACHE_DIR": str(tmp_path),
"EXPORT_DIR": str(tmp_path / "exports"),
"CODEX_DIR": str(tmp_path / "sessions"),
}
def test_skip_joplin_exits_zero(self, tmp_path):
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--skip-joplin"],
env=self._env(tmp_path),
)
assert result.exit_code == 0
assert "Skipping Joplin sync" in result.output
assert "Sync complete" in result.output
def test_joplin_optional_survives_unreachable_joplin(self, tmp_path):
"""Joplin being closed must not fail a scheduled run — the export is done."""
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--joplin-optional"],
env={**self._env(tmp_path), "JOPLIN_API_URL": "http://127.0.0.1:9"},
)
assert result.exit_code == 0
assert "Joplin sync skipped" in result.output
def test_unreachable_joplin_fails_without_the_flag(self, tmp_path):
self._cache(tmp_path)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex"],
env={**self._env(tmp_path), "JOPLIN_API_URL": "http://127.0.0.1:9"},
)
assert result.exit_code == 1
def test_export_failures_set_nonzero_exit(self, tmp_path, monkeypatch):
"""A scheduler must be able to tell a real run from a silent no-op."""
self._cache(tmp_path)
import src.main as main_mod
real_export = main_mod.export.callback
def fake_export(*args, **kwargs):
import click
ctx = click.get_current_context()
ctx.obj["last_export_summary"] = {
"codex": {"exported": 0, "skipped": 0, "failed": 3}
}
monkeypatch.setattr(main_mod.export, "callback", fake_export)
try:
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "sync", "--provider", "codex", "--skip-joplin"],
env=self._env(tmp_path),
)
finally:
monkeypatch.setattr(main_mod.export, "callback", real_export)
assert result.exit_code == 1
assert "3 conversation(s) failed to export" in result.output
class TestNonInteractiveTosGate:
"""Without a TTY the gate must fail loudly, not exit 0 having done nothing."""
def test_no_tty_exits_one_with_explanation(self, tmp_path, monkeypatch):
Cache(tmp_path) # fresh cache: ToS not acknowledged
monkeypatch.setattr("sys.stdin.isatty", lambda: False)
result = CliRunner(mix_stderr=True).invoke(
cli,
["--no-log-file", "doctor"],
env={"CACHE_DIR": str(tmp_path), "EXPORT_DIR": str(tmp_path / "exports")},
)
assert result.exit_code == 1
assert "no terminal to" in result.output