From 23c6e1f51249607bf298b19bb62569cedd4202bc Mon Sep 17 00:00:00 2001 From: JesseMarkowitz Date: Mon, 5 Oct 2026 07:14:16 -0400 Subject: [PATCH] fix: push a FAILED notification when a scheduled run dies without reporting sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF --- CHANGELOG.md | 6 +++ README.md | 13 ++++- scheduling/install-systemd-timer.sh | 20 +++---- scheduling/run-sync.sh | 83 +++++++++++++++++++++++++++++ 4 files changed, 111 insertions(+), 11 deletions(-) create mode 100755 scheduling/run-sync.sh diff --git a/CHANGELOG.md b/CHANGELOG.md index 53191ce..faef6df 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,12 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). `TestSubagentFold.test_fork_containing_its_own_spawn_call` reproduces the on-disk shape (a `fork-context-ref` record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original `RecursionError` against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block. +- **A scheduled run that crashed sent no notification, and the providers that survived said "OK".** `sync` pushes its ntfy result from the end of a run it finished, so the `RecursionError` above — and any crash, any exit before the sync starts (the ToS gate, a cache error), a launcher that cannot build its venv — sent nothing at all. Worse, the unit runs each provider as its own `sync`, so codex kept pushing a low-priority "AI archive OK" every morning for the eleven days claude-code was dead: a broken provider was indistinguishable from a quiet day. + + The systemd unit's `ExecStart` is now `scheduling/run-sync.sh`, which carries the old per-provider loop and pushes a high-priority **FAILED** notification for any run that exits non-zero without the app's "Sync completed with failures" banner (printed right after the app's own push, so an app-reported failure is not reported twice). The push names the provider and, for a crash, the exception class only — `claude-code: crashed (RecursionError)` — never its message, holding to `src/notify.py`'s counts-only rule for a topic anyone can read. It honours `NTFY_NOTIFY=off` and reads `NTFY_*` the way the app does: environment first, then `.env`. Re-run `install-systemd-timer.sh` to pick it up. The Windows task has no equivalent yet. + + Verified against a fake launcher and a local capture server: a crash pushes with the class and without the message, an exit before the sync pushes, an app-reported failure and a success push nothing extra, `off` pushes nothing, and the run still exits 1 if any provider failed. + - **An expired Claude session key reported a raw JSON dump instead of how to fix it.** `_make_request` routed only **401** to the auth handler (`src/providers/base.py`), and claude.ai does not use 401 — an invalid or expired `sessionKey` comes back as `403 permission_error` with `details.error_code = account_session_invalid`. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: `HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}`, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it." Measured live 2026-09-20 against `GET /api/organizations`: a valid key returns 200, while an expired key, a deliberately malformed key and **no cookie at all** return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated." diff --git a/README.md b/README.md index c529d80..f9b8d71 100644 --- a/README.md +++ b/README.md @@ -429,7 +429,8 @@ cache on the next run that finds Joplin up. providers in a single action rather than one action each, because systemd `oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only the last action's result. Every provider is attempted; the run still exits -non-zero if any failed. +non-zero if any failed. On Linux the loop is `scheduling/run-sync.sh`, which the +unit's `ExecStart` calls. ### Getting notified @@ -457,6 +458,16 @@ high priority with an alert tag, so a failed archive is distinguishable from a quiet one on your phone. `NTFY_NOTIFY=failure` notifies only on failure; `off` disables it; `--notify` / `--no-notify` override per run. +`sync` can only push from the end of a run it finished. A crash, an exit before +the sync starts (the ToS gate, a cache error) or a launcher that can't build its +venv sends nothing — and since each provider pushes separately, the ones that +succeeded still say "OK", so a dead provider looks like a quiet day. On Linux, +`scheduling/run-sync.sh` closes that gap: any run that exits non-zero without the +app having reported it gets a high-priority **FAILED** push naming the provider +and, for a crash, the exception's class (`claude-code: crashed (RecursionError)`) +— the class only, never its message, which can carry a conversation title. The +traceback is in the journal. The Windows task has no such backstop yet. + The message includes the **machine name**, which matters because both machines archive into one topic. It contains counts only — never conversation titles. A topic on public ntfy.sh is readable by anyone who knows its name, so if you want diff --git a/scheduling/install-systemd-timer.sh b/scheduling/install-systemd-timer.sh index 053a6e6..1dfbe95 100755 --- a/scheduling/install-systemd-timer.sh +++ b/scheduling/install-systemd-timer.sh @@ -44,10 +44,12 @@ fi [ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all") -if [ ! -x "$REPO/ai-chat-exporter" ]; then - echo "error: $REPO/ai-chat-exporter is missing or not executable." >&2 - exit 1 -fi +for f in "$REPO/ai-chat-exporter" "$REPO/scheduling/run-sync.sh"; do + if [ ! -x "$f" ]; then + echo "error: $f is missing or not executable." >&2 + exit 1 + fi +done mkdir -p "$UNIT_DIR" @@ -65,12 +67,10 @@ mkdir -p "$UNIT_DIR" echo "Type=oneshot" echo "WorkingDirectory=$REPO" echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1" - # One ExecStart per provider would stop at the first failure, silently - # skipping the rest — an expired ChatGPT token would mean codex never runs. - # Loop instead, so every provider is attempted and the unit still reports - # failure if any of them failed. - printf 'ExecStart=/bin/sh -c '\''rc=0; for p in %s; do "$0" sync --provider "$p" --joplin-optional || rc=1; done; exit $rc'\'' %s\n' \ - "${PROVIDERS[*]}" "$REPO/ai-chat-exporter" + echo "Environment=AICHAT_SYNC_UNIT=$NAME" + # run-sync.sh attempts every provider even after one fails, and pushes a + # FAILED notification for any run that died without sending its own. + echo "ExecStart=$REPO/scheduling/run-sync.sh ${PROVIDERS[*]}" } > "$UNIT_DIR/$NAME.service" # Persistent=true runs a missed schedule at the next boot — the machine being diff --git a/scheduling/run-sync.sh b/scheduling/run-sync.sh new file mode 100755 index 0000000..ae232e4 --- /dev/null +++ b/scheduling/run-sync.sh @@ -0,0 +1,83 @@ +#!/usr/bin/env bash +# Run `ai-chat-exporter sync` once per provider — the ExecStart of the systemd +# unit that install-systemd-timer.sh writes. +# +# ./scheduling/run-sync.sh claude-code codex +# +# Every provider is attempted even after one fails, and the exit code is +# non-zero if any failed — one ExecStart per provider would stop at the first. +# +# The app pushes its own ntfy result, but only from the end of a run it +# finished. A crash, a non-zero exit before the sync starts (the terms-of-service +# gate, a cache error) or a launcher that can't build its venv sends nothing, and +# because each provider pushes separately, the providers that did succeed still +# send "OK" — so a broken one looks like a quiet day. This script pushes a FAILED +# notification for any run that exited non-zero without the app having reported +# it. (Its "Sync completed with failures" banner prints right after its push.) +# +# The push carries the provider, the exit code and, for a crash, the exception's +# class name — never its message. Same counts-only rule as src/notify.py: on a +# public ntfy topic anyone who guesses the name can read it, and exception text +# can carry conversation titles. The full traceback is in the journal. + +set -uo pipefail + +REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)" +LAUNCHER="$REPO/ai-chat-exporter" + +# NTFY_* as the app resolves them: the environment wins, then .env. +env_value() { + local name=$1 value=${!1:-} + if [ -z "$value" ] && [ -f "$REPO/.env" ]; then + value=$(sed -n "s/^[[:space:]]*$name[[:space:]]*=[[:space:]]*//p" "$REPO/.env" | tail -n 1) + value=${value%%[[:space:]]#*} + value=${value%"${value##*[![:space:]]}"} + value=${value#[\"\']} + value=${value%[\"\']} + fi + printf '%s' "$value" +} + +push_failure() { + local body=$1 topic server token policy + topic=$(env_value NTFY_TOPIC) + policy=$(env_value NTFY_NOTIFY | tr '[:upper:]' '[:lower:]') + if [ -z "$topic" ] || [ "$policy" = "off" ]; then + return 0 + fi + server=$(env_value NTFY_SERVER) + server=${server:-https://ntfy.sh} + token=$(env_value NTFY_TOKEN) + + local args=(-fsS --max-time 15 -o /dev/null + -H "Title: AI archive FAILED - $(hostname -s)" + -H "Tags: rotating_light" -H "Priority: high" + --data-binary "$body") + [ -n "$token" ] && args+=(-H "Authorization: Bearer $token") + curl "${args[@]}" "${server%/}/$topic" \ + || echo "run-sync: could not send the failure notification" >&2 +} + +[ $# -eq 0 ] && set -- all + +rc=0 +for provider in "$@"; do + out=$(mktemp) + "$LAUNCHER" sync --provider "$provider" --joplin-optional 2>&1 | tee "$out" + status=${PIPESTATUS[0]} + if [ "$status" -ne 0 ]; then + rc=1 + if ! grep -q "Sync completed with failures" "$out"; then + crash=$(grep -oE '^[A-Za-z_][A-Za-z0-9_.]*(Error|Exception)\b' "$out" | tail -n 1) + if [ -n "$crash" ]; then + reason="crashed ($crash)" + else + reason="exited $status before reporting a result" + fi + push_failure "$provider: $reason +journalctl --user -u ${AICHAT_SYNC_UNIT:-aichat-sync} -n 100" + fi + fi + rm -f "$out" +done +exit "$rc"