fix: push a FAILED notification when a scheduled run dies without reporting

sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
This commit is contained in:
JesseMarkowitz
2026-10-05 07:14:16 -04:00
co-authored by Claude Opus 5.5
parent 64068bb19b
commit 23c6e1f512
4 changed files with 111 additions and 11 deletions
+6
View File
@@ -12,6 +12,12 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
`TestSubagentFold.test_fork_containing_its_own_spawn_call` reproduces the on-disk shape (a `fork-context-ref` record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original `RecursionError` against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block. `TestSubagentFold.test_fork_containing_its_own_spawn_call` reproduces the on-disk shape (a `fork-context-ref` record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original `RecursionError` against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block.
- **A scheduled run that crashed sent no notification, and the providers that survived said "OK".** `sync` pushes its ntfy result from the end of a run it finished, so the `RecursionError` above — and any crash, any exit before the sync starts (the ToS gate, a cache error), a launcher that cannot build its venv — sent nothing at all. Worse, the unit runs each provider as its own `sync`, so codex kept pushing a low-priority "AI archive OK" every morning for the eleven days claude-code was dead: a broken provider was indistinguishable from a quiet day.
The systemd unit's `ExecStart` is now `scheduling/run-sync.sh`, which carries the old per-provider loop and pushes a high-priority **FAILED** notification for any run that exits non-zero without the app's "Sync completed with failures" banner (printed right after the app's own push, so an app-reported failure is not reported twice). The push names the provider and, for a crash, the exception class only — `claude-code: crashed (RecursionError)` — never its message, holding to `src/notify.py`'s counts-only rule for a topic anyone can read. It honours `NTFY_NOTIFY=off` and reads `NTFY_*` the way the app does: environment first, then `.env`. Re-run `install-systemd-timer.sh` to pick it up. The Windows task has no equivalent yet.
Verified against a fake launcher and a local capture server: a crash pushes with the class and without the message, an exit before the sync pushes, an app-reported failure and a success push nothing extra, `off` pushes nothing, and the run still exits 1 if any provider failed.
- **An expired Claude session key reported a raw JSON dump instead of how to fix it.** `_make_request` routed only **401** to the auth handler (`src/providers/base.py`), and claude.ai does not use 401 — an invalid or expired `sessionKey` comes back as `403 permission_error` with `details.error_code = account_session_invalid`. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: `HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}`, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it." - **An expired Claude session key reported a raw JSON dump instead of how to fix it.** `_make_request` routed only **401** to the auth handler (`src/providers/base.py`), and claude.ai does not use 401 — an invalid or expired `sessionKey` comes back as `403 permission_error` with `details.error_code = account_session_invalid`. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: `HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}`, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it."
Measured live 2026-09-20 against `GET /api/organizations`: a valid key returns 200, while an expired key, a deliberately malformed key and **no cookie at all** return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated." Measured live 2026-09-20 against `GET /api/organizations`: a valid key returns 200, while an expired key, a deliberately malformed key and **no cookie at all** return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated."
+12 -1
View File
@@ -429,7 +429,8 @@ cache on the next run that finds Joplin up.
providers in a single action rather than one action each, because systemd providers in a single action rather than one action each, because systemd
`oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only `oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only
the last action's result. Every provider is attempted; the run still exits the last action's result. Every provider is attempted; the run still exits
non-zero if any failed. non-zero if any failed. On Linux the loop is `scheduling/run-sync.sh`, which the
unit's `ExecStart` calls.
### Getting notified ### Getting notified
@@ -457,6 +458,16 @@ high priority with an alert tag, so a failed archive is distinguishable from a
quiet one on your phone. `NTFY_NOTIFY=failure` notifies only on failure; `off` quiet one on your phone. `NTFY_NOTIFY=failure` notifies only on failure; `off`
disables it; `--notify` / `--no-notify` override per run. disables it; `--notify` / `--no-notify` override per run.
`sync` can only push from the end of a run it finished. A crash, an exit before
the sync starts (the ToS gate, a cache error) or a launcher that can't build its
venv sends nothing — and since each provider pushes separately, the ones that
succeeded still say "OK", so a dead provider looks like a quiet day. On Linux,
`scheduling/run-sync.sh` closes that gap: any run that exits non-zero without the
app having reported it gets a high-priority **FAILED** push naming the provider
and, for a crash, the exception's class (`claude-code: crashed (RecursionError)`)
— the class only, never its message, which can carry a conversation title. The
traceback is in the journal. The Windows task has no such backstop yet.
The message includes the **machine name**, which matters because both machines The message includes the **machine name**, which matters because both machines
archive into one topic. It contains counts only — never conversation titles. A archive into one topic. It contains counts only — never conversation titles. A
topic on public ntfy.sh is readable by anyone who knows its name, so if you want topic on public ntfy.sh is readable by anyone who knows its name, so if you want
+8 -8
View File
@@ -44,10 +44,12 @@ fi
[ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all") [ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all")
if [ ! -x "$REPO/ai-chat-exporter" ]; then for f in "$REPO/ai-chat-exporter" "$REPO/scheduling/run-sync.sh"; do
echo "error: $REPO/ai-chat-exporter is missing or not executable." >&2 if [ ! -x "$f" ]; then
echo "error: $f is missing or not executable." >&2
exit 1 exit 1
fi fi
done
mkdir -p "$UNIT_DIR" mkdir -p "$UNIT_DIR"
@@ -65,12 +67,10 @@ mkdir -p "$UNIT_DIR"
echo "Type=oneshot" echo "Type=oneshot"
echo "WorkingDirectory=$REPO" echo "WorkingDirectory=$REPO"
echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1" echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1"
# One ExecStart per provider would stop at the first failure, silently echo "Environment=AICHAT_SYNC_UNIT=$NAME"
# skipping the rest — an expired ChatGPT token would mean codex never runs. # run-sync.sh attempts every provider even after one fails, and pushes a
# Loop instead, so every provider is attempted and the unit still reports # FAILED notification for any run that died without sending its own.
# failure if any of them failed. echo "ExecStart=$REPO/scheduling/run-sync.sh ${PROVIDERS[*]}"
printf 'ExecStart=/bin/sh -c '\''rc=0; for p in %s; do "$0" sync --provider "$p" --joplin-optional || rc=1; done; exit $rc'\'' %s\n' \
"${PROVIDERS[*]}" "$REPO/ai-chat-exporter"
} > "$UNIT_DIR/$NAME.service" } > "$UNIT_DIR/$NAME.service"
# Persistent=true runs a missed schedule at the next boot — the machine being # Persistent=true runs a missed schedule at the next boot — the machine being
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env bash
# Run `ai-chat-exporter sync` once per provider — the ExecStart of the systemd
# unit that install-systemd-timer.sh writes.
#
# ./scheduling/run-sync.sh claude-code codex
#
# Every provider is attempted even after one fails, and the exit code is
# non-zero if any failed — one ExecStart per provider would stop at the first.
#
# The app pushes its own ntfy result, but only from the end of a run it
# finished. A crash, a non-zero exit before the sync starts (the terms-of-service
# gate, a cache error) or a launcher that can't build its venv sends nothing, and
# because each provider pushes separately, the providers that did succeed still
# send "OK" — so a broken one looks like a quiet day. This script pushes a FAILED
# notification for any run that exited non-zero without the app having reported
# it. (Its "Sync completed with failures" banner prints right after its push.)
#
# The push carries the provider, the exit code and, for a crash, the exception's
# class name — never its message. Same counts-only rule as src/notify.py: on a
# public ntfy topic anyone who guesses the name can read it, and exception text
# can carry conversation titles. The full traceback is in the journal.
set -uo pipefail
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
LAUNCHER="$REPO/ai-chat-exporter"
# NTFY_* as the app resolves them: the environment wins, then .env.
env_value() {
local name=$1 value=${!1:-}
if [ -z "$value" ] && [ -f "$REPO/.env" ]; then
value=$(sed -n "s/^[[:space:]]*$name[[:space:]]*=[[:space:]]*//p" "$REPO/.env" | tail -n 1)
value=${value%%[[:space:]]#*}
value=${value%"${value##*[![:space:]]}"}
value=${value#[\"\']}
value=${value%[\"\']}
fi
printf '%s' "$value"
}
push_failure() {
local body=$1 topic server token policy
topic=$(env_value NTFY_TOPIC)
policy=$(env_value NTFY_NOTIFY | tr '[:upper:]' '[:lower:]')
if [ -z "$topic" ] || [ "$policy" = "off" ]; then
return 0
fi
server=$(env_value NTFY_SERVER)
server=${server:-https://ntfy.sh}
token=$(env_value NTFY_TOKEN)
local args=(-fsS --max-time 15 -o /dev/null
-H "Title: AI archive FAILED - $(hostname -s)"
-H "Tags: rotating_light" -H "Priority: high"
--data-binary "$body")
[ -n "$token" ] && args+=(-H "Authorization: Bearer $token")
curl "${args[@]}" "${server%/}/$topic" \
|| echo "run-sync: could not send the failure notification" >&2
}
[ $# -eq 0 ] && set -- all
rc=0
for provider in "$@"; do
out=$(mktemp)
"$LAUNCHER" sync --provider "$provider" --joplin-optional 2>&1 | tee "$out"
status=${PIPESTATUS[0]}
if [ "$status" -ne 0 ]; then
rc=1
if ! grep -q "Sync completed with failures" "$out"; then
crash=$(grep -oE '^[A-Za-z_][A-Za-z0-9_.]*(Error|Exception)\b' "$out" | tail -n 1)
if [ -n "$crash" ]; then
reason="crashed ($crash)"
else
reason="exited $status before reporting a result"
fi
push_failure "$provider: $reason
journalctl --user -u ${AICHAT_SYNC_UNIT:-aichat-sync} -n 100"
fi
fi
rm -f "$out"
done
exit "$rc"