fix: push a FAILED notification when a scheduled run dies without reporting
sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
64068bb19b
commit
23c6e1f512
@@ -12,6 +12,12 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
||||
|
||||
`TestSubagentFold.test_fork_containing_its_own_spawn_call` reproduces the on-disk shape (a `fork-context-ref` record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original `RecursionError` against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block.
|
||||
|
||||
- **A scheduled run that crashed sent no notification, and the providers that survived said "OK".** `sync` pushes its ntfy result from the end of a run it finished, so the `RecursionError` above — and any crash, any exit before the sync starts (the ToS gate, a cache error), a launcher that cannot build its venv — sent nothing at all. Worse, the unit runs each provider as its own `sync`, so codex kept pushing a low-priority "AI archive OK" every morning for the eleven days claude-code was dead: a broken provider was indistinguishable from a quiet day.
|
||||
|
||||
The systemd unit's `ExecStart` is now `scheduling/run-sync.sh`, which carries the old per-provider loop and pushes a high-priority **FAILED** notification for any run that exits non-zero without the app's "Sync completed with failures" banner (printed right after the app's own push, so an app-reported failure is not reported twice). The push names the provider and, for a crash, the exception class only — `claude-code: crashed (RecursionError)` — never its message, holding to `src/notify.py`'s counts-only rule for a topic anyone can read. It honours `NTFY_NOTIFY=off` and reads `NTFY_*` the way the app does: environment first, then `.env`. Re-run `install-systemd-timer.sh` to pick it up. The Windows task has no equivalent yet.
|
||||
|
||||
Verified against a fake launcher and a local capture server: a crash pushes with the class and without the message, an exit before the sync pushes, an app-reported failure and a success push nothing extra, `off` pushes nothing, and the run still exits 1 if any provider failed.
|
||||
|
||||
- **An expired Claude session key reported a raw JSON dump instead of how to fix it.** `_make_request` routed only **401** to the auth handler (`src/providers/base.py`), and claude.ai does not use 401 — an invalid or expired `sessionKey` comes back as `403 permission_error` with `details.error_code = account_session_invalid`. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: `HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}`, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it."
|
||||
|
||||
Measured live 2026-09-20 against `GET /api/organizations`: a valid key returns 200, while an expired key, a deliberately malformed key and **no cookie at all** return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated."
|
||||
|
||||
@@ -429,7 +429,8 @@ cache on the next run that finds Joplin up.
|
||||
providers in a single action rather than one action each, because systemd
|
||||
`oneshot` stops at the first failing `ExecStart` and Task Scheduler reports only
|
||||
the last action's result. Every provider is attempted; the run still exits
|
||||
non-zero if any failed.
|
||||
non-zero if any failed. On Linux the loop is `scheduling/run-sync.sh`, which the
|
||||
unit's `ExecStart` calls.
|
||||
|
||||
### Getting notified
|
||||
|
||||
@@ -457,6 +458,16 @@ high priority with an alert tag, so a failed archive is distinguishable from a
|
||||
quiet one on your phone. `NTFY_NOTIFY=failure` notifies only on failure; `off`
|
||||
disables it; `--notify` / `--no-notify` override per run.
|
||||
|
||||
`sync` can only push from the end of a run it finished. A crash, an exit before
|
||||
the sync starts (the ToS gate, a cache error) or a launcher that can't build its
|
||||
venv sends nothing — and since each provider pushes separately, the ones that
|
||||
succeeded still say "OK", so a dead provider looks like a quiet day. On Linux,
|
||||
`scheduling/run-sync.sh` closes that gap: any run that exits non-zero without the
|
||||
app having reported it gets a high-priority **FAILED** push naming the provider
|
||||
and, for a crash, the exception's class (`claude-code: crashed (RecursionError)`)
|
||||
— the class only, never its message, which can carry a conversation title. The
|
||||
traceback is in the journal. The Windows task has no such backstop yet.
|
||||
|
||||
The message includes the **machine name**, which matters because both machines
|
||||
archive into one topic. It contains counts only — never conversation titles. A
|
||||
topic on public ntfy.sh is readable by anyone who knows its name, so if you want
|
||||
|
||||
@@ -44,10 +44,12 @@ fi
|
||||
|
||||
[ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all")
|
||||
|
||||
if [ ! -x "$REPO/ai-chat-exporter" ]; then
|
||||
echo "error: $REPO/ai-chat-exporter is missing or not executable." >&2
|
||||
exit 1
|
||||
fi
|
||||
for f in "$REPO/ai-chat-exporter" "$REPO/scheduling/run-sync.sh"; do
|
||||
if [ ! -x "$f" ]; then
|
||||
echo "error: $f is missing or not executable." >&2
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
mkdir -p "$UNIT_DIR"
|
||||
|
||||
@@ -65,12 +67,10 @@ mkdir -p "$UNIT_DIR"
|
||||
echo "Type=oneshot"
|
||||
echo "WorkingDirectory=$REPO"
|
||||
echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1"
|
||||
# One ExecStart per provider would stop at the first failure, silently
|
||||
# skipping the rest — an expired ChatGPT token would mean codex never runs.
|
||||
# Loop instead, so every provider is attempted and the unit still reports
|
||||
# failure if any of them failed.
|
||||
printf 'ExecStart=/bin/sh -c '\''rc=0; for p in %s; do "$0" sync --provider "$p" --joplin-optional || rc=1; done; exit $rc'\'' %s\n' \
|
||||
"${PROVIDERS[*]}" "$REPO/ai-chat-exporter"
|
||||
echo "Environment=AICHAT_SYNC_UNIT=$NAME"
|
||||
# run-sync.sh attempts every provider even after one fails, and pushes a
|
||||
# FAILED notification for any run that died without sending its own.
|
||||
echo "ExecStart=$REPO/scheduling/run-sync.sh ${PROVIDERS[*]}"
|
||||
} > "$UNIT_DIR/$NAME.service"
|
||||
|
||||
# Persistent=true runs a missed schedule at the next boot — the machine being
|
||||
|
||||
Executable
+83
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env bash
|
||||
# Run `ai-chat-exporter sync` once per provider — the ExecStart of the systemd
|
||||
# unit that install-systemd-timer.sh writes.
|
||||
#
|
||||
# ./scheduling/run-sync.sh claude-code codex
|
||||
#
|
||||
# Every provider is attempted even after one fails, and the exit code is
|
||||
# non-zero if any failed — one ExecStart per provider would stop at the first.
|
||||
#
|
||||
# The app pushes its own ntfy result, but only from the end of a run it
|
||||
# finished. A crash, a non-zero exit before the sync starts (the terms-of-service
|
||||
# gate, a cache error) or a launcher that can't build its venv sends nothing, and
|
||||
# because each provider pushes separately, the providers that did succeed still
|
||||
# send "OK" — so a broken one looks like a quiet day. This script pushes a FAILED
|
||||
# notification for any run that exited non-zero without the app having reported
|
||||
# it. (Its "Sync completed with failures" banner prints right after its push.)
|
||||
#
|
||||
# The push carries the provider, the exit code and, for a crash, the exception's
|
||||
# class name — never its message. Same counts-only rule as src/notify.py: on a
|
||||
# public ntfy topic anyone who guesses the name can read it, and exception text
|
||||
# can carry conversation titles. The full traceback is in the journal.
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
LAUNCHER="$REPO/ai-chat-exporter"
|
||||
|
||||
# NTFY_* as the app resolves them: the environment wins, then .env.
|
||||
env_value() {
|
||||
local name=$1 value=${!1:-}
|
||||
if [ -z "$value" ] && [ -f "$REPO/.env" ]; then
|
||||
value=$(sed -n "s/^[[:space:]]*$name[[:space:]]*=[[:space:]]*//p" "$REPO/.env" | tail -n 1)
|
||||
value=${value%%[[:space:]]#*}
|
||||
value=${value%"${value##*[![:space:]]}"}
|
||||
value=${value#[\"\']}
|
||||
value=${value%[\"\']}
|
||||
fi
|
||||
printf '%s' "$value"
|
||||
}
|
||||
|
||||
push_failure() {
|
||||
local body=$1 topic server token policy
|
||||
topic=$(env_value NTFY_TOPIC)
|
||||
policy=$(env_value NTFY_NOTIFY | tr '[:upper:]' '[:lower:]')
|
||||
if [ -z "$topic" ] || [ "$policy" = "off" ]; then
|
||||
return 0
|
||||
fi
|
||||
server=$(env_value NTFY_SERVER)
|
||||
server=${server:-https://ntfy.sh}
|
||||
token=$(env_value NTFY_TOKEN)
|
||||
|
||||
local args=(-fsS --max-time 15 -o /dev/null
|
||||
-H "Title: AI archive FAILED - $(hostname -s)"
|
||||
-H "Tags: rotating_light" -H "Priority: high"
|
||||
--data-binary "$body")
|
||||
[ -n "$token" ] && args+=(-H "Authorization: Bearer $token")
|
||||
curl "${args[@]}" "${server%/}/$topic" \
|
||||
|| echo "run-sync: could not send the failure notification" >&2
|
||||
}
|
||||
|
||||
[ $# -eq 0 ] && set -- all
|
||||
|
||||
rc=0
|
||||
for provider in "$@"; do
|
||||
out=$(mktemp)
|
||||
"$LAUNCHER" sync --provider "$provider" --joplin-optional 2>&1 | tee "$out"
|
||||
status=${PIPESTATUS[0]}
|
||||
if [ "$status" -ne 0 ]; then
|
||||
rc=1
|
||||
if ! grep -q "Sync completed with failures" "$out"; then
|
||||
crash=$(grep -oE '^[A-Za-z_][A-Za-z0-9_.]*(Error|Exception)\b' "$out" | tail -n 1)
|
||||
if [ -n "$crash" ]; then
|
||||
reason="crashed ($crash)"
|
||||
else
|
||||
reason="exited $status before reporting a result"
|
||||
fi
|
||||
push_failure "$provider: $reason
|
||||
journalctl --user -u ${AICHAT_SYNC_UNIT:-aichat-sync} -n 100"
|
||||
fi
|
||||
fi
|
||||
rm -f "$out"
|
||||
done
|
||||
exit "$rc"
|
||||
Reference in New Issue
Block a user