fix: push a FAILED notification when a scheduled run dies without reporting

sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
This commit is contained in:
JesseMarkowitz
2026-10-05 07:14:16 -04:00
co-authored by Claude Opus 5.5
parent 64068bb19b
commit 23c6e1f512
4 changed files with 111 additions and 11 deletions
+10 -10
View File
@@ -44,10 +44,12 @@ fi
[ ${#PROVIDERS[@]} -eq 0 ] && PROVIDERS=("all")
if [ ! -x "$REPO/ai-chat-exporter" ]; then
echo "error: $REPO/ai-chat-exporter is missing or not executable." >&2
exit 1
fi
for f in "$REPO/ai-chat-exporter" "$REPO/scheduling/run-sync.sh"; do
if [ ! -x "$f" ]; then
echo "error: $f is missing or not executable." >&2
exit 1
fi
done
mkdir -p "$UNIT_DIR"
@@ -65,12 +67,10 @@ mkdir -p "$UNIT_DIR"
echo "Type=oneshot"
echo "WorkingDirectory=$REPO"
echo "Environment=AI_CHAT_EXPORTER_QUIET_CWD=1"
# One ExecStart per provider would stop at the first failure, silently
# skipping the rest — an expired ChatGPT token would mean codex never runs.
# Loop instead, so every provider is attempted and the unit still reports
# failure if any of them failed.
printf 'ExecStart=/bin/sh -c '\''rc=0; for p in %s; do "$0" sync --provider "$p" --joplin-optional || rc=1; done; exit $rc'\'' %s\n' \
"${PROVIDERS[*]}" "$REPO/ai-chat-exporter"
echo "Environment=AICHAT_SYNC_UNIT=$NAME"
# run-sync.sh attempts every provider even after one fails, and pushes a
# FAILED notification for any run that died without sending its own.
echo "ExecStart=$REPO/scheduling/run-sync.sh ${PROVIDERS[*]}"
} > "$UNIT_DIR/$NAME.service"
# Persistent=true runs a missed schedule at the next boot — the machine being
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env bash
# Run `ai-chat-exporter sync` once per provider — the ExecStart of the systemd
# unit that install-systemd-timer.sh writes.
#
# ./scheduling/run-sync.sh claude-code codex
#
# Every provider is attempted even after one fails, and the exit code is
# non-zero if any failed — one ExecStart per provider would stop at the first.
#
# The app pushes its own ntfy result, but only from the end of a run it
# finished. A crash, a non-zero exit before the sync starts (the terms-of-service
# gate, a cache error) or a launcher that can't build its venv sends nothing, and
# because each provider pushes separately, the providers that did succeed still
# send "OK" — so a broken one looks like a quiet day. This script pushes a FAILED
# notification for any run that exited non-zero without the app having reported
# it. (Its "Sync completed with failures" banner prints right after its push.)
#
# The push carries the provider, the exit code and, for a crash, the exception's
# class name — never its message. Same counts-only rule as src/notify.py: on a
# public ntfy topic anyone who guesses the name can read it, and exception text
# can carry conversation titles. The full traceback is in the journal.
set -uo pipefail
REPO="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
LAUNCHER="$REPO/ai-chat-exporter"
# NTFY_* as the app resolves them: the environment wins, then .env.
env_value() {
local name=$1 value=${!1:-}
if [ -z "$value" ] && [ -f "$REPO/.env" ]; then
value=$(sed -n "s/^[[:space:]]*$name[[:space:]]*=[[:space:]]*//p" "$REPO/.env" | tail -n 1)
value=${value%%[[:space:]]#*}
value=${value%"${value##*[![:space:]]}"}
value=${value#[\"\']}
value=${value%[\"\']}
fi
printf '%s' "$value"
}
push_failure() {
local body=$1 topic server token policy
topic=$(env_value NTFY_TOPIC)
policy=$(env_value NTFY_NOTIFY | tr '[:upper:]' '[:lower:]')
if [ -z "$topic" ] || [ "$policy" = "off" ]; then
return 0
fi
server=$(env_value NTFY_SERVER)
server=${server:-https://ntfy.sh}
token=$(env_value NTFY_TOKEN)
local args=(-fsS --max-time 15 -o /dev/null
-H "Title: AI archive FAILED - $(hostname -s)"
-H "Tags: rotating_light" -H "Priority: high"
--data-binary "$body")
[ -n "$token" ] && args+=(-H "Authorization: Bearer $token")
curl "${args[@]}" "${server%/}/$topic" \
|| echo "run-sync: could not send the failure notification" >&2
}
[ $# -eq 0 ] && set -- all
rc=0
for provider in "$@"; do
out=$(mktemp)
"$LAUNCHER" sync --provider "$provider" --joplin-optional 2>&1 | tee "$out"
status=${PIPESTATUS[0]}
if [ "$status" -ne 0 ]; then
rc=1
if ! grep -q "Sync completed with failures" "$out"; then
crash=$(grep -oE '^[A-Za-z_][A-Za-z0-9_.]*(Error|Exception)\b' "$out" | tail -n 1)
if [ -n "$crash" ]; then
reason="crashed ($crash)"
else
reason="exited $status before reporting a result"
fi
push_failure "$provider: $reason
journalctl --user -u ${AICHAT_SYNC_UNIT:-aichat-sync} -n 100"
fi
fi
rm -f "$out"
done
exit "$rc"