fix: push a FAILED notification when a scheduled run dies without reporting

sync pushes its result only from the end of a run it finished, so a crash or an early exit sent nothing, while the providers that survived kept pushing OK. The systemd unit now runs scheduling/run-sync.sh, which keeps the per-provider loop and pushes a high-priority failure, with the exception class only, for any run that exited non-zero without the app's own report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LbnmGHnFqDjyhPcCg1SEfF
This commit is contained in:
JesseMarkowitz
2026-10-05 07:14:16 -04:00
co-authored by Claude Opus 5.5
parent 64068bb19b
commit 23c6e1f512
4 changed files with 111 additions and 11 deletions
+6
View File
@@ -12,6 +12,12 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
`TestSubagentFold.test_fork_containing_its_own_spawn_call` reproduces the on-disk shape (a `fork-context-ref` record, the copied spawn call, the boilerplate-wrapped directive) and fails with the original `RecursionError` against the unfixed code. Verified against the real archive: all 85 sessions normalize, the seven forks each render as one subagent block.
- **A scheduled run that crashed sent no notification, and the providers that survived said "OK".** `sync` pushes its ntfy result from the end of a run it finished, so the `RecursionError` above — and any crash, any exit before the sync starts (the ToS gate, a cache error), a launcher that cannot build its venv — sent nothing at all. Worse, the unit runs each provider as its own `sync`, so codex kept pushing a low-priority "AI archive OK" every morning for the eleven days claude-code was dead: a broken provider was indistinguishable from a quiet day.
The systemd unit's `ExecStart` is now `scheduling/run-sync.sh`, which carries the old per-provider loop and pushes a high-priority **FAILED** notification for any run that exits non-zero without the app's "Sync completed with failures" banner (printed right after the app's own push, so an app-reported failure is not reported twice). The push names the provider and, for a crash, the exception class only — `claude-code: crashed (RecursionError)` — never its message, holding to `src/notify.py`'s counts-only rule for a topic anyone can read. It honours `NTFY_NOTIFY=off` and reads `NTFY_*` the way the app does: environment first, then `.env`. Re-run `install-systemd-timer.sh` to pick it up. The Windows task has no equivalent yet.
Verified against a fake launcher and a local capture server: a crash pushes with the class and without the message, an exit before the sync pushes, an app-reported failure and a success push nothing extra, `off` pushes nothing, and the run still exits 1 if any provider failed.
- **An expired Claude session key reported a raw JSON dump instead of how to fix it.** `_make_request` routed only **401** to the auth handler (`src/providers/base.py`), and claude.ai does not use 401 — an invalid or expired `sessionKey` comes back as `403 permission_error` with `details.error_code = account_session_invalid`. So the one message that names the cookie, its ~30-day lifetime and the DevTools path to refresh it could never fire for Claude. What the user got instead was the generic 4xx path: `HTTP 403 — error: {'type': 'permission_error', 'message': 'Invalid authorization'…}`, which reads like a permissions problem with the account and not like "your key expired, here is how to replace it."
Measured live 2026-09-20 against `GET /api/organizations`: a valid key returns 200, while an expired key, a deliberately malformed key and **no cookie at all** return byte-identical 403s carrying that code — i.e. the API treats a dead session as an absent one. This is the same mistake as the ChatGPT media 403s below: assuming 403 means "forbidden" when the service uses it for "unauthenticated."