add alias to avoid explicit python commands, scheduled runs and bug fixes
This commit is contained in:
@@ -6,6 +6,7 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
||||
## [Unreleased]
|
||||
|
||||
### Fixed
|
||||
- **The terms-of-service gate exited 0 without a terminal.** `click.prompt` raises `Abort` on a closed stdin, which the handler treated as a user Ctrl-C and exited 0 — so a scheduled run on a machine that had never acknowledged the notice would report success having archived nothing. Non-interactive invocations now exit 1 with an explanation of how to clear the gate once by hand. Found by running the new systemd unit rather than by reading the code.
|
||||
- **A single U+0085 in a transcript silently dropped a whole record.** Both local providers split session files with `str.splitlines()`, which breaks not just on `\n` but on U+0085 (NEL), U+2028 and U+2029 — all of which are legal *inside* a JSON string and are written literally by Codex (Rust does not escape non-ASCII). One NEL in captured command output shredded one record into unparseable fragments; the parser logged "skipped 3 unparseable line(s)" and lost the record. Found while exporting a real rollout. Both providers now split on `\n` only, and both have regression tests that write their fixtures with `ensure_ascii=False` — with `json.dumps`' default the hazardous characters are escaped and the bug cannot reproduce.
|
||||
- **Deleted uploads are no longer reported as permission errors.** ChatGPT's `/backend-api/files/{id}/download` answers a *missing* asset with `403 {"detail":"Forbidden"}`, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned `404 {"detail":"File not found"}` on `/files/{id}`, while assets that downloaded fine returned 200 on both in the same session, and `ChatGPT-Account-Id` made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as `expired-or-missing`. A 403 on an asset that *does* still exist is left alone as `forbidden` — that one would be a real problem.
|
||||
- **4xx errors now report why.** `_make_request` ended non-retryable statuses with `raise_for_status()`, whose curl_cffi message is `HTTP Error {code}: {reason}` — and HTTP/2 carries no reason phrase, so a refused request logged as bare `HTTP Error 403:` and the response body (the only explanation the provider gives) was discarded. The body's `detail`/`error`/`message` is now carried into the `ProviderError`, redacted and truncated. This is what made the media 403s on `GET /backend-api/files/{id}/download` undiagnosable.
|
||||
@@ -13,6 +14,9 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
||||
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
|
||||
|
||||
### Added
|
||||
- **`ai-chat-exporter` / `ai-chat-exporter.cmd` launchers — no virtualenv ceremony.** `cd` into the repo and run; the wrapper creates `.venv`, installs dependencies on first run, and reinstalls when `pyproject.toml` changes. A fresh clone goes from nothing to a working command in one step (measured: ~10s), on Linux/macOS and Windows alike, which matters for a tool meant to run on several machines. `cmd.exe` searches the current directory before `PATH`, so Windows needs no `.\` prefix. The working directory is deliberately not changed — `.env`, `cache/` and `exports/` still resolve against it, which is what lets one checkout archive different machines into different places — but the wrapper now warns when you run it from elsewhere, because a different `cache/manifest.json` silently starts a *second* archive rather than failing.
|
||||
- **`sync` command — `export` then `joplin` in one invocation, with a real exit code.** Intended for schedulers (and the "trivial add-on" FUTURE.md §7 anticipated): it exits non-zero if any conversation failed to export or any note failed to sync, so a scheduled run that achieved nothing is distinguishable from one that had nothing to do. `--skip-joplin` exports only; `--joplin-optional` downgrades an unreachable Joplin to a warning, since the export has already captured the local transcripts and the notes rebuild from the cache on the next run that finds Joplin up.
|
||||
- **Daily scheduling for both platforms.** `scheduling/install-systemd-timer.sh` (systemd user timer, `Persistent=true` so a machine that was off catches up at boot) and `scheduling/Register-AiChatSyncTask.ps1` (per-user Task Scheduler entry, `-StartWhenAvailable`). `--provider` is repeatable in both, because the right set differs per machine: the local providers need no credentials and always work unattended, while a web provider whose session token has expired would fail the job every single day and train you to ignore it.
|
||||
- **Codex CLI provider (`--provider codex`).** Archives local Codex agent transcripts from `~/.codex/sessions/**/rollout-*.jsonl` — local-only, like `claude-code`: no tokens, no rate limits, no ToS exposure. Sessions land in their own top-level `AI-Codex` Joplin notebook, with the same prose-only default, repo tags (`CODEX_REPO_TAG_IGNORE`) and multi-root scanning (`CODEX_DIR`, plus `$CODEX_HOME/sessions`).
|
||||
|
||||
Codex writes each session twice in one file and the choice between the two layers is the whole design. `response_item` records are the model-facing wire format, where a tool call arrives as *JavaScript* (`tools.exec_command({...})`) because Codex's `exec` tool is code-mode; `event_msg`/`item_completed` records are Codex's own typed items, already decoded into `CommandExecution`/`FileChange`/`Extension` with argv, cwd, exit code and output as fields. Measured over 7 sessions on 0.147.0 (2026-08-18), the typed layer is 1:1 with the raw layer for prose (91 `AgentMessage` ↔ 91 assistant messages, sharing ids) and additionally omits every piece of harness plumbing — all 51 `developer`-role messages plus the 7 `# AGENTS.md instructions…` and 1 `<environment_context>` injections — which the Claude Code provider has to strip by regex. So the typed layer is parsed for content.
|
||||
|
||||
Reference in New Issue
Block a user