added support for codex provider. fixed bug in claude code export (lines split when shouldn't be)

This commit is contained in:
JesseMarkowitz
2026-08-18 08:01:34 -04:00
parent d083aea135
commit fe5ed341ad
10 changed files with 1435 additions and 40 deletions
+10
View File
@@ -38,6 +38,16 @@ CLAUDE_SESSION_KEY=
# list their names here (comma-separated). # list their names here (comma-separated).
#CLAUDE_CODE_REPO_TAG_IGNORE=some-repo,another-repo #CLAUDE_CODE_REPO_TAG_IGNORE=some-repo,another-repo
# --- Codex (local agent sessions) ---
# The codex provider reads local Codex CLI rollout files. By default it scans
# ~/.codex/sessions/ (plus $CODEX_HOME/sessions when CODEX_HOME is set).
# To scan additional roots, set a ':'-separated list of sessions roots.
#CODEX_DIR=~/.codex/sessions:/mnt/backup/laptop/.codex/sessions
#
# As with Claude Code, session titles are tagged with the git repos they
# touched. To never tag specific repos, list their names here (comma-separated).
#CODEX_REPO_TAG_IGNORE=some-repo,another-repo
# --- Output --- # --- Output ---
# Where exported Markdown files are written (default: ./exports) # Where exported Markdown files are written (default: ./exports)
EXPORT_DIR=./exports EXPORT_DIR=./exports
+11
View File
@@ -6,12 +6,23 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
## [Unreleased] ## [Unreleased]
### Fixed ### Fixed
- **A single U+0085 in a transcript silently dropped a whole record.** Both local providers split session files with `str.splitlines()`, which breaks not just on `\n` but on U+0085 (NEL), U+2028 and U+2029 — all of which are legal *inside* a JSON string and are written literally by Codex (Rust does not escape non-ASCII). One NEL in captured command output shredded one record into unparseable fragments; the parser logged "skipped 3 unparseable line(s)" and lost the record. Found while exporting a real rollout. Both providers now split on `\n` only, and both have regression tests that write their fixtures with `ensure_ascii=False` — with `json.dumps`' default the hazardous characters are escaped and the bug cannot reproduce.
- **Deleted uploads are no longer reported as permission errors.** ChatGPT's `/backend-api/files/{id}/download` answers a *missing* asset with `403 {"detail":"Forbidden"}`, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned `404 {"detail":"File not found"}` on `/files/{id}`, while assets that downloaded fine returned 200 on both in the same session, and `ChatGPT-Account-Id` made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as `expired-or-missing`. A 403 on an asset that *does* still exist is left alone as `forbidden` — that one would be a real problem. - **Deleted uploads are no longer reported as permission errors.** ChatGPT's `/backend-api/files/{id}/download` answers a *missing* asset with `403 {"detail":"Forbidden"}`, which reads like an auth failure and isn't one. Measured live 2026-08-17 across 18 such assets: every one returned `404 {"detail":"File not found"}` on `/files/{id}`, while assets that downloaded fine returned 200 on both in the same session, and `ChatGPT-Account-Id` made no difference. A 403 is now confirmed against the metadata endpoint before being reported (one extra request on the failure path only, none on success) and a confirmed-missing asset is logged as gone and counted as `expired-or-missing`. A 403 on an asset that *does* still exist is left alone as `forbidden` — that one would be a real problem.
- **4xx errors now report why.** `_make_request` ended non-retryable statuses with `raise_for_status()`, whose curl_cffi message is `HTTP Error {code}: {reason}` — and HTTP/2 carries no reason phrase, so a refused request logged as bare `HTTP Error 403:` and the response body (the only explanation the provider gives) was discarded. The body's `detail`/`error`/`message` is now carried into the `ProviderError`, redacted and truncated. This is what made the media 403s on `GET /backend-api/files/{id}/download` undiagnosable. - **4xx errors now report why.** `_make_request` ended non-retryable statuses with `raise_for_status()`, whose curl_cffi message is `HTTP Error {code}: {reason}` — and HTTP/2 carries no reason phrase, so a refused request logged as bare `HTTP Error 403:` and the response body (the only explanation the provider gives) was discarded. The body's `detail`/`error`/`message` is now carried into the `ProviderError`, redacted and truncated. This is what made the media 403s on `GET /backend-api/files/{id}/download` undiagnosable.
- **`redact_secrets` missed compound key names.** It matched keys exactly, so `access_token`, `api_key`, and `session-token` passed through un-redacted into debug-logged response bodies; matching now applies per word ("keywords", "monkey", "tokenizer" stay intact). - **`redact_secrets` missed compound key names.** It matched keys exactly, so `access_token`, `api_key`, and `session-token` passed through un-redacted into debug-logged response bodies; matching now applies per word ("keywords", "monkey", "tokenizer" stay intact).
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery. - **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
### Added ### Added
- **Codex CLI provider (`--provider codex`).** Archives local Codex agent transcripts from `~/.codex/sessions/**/rollout-*.jsonl` — local-only, like `claude-code`: no tokens, no rate limits, no ToS exposure. Sessions land in their own top-level `AI-Codex` Joplin notebook, with the same prose-only default, repo tags (`CODEX_REPO_TAG_IGNORE`) and multi-root scanning (`CODEX_DIR`, plus `$CODEX_HOME/sessions`).
Codex writes each session twice in one file and the choice between the two layers is the whole design. `response_item` records are the model-facing wire format, where a tool call arrives as *JavaScript* (`tools.exec_command({...})`) because Codex's `exec` tool is code-mode; `event_msg`/`item_completed` records are Codex's own typed items, already decoded into `CommandExecution`/`FileChange`/`Extension` with argv, cwd, exit code and output as fields. Measured over 7 sessions on 0.147.0 (2026-08-18), the typed layer is 1:1 with the raw layer for prose (91 `AgentMessage` ↔ 91 assistant messages, sharing ids) and additionally omits every piece of harness plumbing — all 51 `developer`-role messages plus the 7 `# AGENTS.md instructions…` and 1 `<environment_context>` injections — which the Claude Code provider has to strip by regex. So the typed layer is parsed for content.
Its one gap is that it records what *ran*, not what was *attempted*: 26 of 180 `exec_command` calls produced no item (14 sandbox launch failures, 6 user aborts, ~5 still running at turn end, 1 failure). The raw layer is therefore read for a call count only, and placeholders report the shortfall — `3 calls: exec_command ×3 (+2 did not complete)` — instead of silently under-reporting. `wait` calls are process polls, not attempts, and are excluded.
**Reasoning is not exportable from Codex.** All 345 reasoning records carry `encrypted_content`, with `summary` empty in the raw layer and `summary_text`/`raw_content` empty in the typed layer, in every session. It is always dropped and counted; unlike Claude Code, `EXPORTER_HIDDEN_CONTENT=full` cannot surface it. The sidecar SQLite databases (`state_5.sqlite`, `thread_history_1.sqlite`) are deliberately not read: `thread_history_projection_state` tracks a byte offset into the rollout file, so the JSONL is canonical and the DB derived, and its `title` column is just the first user message truncated.
**Codex Cloud is out of scope, verified rather than assumed.** Cloud tasks are reachable at `chatgpt.com/backend-api/api/codex/tasks{,/list}` — the same host and `/backend-api` root the ChatGPT provider already uses — but local CLI sessions are never uploaded there, so it is not an alternative source for these transcripts. `codex cloud list` confirmed the account holds no cloud tasks. The provider makes no network calls.
- **`projects` command — discover the project IDs your config is missing.** `CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is invisible to the listing pass, so its conversations are never fetched. The command reports every project your conversations belong to, marks which are absent from `.env`, and prints a paste-ready line (`--write` applies it). It reads project ids from the conversation listing when they are there and falls back to `--deep`, one detail request per conversation, when they are not. - **`projects` command — discover the project IDs your config is missing.** `CHATGPT_PROJECT_IDS` is maintained by hand, and a project missing from it is invisible to the listing pass, so its conversations are never fetched. The command reports every project your conversations belong to, marks which are absent from `.env`, and prints a paste-ready line (`--write` applies it). It reads project ids from the conversation listing when they are there and falls back to `--deep`, one detail request per conversation, when they are not.
- **Project attribution now reads the conversation's own `gizmo_id`.** Previously the project name came only from `CHATGPT_PROJECT_IDS`, so a conversation in a project you had not listed exported into `no-project/` even though its payload names its project. The detail response carries `gizmo_id`, so it is used as a fallback after the listing annotation and the project map — attribution stays correct without maintaining a list, and moving a chat into a new project no longer silently misfiles it. Only `g-p-` ids count: a custom GPT is not a project and must not become a folder. Each unconfigured project is reported once per run, naming the id to add, because the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that live only inside a project never appear in the default listing. - **Project attribution now reads the conversation's own `gizmo_id`.** Previously the project name came only from `CHATGPT_PROJECT_IDS`, so a conversation in a project you had not listed exported into `no-project/` even though its payload names its project. The detail response carries `gizmo_id`, so it is used as a fallback after the listing annotation and the project map — attribution stays correct without maintaining a list, and moving a chat into a new project no longer silently misfiles it. Only `g-p-` ids count: a custom GPT is not a project and must not become a folder. Each unconfigured project is reported once per run, naming the id to add, because the *listing* pass still needs `CHATGPT_PROJECT_IDS` — conversations that live only inside a project never appear in the default listing.
+22 -1
View File
@@ -263,6 +263,27 @@ Sessions from all roots are merged by folder (no per-machine label); if the same
--- ---
## Codex Sessions
The `codex` provider archives your local [Codex CLI](https://chatgpt.com/codex) agent transcripts — same deal as Claude Code: no tokens, no API, no ToS exposure. Rollout files are read from `~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl`.
```bash
ai-chat-exporter export --provider codex
ai-chat-exporter joplin --provider codex
```
Exports are prose-only by default, with the same placeholder format as Claude Code, into their own top-level **`AI-Codex`** Joplin notebook. Repo tags work the same way (`CODEX_REPO_TAG_IGNORE` to suppress), and additional roots can be scanned with `CODEX_DIR` (`:`-separated); `CODEX_HOME`'s `sessions/` is picked up automatically when that variable is set.
Three things differ from Claude Code, all forced by how Codex stores its data:
**Reasoning cannot be exported.** Codex encrypts it at rest — every reasoning record carries `encrypted_content` with no plaintext summary in any layer of the file. It is always dropped and counted; `EXPORTER_HIDDEN_CONTENT=full` cannot bring it back.
**Incomplete tool calls are reported.** Codex writes each session twice in one file: its own typed items (what actually ran) and the raw model-facing wire format (everything attempted). The provider reads the typed layer — it is already decoded, and it omits harness plumbing that would otherwise need stripping — but cross-checks the raw layer for calls that never produced a result, so placeholders read `3 calls: exec_command ×3 (+2 did not complete)`. Those are commands that failed to launch, that you aborted, or that were still running when the turn ended.
**Cloud tasks are out of scope.** `codex cloud` tasks run server-side and are reachable at `chatgpt.com/backend-api/api/codex/tasks`, but local CLI sessions are never uploaded there, so the cloud API is not an alternative source for these transcripts and this provider stays entirely offline. If you start using `codex cloud exec`, those transcripts would be cloud-only and would need separate work.
---
## Output Structure ## Output Structure
All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks. All exported files go under `EXPORT_DIR`. The folder structure maps directly to Joplin notebooks.
@@ -371,7 +392,7 @@ ai-chat-exporter export --output /path/to/my/notes
ai-chat-exporter export --dry-run ai-chat-exporter export --dry-run
``` ```
Options: `--provider [chatgpt|claude|claude-code|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--hidden-content [full|placeholder|omit]`, `--download-media [images|all|off]`, `--max-conversations N`, `--force`, `--dry-run` Options: `--provider [chatgpt|claude|claude-code|codex|all]`, `--format [markdown|json|both]`, `--output PATH`, `--since YYYY-MM-DD`, `--project NAME`, `--hidden-content [full|placeholder|omit]`, `--download-media [images|all|off]`, `--max-conversations N`, `--force`, `--dry-run`
**Re-rendering the whole archive after an upgrade.** New formatting or features (collapse policy, media downloads) only change conversations as they're re-exported. To re-render everything you already have, use `--force` — it re-exports every conversation even if unchanged, **without** `cache --clear`, so your Joplin note links are preserved (a later `joplin` run updates the existing notes instead of duplicating them). **Re-rendering the whole archive after an upgrade.** New formatting or features (collapse policy, media downloads) only change conversations as they're re-exported. To re-render everything you already have, use `--force` — it re-exports every conversation even if unchanged, **without** `cache --clear`, so your Joplin note links are preserved (a later `joplin` run updates the existing notes instead of duplicating them).
+4
View File
@@ -409,6 +409,10 @@ _PROVIDER_DISPLAY = {
# intermixed with Claude web projects and named by dev folder). Existing # intermixed with Claude web projects and named by dev folder). Existing
# notes self-heal into here on the next sync — update_note moves them. # notes self-heal into here on the next sync — update_note moves them.
"claude-code": "AI-ClaudeCode", "claude-code": "AI-ClaudeCode",
# Codex CLI coding sessions get their own top-level notebook for the same
# reason Claude Code does — they are a distinct surface, not a ChatGPT
# project, and burying them under AI-ChatGPT makes both harder to browse.
"codex": "AI-Codex",
} }
+20 -5
View File
@@ -684,7 +684,7 @@ def _print_doctor_table(checks: list[dict]) -> None:
@cli.command() @cli.command()
@click.option( @click.option(
"--provider", "--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False), type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all", default="all",
show_default=True, show_default=True,
help="Which provider to export.", help="Which provider to export.",
@@ -1076,6 +1076,21 @@ def _resolve_providers(provider: str, cfg) -> list[tuple[str, object]]:
", ".join(str(r) for r in cc_roots), ", ".join(str(r) for r in cc_roots),
) )
if provider in ("codex", "all"):
from src.providers.codex import CodexProvider, resolve_roots as codex_roots_fn
cx_roots = codex_roots_fn()
if any(r.is_dir() for r in cx_roots):
result.append((
"codex",
CodexProvider(hidden_content=cfg.hidden_content),
))
elif provider == "codex":
logging.getLogger(__name__).warning(
"[codex] Skipping — none of these roots found (set "
"CODEX_DIR, ':'-separated for multiple): %s",
", ".join(str(r) for r in cx_roots),
)
return result return result
@@ -1172,7 +1187,7 @@ def _print_export_summary(summary: dict[str, dict[str, int]]) -> None:
@cli.command(name="list") @cli.command(name="list")
@click.option( @click.option(
"--provider", "--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False), type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all", default="all",
show_default=True, show_default=True,
) )
@@ -1238,7 +1253,7 @@ def list_conversations(ctx: click.Context, provider: str, project_filter: str |
@click.option("--clear", is_flag=True, help="Clear cached entries.") @click.option("--clear", is_flag=True, help="Clear cached entries.")
@click.option( @click.option(
"--provider", "--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False), type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all", default="all",
help="Provider to target (used with --clear).", help="Provider to target (used with --clear).",
) )
@@ -1376,7 +1391,7 @@ def prune(ctx: click.Context, dry_run: bool, yes: bool) -> None:
@cli.command() @cli.command()
@click.option( @click.option(
"--provider", "--provider",
type=click.Choice(["chatgpt", "claude", "claude-code", "all"], case_sensitive=False), type=click.Choice(["chatgpt", "claude", "claude-code", "codex", "all"], case_sensitive=False),
default="all", default="all",
show_default=True, show_default=True,
help="Which provider's conversations to sync to Joplin.", help="Which provider's conversations to sync to Joplin.",
@@ -1449,7 +1464,7 @@ def joplin(ctx: click.Context, provider: str, project_filter: str | None, dry_ru
# Determine which providers to process # Determine which providers to process
providers_to_sync: list[str] = [] providers_to_sync: list[str] = []
for prov in ("chatgpt", "claude", "claude-code"): for prov in ("chatgpt", "claude", "claude-code", "codex"):
if provider in (prov, "all"): if provider in (prov, "all"):
providers_to_sync.append(prov) providers_to_sync.append(prov)
+7 -33
View File
@@ -54,6 +54,7 @@ from src.blocks import (
make_unknown_block, make_unknown_block,
) )
from src.loss_report import LossReport from src.loss_report import LossReport
from src.utils import git_root_name
from src.providers.base import ( from src.providers.base import (
BaseProvider, BaseProvider,
HIDDEN_CONTENT_FULL, HIDDEN_CONTENT_FULL,
@@ -395,7 +396,11 @@ def _parse_jsonl(path: Path) -> list[dict]:
except OSError as e: except OSError as e:
logger.warning("[claude-code] Could not read %s: %s", path, e) logger.warning("[claude-code] Could not read %s: %s", path, e)
return records return records
for line in text.splitlines(): # split("\n"), not splitlines(): splitlines() also breaks on U+0085,
# U+2028/9 and friends, which are legal *inside* a JSON string. A NEL in
# captured command output shreds one record into unparseable fragments and
# loses it silently (observed 2026-08-18 in a real Codex rollout).
for line in text.split("\n"):
line = line.strip() line = line.strip()
if not line: if not line:
continue continue
@@ -458,37 +463,6 @@ def _ignored_repos() -> set[str]:
return {s.strip() for s in env.split(",") if s.strip()} return {s.strip() for s in env.split(",") if s.strip()}
# dir Path → git-repo name it belongs to (or None). Process-wide; the working
# tree doesn't change under us mid-run, so caching walked dirs is safe.
_GIT_ROOT_CACHE: dict[Path, str | None] = {}
_CACHE_MISS = object()
def _git_root_name(path: Path, max_steps: int = 25) -> str | None:
"""Name of the git repo ``path`` lives in — nearest ancestor with ``.git``.
Walks up from ``path`` until a ``.git`` entry is found (returns that dir's
basename) or the filesystem root is reached (returns ``None``). Disk-based:
a path in no git repo, or a repo no longer on disk, yields ``None``.
"""
cur = path
for _ in range(max_steps):
cached = _GIT_ROOT_CACHE.get(cur, _CACHE_MISS)
if cached is not _CACHE_MISS:
return cached
try:
if (cur / ".git").exists():
_GIT_ROOT_CACHE[cur] = cur.name
return cur.name
except OSError:
break
if cur.parent == cur: # filesystem root
break
cur = cur.parent
_GIT_ROOT_CACHE[path] = None
return None
# Absolute path-like tokens inside Bash command strings (file_path/path keys are # Absolute path-like tokens inside Bash command strings (file_path/path keys are
# matched directly). git-root resolution short-circuits on non-repo paths. # matched directly). git-root resolution short-circuits on non-repo paths.
_ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+") _ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+")
@@ -521,7 +495,7 @@ def _repos_touched(
if p in seen: if p in seen:
return return
seen.add(p) seen.add(p)
name = _git_root_name(Path(p)) name = git_root_name(Path(p))
if name and not name.startswith(".") and name not in ignore: if name and not name.startswith(".") and name not in ignore:
counts[name] += 1 counts[name] += 1
+881
View File
@@ -0,0 +1,881 @@
"""Codex CLI session provider — archives local agent transcripts.
Reads JSONL rollout files from ``~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl``
(override with ``CODEX_DIR``). Like Claude Code: no tokens, no rate limits, no
ToS risk — the data is local and Codex may prune it.
Local-only, deliberately
------------------------
Codex Cloud tasks (``codex cloud``) live server-side at
``https://chatgpt.com/backend-api/api/codex/tasks{,/list}`` — the same host and
``/backend-api`` root the ChatGPT provider already speaks. **CLI sessions are
never uploaded there**, so the cloud API is not an alternative source for these
transcripts and this provider does not talk to the network. Verified 2026-08-18
against Codex 0.147.0. If ``codex cloud exec`` ever enters regular use, those
transcripts *would* be cloud-only and would need a separate provider.
Two representations, one file (measured 2026-08-18 over 7 sessions / 0.147.0)
----------------------------------------------------------------------------
Every rollout line is ``{timestamp, ordinal, type, payload}``. Dialogue appears
twice, in two different shapes, and we parse the **typed** one:
* ``response_item`` — the model-facing wire format (mirrors the OpenAI Responses
API). Tool calls arrive as *JavaScript source* because Codex's ``exec`` tool is
code-mode::
const r = await tools.exec_command({"cmd":"git status","workdir":"/x", …});
text(r.output);
Exactly one ``tools.*`` call per invocation; three functions observed:
``exec_command`` (180), ``web__run`` (15), ``apply_patch`` (13).
* ``event_msg`` / ``item_completed`` — Codex's own typed items, already decoded:
``UserMessage``, ``AgentMessage``, ``Reasoning``, ``CommandExecution``,
``FileChange``, ``Extension``, ``ContextCompaction``.
The typed layer wins on every axis that matters here. It is 1:1 with the raw
layer for prose (91 ``AgentMessage`` ↔ 91 assistant messages, same ids; 345
``Reasoning`` ↔ 345), it hands us structured command/exit-code/output fields
instead of JS we would have to regex, and it pre-filters harness plumbing for
free: all 51 ``developer``-role messages (skills manifests, ``<multi_agent_mode>``,
"Approved command prefix saved") plus the 7 ``# AGENTS.md instructions…``
injections and 1 ``<environment_context>`` have no typed item. That is the same
noise ``claude_code._HARNESS_TAG_RE`` strips by hand.
Its one weakness: it records what *ran*, not what was *attempted*. 26 of 180
``exec_command`` calls produced no ``CommandExecution`` item — 14 sandbox launch
failures (``bwrap: loopback: Failed RTM_NEWADDR``), 6 user aborts ("aborted by
user after 504.6s"), ~5 still running at turn end, 1 "Script failed". So we read
the raw layer *only* to count attempts, and the collapsed placeholder reports the
shortfall ("3 did not complete") rather than silently under-reporting. ``wait``
function calls (50) are process polls, not attempts, and are not counted.
Reasoning is unrecoverable
--------------------------
All 345 reasoning items carry ``encrypted_content``; ``summary`` is ``[]`` in the
raw layer and ``summary_text``/``raw_content`` are empty in the typed layer, in
every session. Codex does not persist readable reasoning locally. Thinking is
therefore always dropped and counted — the same end state as the Claude Code
policy (decision 2026-06-12), but by necessity rather than by choice, so even
``full`` cannot surface it.
Why the sidecar SQLite is not read
----------------------------------
``~/.codex/state_5.sqlite`` carries a ``threads`` table (title, cwd, model,
tokens_used, rollout_path) and ``thread_history_1.sqlite`` a projection of the
items — but ``thread_history_projection_state`` tracks a byte offset *into the
rollout file*, i.e. the JSONL is canonical and SQLite is derived. Its ``title``
is just the first user message truncated (identical to ``first_user_message`` and
``preview``), so it offers nothing the JSONL lacks, and its filename carries a
schema version that will churn. We read the files.
Rendering follows the EXPORTER_HIDDEN_CONTENT policy: prose-only by default
(dialogue kept, tool traffic collapsed to one grouped placeholder per activity
run, reasoning dropped); ``full`` keeps the decoded tool calls and their output.
Subagents: Codex 0.147.0's ``thread_spawn_edges`` table exists but is empty and no
sub-transcripts were observed, so there is no subagent folding here (contrast
``claude_code._load_subagents``). If spawned agents start appearing they will
arrive as new item types and land in the loss report as unknowns.
"""
import json
import logging
import os
import re
from collections import Counter
from datetime import datetime, timezone
from pathlib import Path
from src.blocks import (
COLLAPSED_KIND_HIDDEN_CONTEXT,
COLLAPSED_KIND_TOOL_DUMP,
UNKNOWN_REASON_UNKNOWN_TYPE,
make_collapsed_block,
make_text_block,
make_tool_result_block,
make_tool_use_block,
make_unknown_block,
)
from src.loss_report import LossReport
from src.providers.base import (
BaseProvider,
HIDDEN_CONTENT_FULL,
ProviderError,
VALID_HIDDEN_CONTENT_POLICIES,
resolve_hidden_content_policy,
)
from src.utils import git_root_name
logger = logging.getLogger(__name__)
DEFAULT_SESSIONS_DIR = "~/.codex/sessions"
# rollout-<ISO-ish timestamp>-<uuid>.jsonl — the trailing UUID is the thread id.
_ROLLOUT_RE = re.compile(
r"^rollout-\d{4}-\d{2}-\d{2}T[\d-]+-"
r"([0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12})$"
)
# The single tools.* call inside an `exec` custom_tool_call's JavaScript body.
_TOOLS_CALL_RE = re.compile(r"tools\.([A-Za-z_][\w.]*)\s*\(")
# Typed items that represent tool traffic, mapped to the label used in the
# collapsed placeholder. These labels must match the ones derived from the raw
# layer in _attempt_label, or the attempted-vs-completed delta is meaningless.
_TOOL_ITEM_LABELS = {
"CommandExecution": "exec_command",
"FileChange": "apply_patch",
# Extension is labelled by its `kind` (e.g. "web.search"); see _tool_label.
}
# function_call names that poll an already-running process rather than starting
# new work. Counting them as attempts would inflate the shortfall.
_POLLING_FUNCTIONS = {"wait"}
# How many distinct tool names to list in a collapsed-activity placeholder.
_TOOL_NAMES_SHOWN = 4
# Harness-injected user text. The typed layer already omits these (they have no
# UserMessage item), so this is a belt-and-braces guard for other Codex versions.
_HARNESS_USER_RE = re.compile(
r"^\s*(?:#\s*AGENTS\.md instructions for\b|<environment_context>)",
)
def resolve_roots(sessions_dir=None) -> list[Path]:
"""Resolve the ordered list of Codex ``sessions/`` roots to scan.
Precedence:
1. Explicit ``sessions_dir`` (a single path or a list) — used by tests.
2. ``CODEX_DIR`` env, split on ``os.pathsep`` (``:``) so multiple roots can
be given; a single path (no separator) stays backward compatible. Falls
back to the default root when unset.
3. Additionally, if ``CODEX_HOME`` is set (Codex's own name for its state
directory), its ``sessions/`` subdir is appended.
Roots are expanded and de-duplicated (order preserved); existence is checked
by the caller / ``_scan``.
"""
raw: list[str]
if sessions_dir is not None:
raw = (
[str(p) for p in sessions_dir]
if isinstance(sessions_dir, (list, tuple))
else [str(sessions_dir)]
)
else:
env = os.getenv("CODEX_DIR")
raw = env.split(os.pathsep) if env else [DEFAULT_SESSIONS_DIR]
codex_home = os.getenv("CODEX_HOME")
if codex_home:
raw.append(str(Path(codex_home) / "sessions"))
roots: list[Path] = []
for p in raw:
p = p.strip()
if not p:
continue
path = Path(p).expanduser()
if path not in roots:
roots.append(path)
return roots
class CodexProvider(BaseProvider):
"""Local-file provider over Codex CLI rollout transcripts."""
provider_name = "codex"
def __init__(
self,
sessions_dir: str | Path | None = None,
hidden_content: str | None = None,
) -> None:
super().__init__()
self._sessions_dirs = resolve_roots(sessions_dir)
self._hidden_content = (
hidden_content
if hidden_content in VALID_HIDDEN_CONTENT_POLICIES
else resolve_hidden_content_policy()
)
# conv_id → rollout file path, populated by _scan()
self._path_map: dict[str, Path] = {}
# ------------------------------------------------------------------
# BaseProvider interface
# ------------------------------------------------------------------
def list_conversations(self, offset: int = 0, limit: int = 100) -> list[dict]:
full = self._scan()
return full[offset : offset + limit]
def fetch_all_conversations(self, since: datetime | None = None) -> list[dict]:
convs = self._scan()
if since is not None:
since_aware = since if since.tzinfo else since.replace(tzinfo=timezone.utc)
convs = [
c for c in convs
if datetime.fromisoformat(c["updated_at"]) >= since_aware
]
logger.info(
"[codex] Found %d session(s) under %s",
len(convs),
", ".join(str(d) for d in self._sessions_dirs),
)
return convs
def get_conversation(self, conv_id: str) -> dict:
path = self._path_map.get(conv_id)
if path is None:
# Direct call without a prior listing (e.g. tests) — scan first.
self._scan()
path = self._path_map.get(conv_id)
if path is None or not path.exists():
raise ProviderError(
self.provider_name,
f"get_conversation({conv_id[:8]})",
FileNotFoundError(f"No rollout file for id {conv_id}"),
)
records = _parse_jsonl(path)
mtime = datetime.fromtimestamp(path.stat().st_mtime, tz=timezone.utc)
return {
"id": conv_id,
"_path": str(path),
"_records": records,
# Listing and normalized updated_at must match, or the cache
# staleness comparison would re-export every session every run.
"_mtime_iso": mtime.isoformat(),
}
def normalize_conversation(self, raw: dict, loss_report: LossReport | None = None) -> dict:
report = loss_report if loss_report is not None else LossReport()
policy = getattr(self, "_hidden_content", None) or resolve_hidden_content_policy()
conv_id = raw.get("id") or ""
records: list[dict] = raw.get("_records") or []
title = _extract_title(records)
# Append the repos this session touched, e.g. "… [repo-a, repo-b]".
# Codex sessions are commonly all launched from one workspace root, so
# without this every session lands in the same notebook with no way to
# tell them apart.
launch_cwd = _extract_launch_cwd(records)
if launch_cwd:
repos = _repos_touched(records, launch_cwd)
if repos:
title = f"{title} [{', '.join(repos)}]"
project = _extract_project(records)
created_at = _extract_created_at(records)
updated_at = raw.get("_mtime_iso") or next(
(r.get("timestamp") for r in reversed(records) if r.get("timestamp")), ""
)
messages = _extract_messages(records, conv_id, report, policy)
for _ in messages:
report.record_message()
report.record_conversation()
return {
"id": conv_id,
"title": title,
"provider": self.provider_name,
"project": project,
"created_at": created_at or "",
"updated_at": updated_at or "",
"message_count": len(messages),
"messages": messages,
}
# ------------------------------------------------------------------
# Scanning
# ------------------------------------------------------------------
def _scan(self) -> list[dict]:
existing = [d for d in self._sessions_dirs if d.is_dir()]
if not existing:
logger.warning(
"[codex] No sessions directory exists: %s",
", ".join(str(d) for d in self._sessions_dirs),
)
return []
# conv_id → (mtime, conv dict). If the same thread id appears under two
# roots (e.g. a live dir and a backup), the newer-mtime copy wins.
by_id: dict[str, tuple[float, dict]] = {}
for root in existing:
# Rollouts are filed under YYYY/MM/DD; rglob keeps us agnostic to
# that layout in case Codex reorganises it.
for session_file in sorted(root.rglob("rollout-*.jsonl")):
match = _ROLLOUT_RE.match(session_file.stem)
if not match:
logger.debug("[codex] Skipping unrecognised filename %s", session_file.name)
continue
try:
stat = session_file.stat()
except OSError:
continue
if stat.st_size == 0:
continue
conv_id = match.group(1)
prev = by_id.get(conv_id)
if prev is not None and prev[0] >= stat.st_mtime:
logger.debug(
"[codex] Duplicate session %s in %s; keeping newer copy",
conv_id[:8], root,
)
continue
self._path_map[conv_id] = session_file
title, project, created = _read_session_meta(session_file)
by_id[conv_id] = (
stat.st_mtime,
{
"id": conv_id,
"title": title,
"project": project,
# The --project filter and dry-run table read the
# listing dict, not the normalized conversation.
"_project_name": project,
"created_at": created,
"updated_at": datetime.fromtimestamp(
stat.st_mtime, tz=timezone.utc
).isoformat(),
"_path": str(session_file),
},
)
# Deterministic order: by creation date bucket, then thread id.
return sorted(
(conv for _, conv in by_id.values()),
key=lambda c: (c["created_at"], c["id"]),
)
# ---------------------------------------------------------------------------
# Internal helpers
# ---------------------------------------------------------------------------
def _parse_jsonl(path: Path) -> list[dict]:
"""Read a JSONL file into records, tolerating (and logging) bad lines."""
records: list[dict] = []
bad_lines = 0
try:
text = path.read_text(encoding="utf-8")
except OSError as e:
logger.warning("[codex] Could not read %s: %s", path, e)
return records
# split("\n"), not splitlines(): splitlines() also breaks on U+0085,
# U+2028/9 and friends, which are legal *inside* a JSON string. A NEL in
# captured command output shreds one record into unparseable fragments and
# loses it silently (observed 2026-08-18 in a real Codex rollout).
for line in text.split("\n"):
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError:
bad_lines += 1
if bad_lines:
logger.warning("[codex] %s: skipped %d unparseable line(s)", path.name, bad_lines)
return records
def _item(rec: dict) -> dict | None:
"""Return the typed item from an ``event_msg``/``item_completed`` record."""
if rec.get("type") != "event_msg":
return None
payload = rec.get("payload") or {}
if payload.get("type") != "item_completed":
return None
item = payload.get("item")
return item if isinstance(item, dict) else None
def _item_kind(item: dict) -> str:
"""Typed item discriminator. 0.147.0 uses ``type``; ``item_type`` is a hedge."""
return str(item.get("item_type") or item.get("type") or "")
def _item_text(item: dict) -> str:
"""Concatenate the text of a UserMessage / AgentMessage item.
The two disagree on case — ``UserMessage`` blocks are ``"text"`` and
``AgentMessage`` blocks are ``"Text"`` — so the comparison is case-folded.
"""
content = item.get("content")
if isinstance(content, str):
return content.strip()
if not isinstance(content, list):
return ""
parts = []
for block in content:
if isinstance(block, dict) and str(block.get("type", "")).lower() == "text":
text = block.get("text")
if isinstance(text, str):
parts.append(text)
return "".join(parts).strip()
def _read_session_meta(path: Path) -> tuple[str, str | None, str]:
"""Light single-pass scan for listing metadata: (title, project, created_at).
Substring guards keep this cheap — only candidate lines are JSON-parsed, and
the scan stops as soon as the title is found (the first UserMessage is
usually within the first few dozen lines of a multi-megabyte file).
"""
title = ""
project: str | None = None
created = ""
try:
with path.open(encoding="utf-8") as fh:
for line in fh:
if not created and '"session_meta"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
payload = rec.get("payload") or {}
created = str(payload.get("timestamp") or rec.get("timestamp") or "")
cwd = payload.get("cwd")
if cwd:
project = Path(cwd).name or None
continue
if not title and '"UserMessage"' in line:
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
item = _item(rec)
if item is None or _item_kind(item) != "UserMessage":
continue
text = _item_text(item)
if text and not _HARNESS_USER_RE.match(text):
title = text[:80]
break
except OSError as e:
logger.warning("[codex] Could not read %s: %s", path, e)
return title or "Untitled session", project, created
def _extract_title(records: list[dict]) -> str:
"""Codex has no AI-generated title — the first real user prompt is the title.
(``state_5.sqlite``'s ``title`` column is this same string truncated; see the
module docstring on why the sidecar DB is not consulted.)
"""
for rec in records:
item = _item(rec)
if item is None or _item_kind(item) != "UserMessage":
continue
text = _item_text(item)
if text and not _HARNESS_USER_RE.match(text):
return text[:80]
return "Untitled session"
def _session_meta_payload(records: list[dict]) -> dict:
for rec in records:
if rec.get("type") == "session_meta":
payload = rec.get("payload")
if isinstance(payload, dict):
return payload
return {}
def _extract_launch_cwd(records: list[dict]) -> str | None:
"""The session's working directory (constant per session)."""
cwd = _session_meta_payload(records).get("cwd")
return cwd if isinstance(cwd, str) and cwd else None
def _extract_project(records: list[dict]) -> str | None:
"""Project = basename of the session's working directory."""
cwd = _extract_launch_cwd(records)
if cwd:
return Path(cwd).name or None
return None
def _extract_created_at(records: list[dict]) -> str:
"""Session start time.
Prefers ``session_meta.payload.timestamp`` — the outer line ``timestamp`` is
when the record was *flushed*, which can trail the true start by minutes.
"""
payload = _session_meta_payload(records)
ts = payload.get("timestamp")
if isinstance(ts, str) and ts:
return ts
return next((r.get("timestamp") for r in records if r.get("timestamp")), "") or ""
def _ignored_repos() -> set[str]:
"""Optional ignore-list: git repos to never tag (comma-separated names)."""
env = os.getenv("CODEX_REPO_TAG_IGNORE", "")
return {s.strip() for s in env.split(",") if s.strip()}
def _strip_file_uri(path: str) -> str:
"""``CommandExecution.cwd`` is a ``file://`` URI; every other path is plain."""
return path[7:] if path.startswith("file://") else path
# Absolute path-like tokens inside command strings.
_ABS_PATH_RE = re.compile(r"/(?:[\w.\-]+/)*[\w.\-]+")
def _repos_touched(
records: list[dict], launch_cwd: str, cap: int = 3, ignore: set[str] | None = None
) -> list[str]:
"""Repos a session touched, for the title's ``[repo-a, repo-b]`` tag.
A file's repo is the git repository it lives in (nearest ancestor with a
``.git``). Paths come from the typed items: ``FileChange.changes`` keys
(absolute), and ``CommandExecution``'s argv plus its ``cwd``. Ordered by
touch frequency, capped with a trailing ``…``.
"""
ignore = _ignored_repos() if ignore is None else ignore
counts: Counter = Counter()
seen: set[str] = set()
def note(p) -> None:
if not isinstance(p, str) or not p:
return
p = _strip_file_uri(p)
if not p.startswith("/"): # resolve relative paths against the launch cwd
if not launch_cwd:
return
p = str(Path(launch_cwd) / p)
if p in seen:
return
seen.add(p)
name = git_root_name(Path(p))
if name and not name.startswith(".") and name not in ignore:
counts[name] += 1
for rec in records:
item = _item(rec)
if item is None:
continue
kind = _item_kind(item)
if kind == "FileChange":
changes = item.get("changes")
if isinstance(changes, dict):
for file_path in changes:
note(file_path)
elif kind == "CommandExecution":
note(item.get("cwd"))
command = item.get("command")
if isinstance(command, list) and command:
tail = command[-1]
if isinstance(tail, str):
for m in _ABS_PATH_RE.finditer(tail):
note(m.group(0))
ordered = [name for name, _ in counts.most_common()]
if len(ordered) > cap:
return ordered[:cap] + ["…"]
return ordered
def _tool_label(item: dict, kind: str) -> str:
"""Placeholder label for a completed tool item.
``Extension`` covers everything routed through the model's own extensions
(``web.search`` so far), so its ``kind`` field is the useful name.
"""
if kind == "Extension":
return str(item.get("kind") or "extension")
return _TOOL_ITEM_LABELS.get(kind, kind)
def _attempt_label(payload: dict) -> str | None:
"""Label for a raw tool call, or None if it should not count as an attempt.
``exec`` custom_tool_calls carry JavaScript; the inner ``tools.<fn>`` name is
what lines up with the typed items' labels. ``wait`` polls an already-running
process and starts no new work.
"""
ptype = payload.get("type")
name = payload.get("name") or ""
if ptype == "custom_tool_call":
if name == "exec":
match = _TOOLS_CALL_RE.search(payload.get("input") or "")
if match:
# tools.web__run → the Extension item calls itself "web.search";
# both are "one web call", which is all the count claims.
return match.group(1)
return "exec"
return name or "tool"
if ptype == "function_call":
if name in _POLLING_FUNCTIONS:
return None
return name or "function"
return None
def _command_string(item: dict) -> str:
"""Human-readable command from a ``CommandExecution`` argv list.
The argv is ``["/bin/bash", "-lc", "<script>"]``; the script is the part
worth showing.
"""
command = item.get("command")
if isinstance(command, list):
if len(command) >= 3 and command[0].endswith("sh") and command[1] in ("-lc", "-c"):
return str(command[-1])
return " ".join(str(c) for c in command)
return str(command or "")
def _extract_messages(
records: list[dict],
conv_id: str,
report: LossReport,
policy: str,
) -> list[dict]:
"""Normalize Codex rollout records into messages.
Reads the typed ``item_completed`` layer for content and the raw
``response_item`` layer only to count attempted tool calls, so a collapsed
placeholder can report calls that never produced an item (sandbox failures,
user aborts, still-running processes). See the module docstring.
"""
messages: list[dict] = []
# Pending collapsed tool activity for the current run.
pending_tools: Counter = Counter()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
def flush_pending() -> None:
nonlocal pending_bytes, pending_attempted, pending_completed
# An activity run with only failed attempts still deserves a placeholder
# — silence would imply nothing happened.
shortfall = max(pending_attempted - pending_completed, 0)
if not pending_tools and not shortfall:
return
if policy == HIDDEN_CONTENT_FULL:
# Under `full` the completed calls are already rendered as tool
# blocks, so only the shortfall is left to report — and a *collapsed*
# block would render the "omitted (set full to keep)" suffix, which
# contradicts the policy in force. Say it plainly instead.
report.record_collapsed("tool_call_incomplete", 0)
note = make_tool_result_block(
f"{shortfall} tool call(s) produced no result — the process failed to "
"launch, was aborted, or was still running when the turn ended.",
tool_name="incomplete",
is_error=True,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(note)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None,
"blocks": [note]}
)
pending_tools.clear()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
return
if pending_tools:
shown = ", ".join(
f"{name} ×{count}"
for name, count in pending_tools.most_common(_TOOL_NAMES_SHOWN)
)
if len(pending_tools) > _TOOL_NAMES_SHOWN:
shown += ", …"
origin = f"{sum(pending_tools.values())} calls: {shown}"
else:
origin = "0 calls"
if shortfall:
origin += f" (+{shortfall} did not complete)"
report.record_collapsed("tool_call_incomplete", 0)
block = make_collapsed_block(
origin=origin,
content_type="tool_activity",
size_bytes=pending_bytes,
kind=COLLAPSED_KIND_TOOL_DUMP,
)
if messages and messages[-1]["role"] == "assistant":
messages[-1]["blocks"].append(block)
else:
messages.append(
{"role": "tool", "content_type": "text", "timestamp": None, "blocks": [block]}
)
pending_tools.clear()
pending_bytes = 0
pending_attempted = 0
pending_completed = 0
def append_message(role: str, blocks: list[dict], timestamp) -> None:
# Attach the preceding run's tool activity to the previous message
# before starting a new one, so the placeholder lands between the
# dialogue turns it actually occurred between.
flush_pending()
messages.append(
{
"role": role,
"content_type": "text",
"timestamp": timestamp,
"blocks": blocks,
}
)
for rec in records:
rec_type = rec.get("type")
payload = rec.get("payload") or {}
timestamp = rec.get("timestamp")
# ── Raw layer: attempt counting only ──────────────────────────────
if rec_type == "response_item":
if payload.get("type") in ("custom_tool_call", "function_call"):
if _attempt_label(payload) is not None:
pending_attempted += 1
continue
item = _item(rec)
if item is None:
continue
kind = _item_kind(item)
# ── Dialogue ──────────────────────────────────────────────────────
if kind in ("UserMessage", "AgentMessage"):
text = _item_text(item)
if kind == "UserMessage" and _HARNESS_USER_RE.match(text):
# Defensive: 0.147.0 gives these no typed item at all.
report.record_filtered_role("codex.harness_injection")
continue
block = make_text_block(text)
if block:
append_message(
"user" if kind == "UserMessage" else "assistant", [block], timestamp
)
continue
# ── Reasoning: encrypted at rest, nothing to keep under any policy ─
if kind == "Reasoning":
report.record_collapsed("reasoning", len(json.dumps(item, default=str)))
continue
# ── Context compaction: a visible marker, not a silent drop ────────
if kind == "ContextCompaction":
flush_pending()
marker = make_collapsed_block(
origin="context compacted — earlier turns dropped from the model's context",
content_type="context_compaction",
size_bytes=0,
kind=COLLAPSED_KIND_HIDDEN_CONTEXT,
)
report.record_collapsed("context_compaction", 0)
messages.append(
{"role": "tool", "content_type": "text", "timestamp": timestamp,
"blocks": [marker]}
)
continue
# ── Tool traffic ──────────────────────────────────────────────────
if kind in ("CommandExecution", "FileChange", "Extension"):
pending_completed += 1
label = _tool_label(item, kind)
size = len(json.dumps(item, default=str))
if policy == HIDDEN_CONTENT_FULL:
use_block, result_block = _full_tool_blocks(item, kind, label)
target = messages[-1] if messages and messages[-1]["role"] == "assistant" else None
if target is None:
messages.append(
{"role": "tool", "content_type": "text",
"timestamp": timestamp, "blocks": []}
)
target = messages[-1]
target["blocks"].append(use_block)
if result_block:
target["blocks"].append(result_block)
else:
pending_tools[label] += 1
pending_bytes += size
report.record_collapsed(label, size)
continue
# ── Anything new in a future Codex version ────────────────────────
logger.warning("[codex] Unknown item type %r in session %s", kind, conv_id[:8])
report.record_unknown(f"codex.{kind or '?'}")
append_message(
"tool",
[
make_unknown_block(
raw_type=f"codex.{kind or '?'}",
observed_keys=list(item.keys()),
reason=UNKNOWN_REASON_UNKNOWN_TYPE,
)
],
timestamp,
)
flush_pending()
return messages
def _full_tool_blocks(item: dict, kind: str, label: str) -> tuple[dict, dict | None]:
"""Decoded tool_use / tool_result blocks for EXPORTER_HIDDEN_CONTENT=full.
The typed item already carries structured fields, so this reads the command,
exit code and output directly rather than parsing the raw layer's JavaScript.
"""
if kind == "CommandExecution":
exit_code = item.get("exit_code")
use = make_tool_use_block(
label,
{
"command": _command_string(item),
"cwd": _strip_file_uri(str(item.get("cwd") or "")),
"status": item.get("status"),
"exit_code": exit_code,
},
item.get("id"),
)
output = item.get("aggregated_output") or item.get("formatted_output") or ""
result = make_tool_result_block(
str(output),
tool_name=label,
is_error=bool(exit_code not in (0, None)),
)
return use, result
if kind == "FileChange":
changes = item.get("changes") if isinstance(item.get("changes"), dict) else {}
use = make_tool_use_block(
label,
{
"files": {
path: (change.get("type") if isinstance(change, dict) else "?")
for path, change in changes.items()
},
"status": item.get("status"),
},
item.get("id"),
)
output = "\n".join(
part for part in (item.get("stdout"), item.get("stderr")) if part
)
result = make_tool_result_block(
output or f"{len(changes)} file(s) changed",
tool_name=label,
is_error=item.get("status") not in (None, "completed", "success"),
)
return use, result
# Extension (web.search and anything else routed through an extension).
use = make_tool_use_block(
label,
{"query": item.get("query"), "action": item.get("action")},
item.get("id"),
)
results = item.get("results")
result = make_tool_result_block(
json.dumps(results, default=str, indent=1) if results else "",
tool_name=label,
)
return use, result
+39
View File
@@ -171,3 +171,42 @@ def _parse_dt(ts: str) -> datetime:
if dt.tzinfo is None: if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc) dt = dt.replace(tzinfo=timezone.utc)
return dt return dt
# ---------------------------------------------------------------------------
# Git-repo resolution (shared by the local agent-transcript providers)
# ---------------------------------------------------------------------------
# dir Path → git-repo name it belongs to (or None). Process-wide; the working
# tree doesn't change under us mid-run, so caching walked dirs is safe.
_GIT_ROOT_CACHE: dict[Path, str | None] = {}
_CACHE_MISS = object()
def git_root_name(path: Path, max_steps: int = 25) -> str | None:
"""Name of the git repo ``path`` lives in — nearest ancestor with ``.git``.
Walks up from ``path`` until a ``.git`` entry is found (returns that dir's
basename) or the filesystem root is reached (returns ``None``). Disk-based:
a path in no git repo, or a repo no longer on disk, yields ``None``.
Used by the Claude Code and Codex providers to tag a session title with the
repos it touched, so sessions launched from a shared workspace root stay
distinguishable.
"""
cur = path
for _ in range(max_steps):
cached = _GIT_ROOT_CACHE.get(cur, _CACHE_MISS)
if cached is not _CACHE_MISS:
return cached
try:
if (cur / ".git").exists():
_GIT_ROOT_CACHE[cur] = cur.name
return cur.name
except OSError:
break
if cur.parent == cur: # filesystem root
break
cur = cur.parent
_GIT_ROOT_CACHE[path] = None
return None
+31 -1
View File
@@ -20,7 +20,11 @@ def _write_session(tmp_path, records, project_dir="-home-jesse-myproj", name="ab
proj = tmp_path / project_dir proj = tmp_path / project_dir
proj.mkdir(parents=True, exist_ok=True) proj.mkdir(parents=True, exist_ok=True)
f = proj / f"{name}.jsonl" f = proj / f"{name}.jsonl"
f.write_text("\n".join(json.dumps(r) for r in records), encoding="utf-8") # ensure_ascii=False so raw U+0085/U+2028 reach the parser — see
# TestExoticLineBreaks.
f.write_text(
"\n".join(json.dumps(r, ensure_ascii=False) for r in records), encoding="utf-8"
)
return f return f
@@ -501,3 +505,29 @@ class TestMultiRoot:
monkeypatch.setenv("CLAUDE_CONFIG_DIR", str(tmp_path / "cfg")) monkeypatch.setenv("CLAUDE_CONFIG_DIR", str(tmp_path / "cfg"))
roots = resolve_roots() roots = resolve_roots()
assert (tmp_path / "cfg" / "projects") in roots assert (tmp_path / "cfg" / "projects") in roots
class TestExoticLineBreaks:
"""Same U+0085 hazard as the Codex provider — see its TestExoticLineBreaks."""
# C0 controls (\x0b, \x0c) are excluded: JSON requires them escaped, so they
# never reach the splitter literally. These three do.
@pytest.mark.parametrize("sep", ["\x85", "
", "
"])
def test_record_with_exotic_break_survives(self, tmp_path, sep):
records = [
{
"type": "user",
"cwd": "/home/jesse/myproj",
"timestamp": "2026-05-01T10:00:00.000Z",
"message": {"role": "user", "content": f"before{sep}after"},
},
]
_write_session(tmp_path, records)
prov = ClaudeCodeProvider(projects_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation("abc-123"))
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert f"before{sep}after" in texts
+410
View File
@@ -0,0 +1,410 @@
"""Unit tests for the Codex CLI session provider.
Fixtures mirror the real 0.147.0 rollout shape observed on 2026-08-18: dialogue
carried twice (typed ``item_completed`` items plus raw ``response_item``s),
code-mode ``exec`` calls whose input is JavaScript, encrypted reasoning, and
harness-injected user messages that exist only in the raw layer.
"""
import json
import pytest
from src.blocks import (
BLOCK_TYPE_COLLAPSED,
BLOCK_TYPE_TEXT,
BLOCK_TYPE_TOOL_RESULT,
BLOCK_TYPE_TOOL_USE,
COLLAPSED_KIND_HIDDEN_CONTEXT,
)
from src.loss_report import LossReport
from src.providers.codex import CodexProvider, resolve_roots
SESSION_ID = "01a00e3f-a309-74a3-bf32-06c2cd87faa3"
FILENAME = f"rollout-2026-08-17T01-44-06-{SESSION_ID}.jsonl"
def _write_session(tmp_path, records, name=FILENAME, day="2026/08/17"):
day_dir = tmp_path / day
day_dir.mkdir(parents=True, exist_ok=True)
f = day_dir / name
# ensure_ascii=False: Codex (Rust) writes non-ASCII literally, so real
# rollouts contain raw U+0085/U+2028 inside JSON strings. Escaping them here
# would hide exactly the hazard TestExoticLineBreaks exists to catch.
f.write_text(
"\n".join(json.dumps(r, ensure_ascii=False) for r in records), encoding="utf-8"
)
return f
def _item(item_type, ts="2026-08-17T05:44:10.000Z", **fields):
# `item_type`, not `kind` — Extension items carry their own `kind` field.
return {
"timestamp": ts,
"type": "event_msg",
"payload": {"type": "item_completed", "item": {"type": item_type, **fields}},
}
def _exec_call(cmd, fn="exec_command"):
"""A raw code-mode custom_tool_call — its input is JavaScript, not JSON."""
return {
"timestamp": "2026-08-17T05:44:11.000Z",
"type": "response_item",
"payload": {
"type": "custom_tool_call",
"name": "exec",
"call_id": "call_1",
"input": (
f'const r = await tools.{fn}({{"cmd":{json.dumps(cmd)},'
f'"workdir":"/home/jesse/ws","yield_time_ms":30000}});\ntext(r.output);'
),
},
}
def _records(cwd="/home/jesse/ws"):
"""A representative session: meta, harness noise, dialogue, tool traffic."""
return [
{
"timestamp": "2026-08-17T05:49:13.629Z",
"ordinal": 0,
"type": "session_meta",
"payload": {
"session_id": SESSION_ID,
"timestamp": "2026-08-17T05:44:06.666Z",
"cwd": cwd,
"cli_version": "0.147.0",
"source": "cli",
},
},
# Harness plumbing: present in the raw layer only, exactly as 0.147.0
# writes it. Must not become a message.
{
"timestamp": "2026-08-17T05:44:07.000Z",
"type": "response_item",
"payload": {
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "# AGENTS.md instructions for /x"}],
},
},
{
"timestamp": "2026-08-17T05:44:08.000Z",
"type": "response_item",
"payload": {
"type": "message",
"role": "developer",
"content": [{"type": "input_text", "text": "<skills_instructions>…"}],
},
},
_item(
"UserMessage",
ts="2026-08-17T05:44:09.000Z",
id="u1",
content=[{"type": "text", "text": "Write the backup guide.", "text_elements": []}],
),
# Encrypted reasoning — nothing recoverable in either layer.
{
"timestamp": "2026-08-17T05:44:09.500Z",
"type": "response_item",
"payload": {"type": "reasoning", "summary": [], "encrypted_content": "gAAAA…"},
},
_item("Reasoning", id="r1", summary_text=[], raw_content=[]),
# AgentMessage content blocks are "Text" (capital T), unlike UserMessage.
_item(
"AgentMessage",
id="a1",
phase="commentary",
content=[{"type": "Text", "text": "I'll inspect the repo first."}],
),
_exec_call("ls -la"),
_item(
"CommandExecution",
id="exec-1",
process_id="123",
command=["/bin/bash", "-lc", "ls -la"],
cwd="file:///home/jesse/ws",
source="unified_exec_startup",
status="completed",
exit_code=0,
stdout="total 4\n",
aggregated_output="total 4\n",
formatted_output="total 4\n",
),
_exec_call("cat missing"),
_item(
"CommandExecution",
id="exec-2",
command=["/bin/bash", "-lc", "cat missing"],
cwd="file:///home/jesse/ws",
status="completed",
exit_code=1,
stderr="No such file\n",
aggregated_output="No such file\n",
),
# An attempt that never produced an item (sandbox failure / abort).
_exec_call("npm test"),
# A `wait` poll: not an attempt, must not inflate the shortfall.
{
"timestamp": "2026-08-17T05:44:12.000Z",
"type": "response_item",
"payload": {"type": "function_call", "name": "wait", "call_id": "call_w"},
},
_item(
"AgentMessage",
ts="2026-08-17T05:45:00.000Z",
id="a2",
phase="final_answer",
content=[{"type": "Text", "text": "Done — the guide is written."}],
),
]
class TestCodexProvider:
def test_scan_lists_session_with_metadata(self, tmp_path):
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
convs = prov.list_conversations()
assert len(convs) == 1
assert convs[0]["id"] == SESSION_ID
assert convs[0]["title"] == "Write the backup guide."
assert convs[0]["project"] == "ws"
# session_meta payload timestamp, not the (later) flush timestamp.
assert convs[0]["created_at"] == "2026-08-17T05:44:06.666Z"
def test_ignores_non_rollout_files(self, tmp_path):
_write_session(tmp_path, _records())
(tmp_path / "2026/08/17/notes.jsonl").write_text("{}", encoding="utf-8")
(tmp_path / "2026/08/17/rollout-garbage.jsonl").write_text("{}", encoding="utf-8")
assert len(CodexProvider(sessions_dir=tmp_path).list_conversations()) == 1
def test_empty_file_skipped(self, tmp_path):
_write_session(tmp_path, _records())
(tmp_path / "2026/08/16").mkdir(parents=True)
(tmp_path / "2026/08/16" / FILENAME.replace("17T01", "16T01")).write_text("")
assert len(CodexProvider(sessions_dir=tmp_path).list_conversations()) == 1
def test_missing_root_returns_empty(self, tmp_path):
prov = CodexProvider(sessions_dir=tmp_path / "nope")
assert prov.list_conversations() == []
def test_get_conversation_unknown_id_raises(self, tmp_path):
from src.providers.base import ProviderError
prov = CodexProvider(sessions_dir=tmp_path)
with pytest.raises(ProviderError):
prov.get_conversation("does-not-exist")
class TestNormalize:
def _normalized(self, tmp_path, policy="placeholder", records=None):
_write_session(tmp_path, records if records is not None else _records())
prov = CodexProvider(sessions_dir=tmp_path, hidden_content=policy)
prov.list_conversations()
report = LossReport()
return prov.normalize_conversation(prov.get_conversation(SESSION_ID), report), report
def test_dialogue_only_by_default(self, tmp_path):
conv, _ = self._normalized(tmp_path)
roles = [m["role"] for m in conv["messages"]]
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert roles[0] == "user"
assert texts == [
"Write the backup guide.",
"I'll inspect the repo first.",
"Done — the guide is written.",
]
def test_harness_injections_never_become_messages(self, tmp_path):
conv, _ = self._normalized(tmp_path)
blob = json.dumps(conv)
assert "AGENTS.md instructions" not in blob
assert "skills_instructions" not in blob
def test_reasoning_is_dropped_and_counted(self, tmp_path):
conv, report = self._normalized(tmp_path)
assert "thinking" not in json.dumps(conv)
assert "reasoning" in report.format_summary()
def test_reasoning_stays_dropped_under_full(self, tmp_path):
# Unlike Claude Code, `full` cannot surface it — it is encrypted at rest.
conv, _ = self._normalized(tmp_path, policy="full")
assert "encrypted" not in json.dumps(conv)
def test_tool_traffic_collapses_with_shortfall(self, tmp_path):
conv, _ = self._normalized(tmp_path)
collapsed = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
]
assert len(collapsed) == 1
origin = collapsed[0]["origin"]
# 2 completed of 3 attempted; the `wait` poll is not an attempt.
assert "2 calls: exec_command ×2" in origin
assert "+1 did not complete" in origin
def test_full_policy_emits_decoded_tool_blocks(self, tmp_path):
conv, _ = self._normalized(tmp_path, policy="full")
uses = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TOOL_USE
]
results = [
b for m in conv["messages"] for b in m["blocks"]
# The "incomplete" note is asserted by TestFullPolicyShortfall.
if b["type"] == BLOCK_TYPE_TOOL_RESULT and b.get("tool_name") != "incomplete"
]
assert len(uses) == 2 and len(results) == 2
# The command is read from the typed item, not parsed out of the JS.
assert uses[0]["input"]["command"] == "ls -la"
assert uses[0]["input"]["cwd"] == "/home/jesse/ws" # file:// stripped
assert results[0]["is_error"] is False
assert results[1]["is_error"] is True # exit_code 1
def test_message_count_matches(self, tmp_path):
conv, _ = self._normalized(tmp_path)
assert conv["message_count"] == len(conv["messages"])
def test_updated_at_matches_listing(self, tmp_path):
"""Cache staleness compares these two; a mismatch re-exports every run."""
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
listed = prov.list_conversations()[0]
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["updated_at"] == listed["updated_at"]
def test_context_compaction_is_visible(self, tmp_path):
records = _records() + [_item("ContextCompaction", id="c1")]
conv, report = self._normalized(tmp_path, records=records)
markers = [
b for m in conv["messages"] for b in m["blocks"]
if b.get("kind") == COLLAPSED_KIND_HIDDEN_CONTEXT
]
assert len(markers) == 1
assert "compacted" in markers[0]["origin"]
assert "context_compaction" in report.format_summary()
def test_unknown_item_type_is_reported(self, tmp_path):
records = _records() + [_item("QuantumMessage", id="q1", mystery=True)]
conv, report = self._normalized(tmp_path, records=records)
unknowns = [
b for m in conv["messages"] for b in m["blocks"] if b["type"] == "unknown"
]
assert len(unknowns) == 1
assert unknowns[0]["raw_type"] == "codex.QuantumMessage"
assert "codex.QuantumMessage" in report.format_summary()
def test_extension_labelled_by_kind(self, tmp_path):
records = _records() + [
_item("Extension", id="e1", kind="web.search", query="hsts", results=[]),
]
conv, _ = self._normalized(tmp_path, records=records)
origins = " ".join(
b["origin"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
)
assert "web.search ×1" in origins
class TestRepoTags:
def test_title_tagged_with_repos_touched(self, tmp_path):
repo = tmp_path / "ws" / "myrepo"
(repo / ".git").mkdir(parents=True)
records = _records(cwd=str(tmp_path / "ws")) + [
_item(
"FileChange",
id="fc1",
status="completed",
changes={str(repo / "README.md"): {"type": "add", "content": "x"}},
)
]
_write_session(tmp_path, records)
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["title"].endswith("[myrepo]")
def test_no_tag_when_nothing_touched(self, tmp_path):
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
assert conv["title"] == "Write the backup guide."
class TestResolveRoots:
def test_default_root(self, monkeypatch):
monkeypatch.delenv("CODEX_DIR", raising=False)
monkeypatch.delenv("CODEX_HOME", raising=False)
assert resolve_roots()[0].name == "sessions"
def test_codex_dir_splits_on_pathsep(self, monkeypatch):
monkeypatch.setenv("CODEX_DIR", "/a/sessions:/b/sessions")
monkeypatch.delenv("CODEX_HOME", raising=False)
assert [str(p) for p in resolve_roots()] == ["/a/sessions", "/b/sessions"]
def test_codex_home_appended_and_deduped(self, monkeypatch):
monkeypatch.setenv("CODEX_DIR", "/a/sessions")
monkeypatch.setenv("CODEX_HOME", "/a")
# /a/sessions is already listed — must not appear twice.
assert [str(p) for p in resolve_roots()] == ["/a/sessions"]
class TestExoticLineBreaks:
"""U+0085 (NEL) and friends are legal inside a JSON string.
``str.splitlines()`` breaks on them, shredding one record into unparseable
fragments and losing it silently. Observed 2026-08-18 in a real rollout,
where captured command output contained two NELs.
"""
# C0 controls (\x0b, \x0c) are excluded: JSON requires them escaped, so they
# never reach the splitter literally. These three do.
@pytest.mark.parametrize("sep", ["\x85", "
", "
"])
def test_record_with_exotic_break_survives(self, tmp_path, sep):
records = _records()
records.append(
_item(
"AgentMessage",
ts="2026-08-17T05:46:00.000Z",
id="a3",
phase="final_answer",
content=[{"type": "Text", "text": f"before{sep}after"}],
)
)
_write_session(tmp_path, records)
prov = CodexProvider(sessions_dir=tmp_path)
prov.list_conversations()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID))
texts = [
b["text"] for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TEXT
]
assert f"before{sep}after" in texts
class TestFullPolicyShortfall:
def test_shortfall_note_is_not_a_collapsed_block(self, tmp_path):
"""Under `full`, a collapsed block would advise setting the policy that
is already in force. The shortfall is stated plainly instead."""
_write_session(tmp_path, _records())
prov = CodexProvider(sessions_dir=tmp_path, hidden_content="full")
prov.list_conversations()
report = LossReport()
conv = prov.normalize_conversation(prov.get_conversation(SESSION_ID), report)
collapsed = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_COLLAPSED
]
assert collapsed == []
notes = [
b for m in conv["messages"] for b in m["blocks"]
if b["type"] == BLOCK_TYPE_TOOL_RESULT and b.get("tool_name") == "incomplete"
]
assert len(notes) == 1
assert "1 tool call(s) produced no result" in notes[0]["output"]
assert "tool_call_incomplete" in report.format_summary()