Files
TheLadder/tools/story-to-pack
JesseMarkowitzandClaude Opus 5 fa3769d0fe Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
2026-09-15 06:59:35 -04:00
..

story-to-pack

Research toward a tool that builds a The Ladder content pack from a corpus of stories.

Independent of the game. Nothing in src/, content/ or test/ imports anything here, and nothing here is part of npm test or npm run build. It lives on the story-to-pack branch so it is versioned alongside the pack format it targets, without touching main.

  • What was decided and why: DESIGN.md
  • How situations are selected, and every test of it: SELECTION.md
  • The 2026-09-10 scripts, recovered from a transcript: recovered/

Status

Selection is settled as gate G2 followed by a structured human review (see below). Every automated alternative tried — three more geometric gates and two model-judged gates — was worse on fresh, blind-labelled data. Nothing downstream of selection (paths and directions, numbers, pack output) is built yet.

Layout

probe/
  corpus/                O. Henry's Four Million cycle, Project Gutenberg (public domain)
  split.py, chunk.py     volumes -> stories.json -> chunks.json (838 scenes)
  ollama.py              the only module that talks to an inference host (host from STP_OLLAMA)
  summarise.py           scene -> one sentence, via the model          -> summaries.json
  strip_names.py         names removed deterministically               -> summaries_clean.json
  embed.py               texts -> vectors                              -> embeddings.json, summary_embeddings.json
  measure.py, signals.py, gate2.py, gate3.py, judge.py …   the selection experiments
  labels-*.json, sheet-*.txt, candidates-*.json, *-map.json   blind labels and what they labelled
  review.py              G2 on a partition -> a review package for a human
  review_format.py       the review form: schema, checks, conversion
  apply_review.py        filled review.json -> situations.json
  test_review.py         tests for the review form
  review/<run>/          review packages: REVIEW.md, review.json, candidates.json, situations.json

Not committed (see .gitignore): the inference host's power and kernel logs, and caches that the scripts rebuild on first use (sim.pkl, null*.pkl, coassoc.json). The embeddings are committed (about 16 MB) because recreating them needs the inference host.

Running

Everything runs from probe/ with Python 3 and no third-party packages.

Anything that calls a model reads the host from the environment and never from a file:

export STP_OLLAMA=http://<inference-host>:11434

Reviewing situations

cd probe
python3 review.py                      # k = 60, seed 24  ->  review/k60-s24/

That writes three files into the run directory:

  • REVIEW.md — what the program found, and exactly what the reviewer is asked to decide for each candidate, with every scene listed. Read this.
  • review.json — the form. The only file the reviewer edits.
  • candidates.json — what the program found, for the checker. Do not edit.

When the form is filled in:

python3 apply_review.py review/k60-s24

It reports every problem at once and writes nothing until the form is complete and consistent; then it writes situations.json, the machine-readable input for the next stage: each situation's wording, roles and stakes, the scenes that show it (with story and summary), and which candidates it came from.

python3 -m unittest test_review        # the review form's checks