Files
TheLadder/tools/story-to-pack/README.md
T
JesseMarkowitzandClaude Opus 5 fa3769d0fe Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
2026-09-15 06:59:35 -04:00

81 lines
3.4 KiB
Markdown

# story-to-pack
Research toward a tool that builds a The Ladder content pack from a corpus of stories.
**Independent of the game.** Nothing in `src/`, `content/` or `test/` imports anything here, and
nothing here is part of `npm test` or `npm run build`. It lives on the `story-to-pack` branch so it is
versioned alongside the pack format it targets, without touching `main`.
- **What was decided and why:** [`DESIGN.md`](DESIGN.md)
- **How situations are selected, and every test of it:** [`SELECTION.md`](SELECTION.md)
- **The 2026-09-10 scripts, recovered from a transcript:** [`recovered/`](recovered/README.md)
## Status
Selection is settled as **gate G2 followed by a structured human review** (see below). Every automated
alternative tried — three more geometric gates and two model-judged gates — was worse on fresh,
blind-labelled data. Nothing downstream of selection (paths and directions, numbers, pack output) is
built yet.
## Layout
```
probe/
corpus/ O. Henry's Four Million cycle, Project Gutenberg (public domain)
split.py, chunk.py volumes -> stories.json -> chunks.json (838 scenes)
ollama.py the only module that talks to an inference host (host from STP_OLLAMA)
summarise.py scene -> one sentence, via the model -> summaries.json
strip_names.py names removed deterministically -> summaries_clean.json
embed.py texts -> vectors -> embeddings.json, summary_embeddings.json
measure.py, signals.py, gate2.py, gate3.py, judge.py … the selection experiments
labels-*.json, sheet-*.txt, candidates-*.json, *-map.json blind labels and what they labelled
review.py G2 on a partition -> a review package for a human
review_format.py the review form: schema, checks, conversion
apply_review.py filled review.json -> situations.json
test_review.py tests for the review form
review/<run>/ review packages: REVIEW.md, review.json, candidates.json, situations.json
```
Not committed (see `.gitignore`): the inference host's power and kernel logs, and caches that the
scripts rebuild on first use (`sim.pkl`, `null*.pkl`, `coassoc.json`). The embeddings are committed
(about 16 MB) because recreating them needs the inference host.
## Running
Everything runs from `probe/` with Python 3 and no third-party packages.
Anything that calls a model reads the host from the environment and never from a file:
```sh
export STP_OLLAMA=http://<inference-host>:11434
```
## Reviewing situations
```sh
cd probe
python3 review.py # k = 60, seed 24 -> review/k60-s24/
```
That writes three files into the run directory:
- **`REVIEW.md`** — what the program found, and exactly what the reviewer is asked to decide for
each candidate, with every scene listed. Read this.
- **`review.json`** — the form. The only file the reviewer edits.
- **`candidates.json`** — what the program found, for the checker. Do not edit.
When the form is filled in:
```sh
python3 apply_review.py review/k60-s24
```
It reports every problem at once and writes nothing until the form is complete and consistent; then
it writes **`situations.json`**, the machine-readable input for the next stage: each situation's
wording, roles and stakes, the scenes that show it (with story and summary), and which candidates it
came from.
```sh
python3 -m unittest test_review # the review form's checks
```