# story-to-pack Research toward a tool that builds a The Ladder content pack from a corpus of stories. **Independent of the game.** Nothing in `src/`, `content/` or `test/` imports anything here, and nothing here is part of `npm test` or `npm run build`. It lives on the `story-to-pack` branch so it is versioned alongside the pack format it targets, without touching `main`. - **What was decided and why:** [`DESIGN.md`](DESIGN.md) - **How situations are selected, and every test of it:** [`SELECTION.md`](SELECTION.md) - **The 2026-09-10 scripts, recovered from a transcript:** [`recovered/`](recovered/README.md) ## Status Selection was settled as **gate G2 followed by a structured human review** (see below). Every automated alternative tried — three more geometric gates and two model-judged gates — was worse on fresh, blind-labelled data. The first real review then found G2's groups were **topics** (artists, police, hotels), not situations. **Now under test:** re-describe every scene as its main person's **predicament** with qwen3:14b, regroup, and read the groups (SELECTION.md, "Predicaments"). Nothing downstream of selection (paths and directions, numbers, pack output) is built yet. ## Layout `probe/` holds the v1 data every G1–G4 result was measured on. `probe/v2/` is the corrected corpus (97 stories, 839 scenes: v1's splitter merged three quoted-title stories into their neighbours), built by running the same scripts from inside `v2/` (`python3 ../split.py`, and so on). ``` probe/ corpus/ O. Henry's Four Million cycle, Project Gutenberg (public domain) split.py, chunk.py volumes -> stories.json -> chunks.json (838 scenes in v1, 839 in v2) carry_summaries.py v1 summaries onto v2 scenes; only new scene text goes to the model redescribe.py scene -> its main person's predicament, via the model -> predicaments.json topic_words.py, topic_share.py how much groups are held together by a job, relationship or place word pilot_page.py side-by-side page for the 3B / 14B predicament pilot -> pilot-compare.html groups_page.py reading page for regrouped predicaments -> v2/groups-*.html ollama.py the only module that talks to an inference host (host from STP_OLLAMA) summarise.py scene -> one sentence, via the model -> summaries.json strip_names.py names removed deterministically -> summaries_clean.json embed.py texts -> vectors -> embeddings.json, summary_embeddings.json measure.py, signals.py, gate2.py, gate3.py, judge.py … the selection experiments labels-*.json, sheet-*.txt, candidates-*.json, *-map.json blind labels and what they labelled review.py G2 on a partition -> REVIEW.html and candidates.json review_page.py renders the page; picks out role words for each group review_page.html the page template review_page_logic.cjs the rules the page applies (also run by the Node tests) review_format.py the saved review's format, checks and conversion apply_review.py saved review -> situations.json test_review.py, test_review_page.mjs tests review// REVIEW.html, candidates.json, and situations.json once applied ``` Not committed (see `.gitignore`): the inference host's power and kernel logs, and caches that the scripts rebuild on first use (`sim.pkl`, `null*.pkl`, `coassoc.json`). The embeddings are committed (about 16 MB) because recreating them needs the inference host. ## Running Everything runs from `probe/` with Python 3 and no third-party packages. Anything that calls a model reads the host from the environment and never from a file: ```sh export STP_OLLAMA=http://:11434 ``` ## Reviewing situations The review happens in a web page. Nobody edits JSON. ```sh cd probe python3 review.py # k = 60, seed 24 -> review/k60-s24/REVIEW.html ``` 1. **Open `review/k60-s24/REVIEW.html` in a browser** (double-click it; nothing to install). The first screen explains the job with examples. 2. **One group per screen.** For each: keep it, drop it, or mark it the same as a group already kept; untick scenes that don't fit; and fill in "A ___ wants ___ from ___", with the people named in its scenes offered as click-to-fill words. The page shows as you go whether a group still has enough scenes, and colours each group in the strip across the top. Keys: K keep, D drop, ← →. 3. **Progress saves in the browser** after every change, so you can close it and come back. 4. **"Finish & save"** lists anything unfinished and saves `review-k60-s24.json`, usually to Downloads. 5. **Then:** ```sh python3 apply_review.py review/k60-s24 ``` It finds the saved review in Downloads by itself, checks everything again, and names any problem the way the page does ("Group 4: kept, but its name is not finished"). Nothing is written until the review passes; then it writes **`situations.json`**: each situation's sentence and roles, the scenes that show it with their stories, and which groups it came from. `candidates.json` in the run directory is the program's record for the checker; leave it alone. ```sh python3 -m unittest test_review # the saved review's checks, the page renderer, apply_review.py node --test test_review_page.mjs # the rules the page applies as you review ``` The page's drawing code is only checked for syntax and for starting up; how it looks and behaves in a real browser has to be tried by a person.