Files
TheLadder/tools/story-to-pack
JesseMarkowitzandClaude Opus 5.5 3a4758172d story-to-pack: add the corpora and the catalogue v2/v3
Fourteen corpora split and classified: aesop, bierce, chekhov, holmes,
keefe, lawson, lorimer, maupassant, nobody, plaintales, poe, torchy,
wallingford and winesburg, each with its splitter and the hand-written
groups and pages; catalogue v2 and v3; and the shared splitters
gutenberg_chunks.py, se_split.py and se_build.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rwKTmug58sEsJ72AuWsEi
2026-10-07 06:05:51 -04:00
..

story-to-pack

Research toward a tool that builds a The Ladder content pack from a corpus of stories.

Independent of the game. Nothing in src/, content/ or test/ imports anything here, and nothing here is part of npm test or npm run build. It lives on the story-to-pack branch so it is versioned alongside the pack format it targets, without touching main.

  • What was decided and why: DESIGN.md
  • How situations are selected, and every test of it: SELECTION.md
  • The 2026-09-10 scripts, recovered from a transcript: recovered/

Status

Selection was settled as gate G2 followed by a structured human review (see below). Every automated alternative tried — three more geometric gates and two model-judged gates — was worse on fresh, blind-labelled data. The first real review then found G2's groups were topics (artists, police, hotels), not situations. Now under test: re-describe every scene as its main person's predicament with qwen3:14b, regroup, and read the groups (SELECTION.md, "Predicaments"). Nothing downstream of selection (paths and directions, numbers, pack output) is built yet.

Layout

probe/ holds the v1 data every G1–G4 result was measured on. probe/v2/ is the corrected corpus (97 stories, 839 scenes: v1's splitter merged three quoted-title stories into their neighbours), built by running the same scripts from inside v2/ (python3 ../split.py, and so on).

probe/
  corpus/                O. Henry's Four Million cycle, Project Gutenberg (public domain)
  split.py, chunk.py     volumes -> stories.json -> chunks.json (838 scenes in v1, 839 in v2)
  carry_summaries.py     v1 summaries onto v2 scenes; only new scene text goes to the model
  redescribe.py          scene -> its main person's predicament, via the model -> predicaments.json
  topic_words.py, topic_share.py   how much groups are held together by a job, relationship or place word
  pilot_page.py          side-by-side page for the 3B / 14B predicament pilot -> pilot-compare.html
  groups_page.py         reading page for regrouped predicaments -> v2/groups-*.html
  ollama.py              the only module that talks to an inference host (host from STP_OLLAMA)
  summarise.py           scene -> one sentence, via the model          -> summaries.json
  strip_names.py         names removed deterministically               -> summaries_clean.json
  embed.py               texts -> vectors                              -> embeddings.json, summary_embeddings.json
  measure.py, signals.py, gate2.py, gate3.py, judge.py …   the selection experiments
  labels-*.json, sheet-*.txt, candidates-*.json, *-map.json   blind labels and what they labelled
  review.py              G2 on a partition -> REVIEW.html and candidates.json
  review_page.py         renders the page; picks out role words for each group
  review_page.html       the page template
  review_page_logic.cjs  the rules the page applies (also run by the Node tests)
  review_format.py       the saved review's format, checks and conversion
  apply_review.py        saved review -> situations.json
  test_review.py, test_review_page.mjs   tests
  review/<run>/          REVIEW.html, candidates.json, and situations.json once applied

Not committed (see .gitignore): the inference host's power and kernel logs, and caches that the scripts rebuild on first use (sim.pkl, null*.pkl, coassoc.json). The embeddings are committed (about 16 MB) because recreating them needs the inference host.

Running

Everything runs from probe/ with Python 3 and no third-party packages.

Anything that calls a model reads the host from the environment and never from a file:

export STP_OLLAMA=http://<inference-host>:11434

Reviewing situations

The review happens in a web page. Nobody edits JSON.

cd probe
python3 review.py                      # k = 60, seed 24  ->  review/k60-s24/REVIEW.html
  1. Open review/k60-s24/REVIEW.html in a browser (double-click it; nothing to install). The first screen explains the job with examples.

  2. One group per screen. For each: keep it, drop it, or mark it the same as a group already kept; untick scenes that don't fit; and fill in "A ___ wants ___ from ___", with the people named in its scenes offered as click-to-fill words. The page shows as you go whether a group still has enough scenes, and colours each group in the strip across the top. Keys: K keep, D drop, ← →.

  3. Progress saves in the browser after every change, so you can close it and come back.

  4. "Finish & save" lists anything unfinished and saves review-k60-s24.json, usually to Downloads.

  5. Then:

    python3 apply_review.py review/k60-s24
    

    It finds the saved review in Downloads by itself, checks everything again, and names any problem the way the page does ("Group 4: kept, but its name is not finished"). Nothing is written until the review passes; then it writes situations.json: each situation's sentence and roles, the scenes that show it with their stories, and which groups it came from.

candidates.json in the run directory is the program's record for the checker; leave it alone.

python3 -m unittest test_review        # the saved review's checks, the page renderer, apply_review.py
node --test test_review_page.mjs       # the rules the page applies as you review

The page's drawing code is only checked for syntax and for starting up; how it looks and behaves in a real browser has to be tried by a person.