Files
TheLadder/tools/story-to-pack/README.md
T
JesseMarkowitzandClaude Opus 5 e5617b86ba Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.

Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.

Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.

Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.

Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.

The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.

Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
2026-09-20 17:19:42 -04:00

104 lines
5.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# story-to-pack
Research toward a tool that builds a The Ladder content pack from a corpus of stories.
**Independent of the game.** Nothing in `src/`, `content/` or `test/` imports anything here, and
nothing here is part of `npm test` or `npm run build`. It lives on the `story-to-pack` branch so it is
versioned alongside the pack format it targets, without touching `main`.
- **What was decided and why:** [`DESIGN.md`](DESIGN.md)
- **How situations are selected, and every test of it:** [`SELECTION.md`](SELECTION.md)
- **The 2026-09-10 scripts, recovered from a transcript:** [`recovered/`](recovered/README.md)
## Status
Selection was settled as **gate G2 followed by a structured human review** (see below). Every automated
alternative tried — three more geometric gates and two model-judged gates — was worse on fresh,
blind-labelled data. The first real review then found G2's groups were **topics** (artists, police,
hotels), not situations. **Now under test:** re-describe every scene as its main person's
**predicament** with qwen3:14b, regroup, and read the groups (SELECTION.md, "Predicaments"). Nothing
downstream of selection (paths and directions, numbers, pack output) is built yet.
## Layout
`probe/` holds the v1 data every G1–G4 result was measured on. `probe/v2/` is the corrected corpus
(97 stories, 839 scenes: v1's splitter merged three quoted-title stories into their neighbours), built by
running the same scripts from inside `v2/` (`python3 ../split.py`, and so on).
```
probe/
corpus/ O. Henry's Four Million cycle, Project Gutenberg (public domain)
split.py, chunk.py volumes -> stories.json -> chunks.json (838 scenes in v1, 839 in v2)
carry_summaries.py v1 summaries onto v2 scenes; only new scene text goes to the model
redescribe.py scene -> its main person's predicament, via the model -> predicaments.json
topic_words.py, topic_share.py how much groups are held together by a job, relationship or place word
pilot_page.py side-by-side page for the 3B / 14B predicament pilot -> pilot-compare.html
groups_page.py reading page for regrouped predicaments -> v2/groups-*.html
ollama.py the only module that talks to an inference host (host from STP_OLLAMA)
summarise.py scene -> one sentence, via the model -> summaries.json
strip_names.py names removed deterministically -> summaries_clean.json
embed.py texts -> vectors -> embeddings.json, summary_embeddings.json
measure.py, signals.py, gate2.py, gate3.py, judge.py … the selection experiments
labels-*.json, sheet-*.txt, candidates-*.json, *-map.json blind labels and what they labelled
review.py G2 on a partition -> REVIEW.html and candidates.json
review_page.py renders the page; picks out role words for each group
review_page.html the page template
review_page_logic.cjs the rules the page applies (also run by the Node tests)
review_format.py the saved review's format, checks and conversion
apply_review.py saved review -> situations.json
test_review.py, test_review_page.mjs tests
review/<run>/ REVIEW.html, candidates.json, and situations.json once applied
```
Not committed (see `.gitignore`): the inference host's power and kernel logs, and caches that the
scripts rebuild on first use (`sim.pkl`, `null*.pkl`, `coassoc.json`). The embeddings are committed
(about 16 MB) because recreating them needs the inference host.
## Running
Everything runs from `probe/` with Python 3 and no third-party packages.
Anything that calls a model reads the host from the environment and never from a file:
```sh
export STP_OLLAMA=http://<inference-host>:11434
```
## Reviewing situations
The review happens in a web page. Nobody edits JSON.
```sh
cd probe
python3 review.py # k = 60, seed 24 -> review/k60-s24/REVIEW.html
```
1. **Open `review/k60-s24/REVIEW.html` in a browser** (double-click it; nothing to install). The first
screen explains the job with examples.
2. **One group per screen.** For each: keep it, drop it, or mark it the same as a group already kept;
untick scenes that don't fit; and fill in "A ___ wants ___ from ___", with the people named in its
scenes offered as click-to-fill words. The page shows as you go whether a group still has enough
scenes, and colours each group in the strip across the top. Keys: K keep, D drop, ← →.
3. **Progress saves in the browser** after every change, so you can close it and come back.
4. **"Finish & save"** lists anything unfinished and saves `review-k60-s24.json`, usually to Downloads.
5. **Then:**
```sh
python3 apply_review.py review/k60-s24
```
It finds the saved review in Downloads by itself, checks everything again, and names any problem the
way the page does ("Group 4: kept, but its name is not finished"). Nothing is written until the
review passes; then it writes **`situations.json`**: each situation's sentence and roles, the scenes
that show it with their stories, and which groups it came from.
`candidates.json` in the run directory is the program's record for the checker; leave it alone.
```sh
python3 -m unittest test_review # the saved review's checks, the page renderer, apply_review.py
node --test test_review_page.mjs # the rules the page applies as you review
```
The page's drawing code is only checked for syntax and for starting up; how it looks and behaves in a
real browser has to be tried by a person.