The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
Jacobs reference set: The Lady of the Barge (1902)
A hand-made comparison baseline for the story-to-pack pipeline, the Jacobs twin of the O. Henry
reference in ../v2/claude/. Every summary, predicament, group and option here was written by
Claude reading the scenes. No local inference was used. Compare the local-model runs against it.
Do not treat it as ground truth: it is one careful reading.
| File | What it is |
|---|---|
corpus/lady-of-the-barge.html |
Standard Ebooks single-page text (public domain) |
split_se.py |
→ stories.json (12 stories, split on <h2>) and chunks.json (157 scenes, same 320-word paragraph rule as ../chunk.py) |
dump.py |
python3 dump.py A B prints scenes A..B for reading |
claude/batches/sNN.json |
per story: {chunk, story, summary, predicament}, the same schema as ../v2/claude/batches/ |
claude/groups_src.py |
the groups, options and per-story protagonist standing |
claude/build.py |
checks everything, then writes claude/records.json and claude/groups.json |
claude/page.py |
→ claude/REFERENCE.html, the readable page |
Rebuild: python3 split_se.py && cd claude && python3 build.py && python3 page.py.
What is in it
- 157 predicaments: one sentence each (median 19 words, max 24), written to the
redescribe.pyrules: no names, jobs, places or story-only objects. Thetopic_wordsleak check flags 23 of 157. All 23 are role words (rival, partner, guest, neighbour, suitor, newcomer), kept on purpose because the design's cast is roles. - 19 groups, each a situation that recurs in two or more stories. Each scene sits in at most one group: 137 are assigned and 20 fit nothing that recurs.
- Options: 3 to 5 per group, 79 in all. 61 come from what someone in a scene did (
scene:N), 10 are archetypes and 8 fill gaps. 8 scene options cite a scene outside their own group. That is looser than DESIGN.md's "drawn from other scenes in the cluster", andgroups.jsonmarks each onein_group: false. - Moves: spine paths with a direction only (
+,-,risk,set). There are no magnitudes, as the design requires. Skill labels proposed for this corpus:nerveandcunning. - Stage comes from protagonist standing, judged per story, not per scene. That is the
coarsest judgement in the set, recorded in
STANDINGingroups_src.py.
How this corpus differs from O. Henry
It is small, 12 stories against 97. It also mixes registers: comic waterfront and village yarns, plus four horror or crime stories (The Monkey's Paw, The Well, In the Library, Captain Rogers). So several groups straddle the two registers. G09 "Your scheme turns on you" holds both a comic fake-drowning and a murderer caught out. A tool that groups by register or mood, rather than by predicament, should split them.