The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
Wharton reference: The Descent of Man and Other Stories
A hand-made comparison target for story-to-pack, the third alongside O. Henry (../v2/claude/)
and Jacobs (../jacobs/). Claude read every scene of Edith Wharton's 1904 collection and wrote
its summary and predicament. The recurring groups, a draft event per group, and a setting were
then written by hand. The local-inference pipeline can be run on the same chunks.json and set
beside it scene for scene.
It is a reference, not an answer key: a different careful reader would draw some group boundaries differently. No inference host was used.
Open claude/REFERENCE.html in a browser to read it.
Files
| File | What |
|---|---|
corpus/descent-of-man.txt |
Project Gutenberg #4519, plain text as downloaded 2026-09-17 |
split_gutenberg.py |
Splits the text into the ten stories and cuts ~320-word scenes, with the same rule as ../chunk.py and ../jacobs/split_se.py |
stories.json, chunks.json |
10 stories, 262 scenes (median 285 words) |
dump.py |
dump.py 0 28 prints scenes for reading |
claude/batches/s00-s09.json |
One file per story: {chunk, story, summary, predicament} per scene, the same shape as the sibling probes |
claude/summaries.json, claude/predicaments.json |
The same, in the {"summaries": [{chunk, summary}]} shape embed.py, topic_share.py and groups_page.py read. Written by claude_export.py |
claude/groups.json |
16 recurring groups, each with its members, a one-sentence predicament and a tightness label |
claude/events.json |
One draft event per group: description variants citing their source scenes, and 3-4 options with effects as a path and a direction only |
claude/setting.json |
Stat-spine labels, roles, stages and a suggested player role |
claude/measure.log |
claude_export.py output: topic share of the hand groups, and word leakage in the predicaments |
claude_page.py → claude/REFERENCE.html |
The reading page |
Rules followed
- Predicament, as SELECTION.md defines it: what the main person is up against and what they want, with no names, jobs, places or story-specific objects. Relationship words ("partner", "spouse", "rival") are kept, as the spec allows.
- Group: at least 3 of the 10 stories, and no story more than 50% of the members. O. Henry uses at least 4 of 97 stories and at most 40%; with only ten stories that bar would leave almost nothing. Each scene is in at most one group.
- Options follow the DESIGN.md recipe: fixed archetypes (18), options drawn from scenes (41), and options filled in where the corpus had none (2).
- No numbers. Effects give a path and a direction. Magnitudes are the numeric search's job.
- Role tokens (
{spouse},{former}…) are left in for code to substitute, because DESIGN.md found a small model should not do entity substitution.
What came out
- 16 groups covering 129 of 262 scenes (49%). 12 are tight (one predicament) and 4 are broad.
Two (
own_words_back,owning_up) read as climaxes rather than things that recur in a run, and are markedrecurs: false. - A third of the options come from a scene next to a member, not from another member. Of the 41 options drawn from scenes, 17 came from outside their group. 13 of those 17 are within two scenes of a member in the same story: 8 come from the scene after, where the character acts, and 5 from the scene before, which sets up what they are responding to. A tool that draws options only from cluster members would miss about a third of these. It should also read the scene or two either side of each member.
- Topic share for the hand groups is 39%, against the O. Henry baselines of 20% (summaries)
and 13% (14B predicaments). This is not a fair like-for-like, since the corpus and the
summariser differ, but the per-group list shows why it is high:
husband,wifeandmothertop 9 of the 16 groups. Wharton's predicaments are domestic, so the relationship word is the predicament ("giving ground to keep the peace" needs a spouse). The topic-share measure cannot tell a relationship that makes up a predicament from one that is only a topic. On a corpus like this one, it would mark good groups as topic-bound. - Ten stories is below the design's 20-30. Nine of the 16 groups sit exactly at the 50% ceiling. The corpus does recur (marriage, compromise, secrets, the past coming back), but thinly. A pack needs ~15 situations a stage, and this gives 16 in all, across two stages.
Known defects
- Two Gutenberg epigraph lines were left inside scene text (scene 45 carries a stray line from
scene 56, and scene 51 carries "Earth's Martyrs. By Stephen Phillips."). They are not removed,
so that
chunks.jsonstays the corpus as split. - Scenes 0-6 and 237-239 are narration with little predicament. They were summarised anyway, and none of them is in a group.
- Two descriptions cite a scene one before their group's members (50, 122). This is recorded in
events.json.