Files
TheLadder/tools/story-to-pack/probe/wharton
JesseMarkowitzandClaude Opus 5 e5617b86ba Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.

Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.

Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.

Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.

Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.

The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.

Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
2026-09-20 17:19:42 -04:00
..

Wharton reference: The Descent of Man and Other Stories

A hand-made comparison target for story-to-pack, the third alongside O. Henry (../v2/claude/) and Jacobs (../jacobs/). Claude read every scene of Edith Wharton's 1904 collection and wrote its summary and predicament. The recurring groups, a draft event per group, and a setting were then written by hand. The local-inference pipeline can be run on the same chunks.json and set beside it scene for scene.

It is a reference, not an answer key: a different careful reader would draw some group boundaries differently. No inference host was used.

Open claude/REFERENCE.html in a browser to read it.

Files

File What
corpus/descent-of-man.txt Project Gutenberg #4519, plain text as downloaded 2026-09-17
split_gutenberg.py Splits the text into the ten stories and cuts ~320-word scenes, with the same rule as ../chunk.py and ../jacobs/split_se.py
stories.json, chunks.json 10 stories, 262 scenes (median 285 words)
dump.py dump.py 0 28 prints scenes for reading
claude/batches/s00-s09.json One file per story: {chunk, story, summary, predicament} per scene, the same shape as the sibling probes
claude/summaries.json, claude/predicaments.json The same, in the {"summaries": [{chunk, summary}]} shape embed.py, topic_share.py and groups_page.py read. Written by claude_export.py
claude/groups.json 16 recurring groups, each with its members, a one-sentence predicament and a tightness label
claude/events.json One draft event per group: description variants citing their source scenes, and 3-4 options with effects as a path and a direction only
claude/setting.json Stat-spine labels, roles, stages and a suggested player role
claude/measure.log claude_export.py output: topic share of the hand groups, and word leakage in the predicaments
claude_page.py → claude/REFERENCE.html The reading page

Rules followed

  • Predicament, as SELECTION.md defines it: what the main person is up against and what they want, with no names, jobs, places or story-specific objects. Relationship words ("partner", "spouse", "rival") are kept, as the spec allows.
  • Group: at least 3 of the 10 stories, and no story more than 50% of the members. O. Henry uses at least 4 of 97 stories and at most 40%; with only ten stories that bar would leave almost nothing. Each scene is in at most one group.
  • Options follow the DESIGN.md recipe: fixed archetypes (18), options drawn from scenes (41), and options filled in where the corpus had none (2).
  • No numbers. Effects give a path and a direction. Magnitudes are the numeric search's job.
  • Role tokens ({spouse}, {former} …) are left in for code to substitute, because DESIGN.md found a small model should not do entity substitution.

What came out

  • 16 groups covering 129 of 262 scenes (49%). 12 are tight (one predicament) and 4 are broad. Two (own_words_back, owning_up) read as climaxes rather than things that recur in a run, and are marked recurs: false.
  • A third of the options come from a scene next to a member, not from another member. Of the 41 options drawn from scenes, 17 came from outside their group. 13 of those 17 are within two scenes of a member in the same story: 8 come from the scene after, where the character acts, and 5 from the scene before, which sets up what they are responding to. A tool that draws options only from cluster members would miss about a third of these. It should also read the scene or two either side of each member.
  • Topic share for the hand groups is 39%, against the O. Henry baselines of 20% (summaries) and 13% (14B predicaments). This is not a fair like-for-like, since the corpus and the summariser differ, but the per-group list shows why it is high: husband, wife and mother top 9 of the 16 groups. Wharton's predicaments are domestic, so the relationship word is the predicament ("giving ground to keep the peace" needs a spouse). The topic-share measure cannot tell a relationship that makes up a predicament from one that is only a topic. On a corpus like this one, it would mark good groups as topic-bound.
  • Ten stories is below the design's 20-30. Nine of the 16 groups sit exactly at the 50% ceiling. The corpus does recur (marriage, compromise, secrets, the past coming back), but thinly. A pack needs ~15 situations a stage, and this gives 16 in all, across two stages.

Known defects

  • Two Gutenberg epigraph lines were left inside scene text (scene 45 carries a stray line from scene 56, and scene 51 carries "Earth's Martyrs. By Stephen Phillips."). They are not removed, so that chunks.json stays the corpus as split.
  • Scenes 0-6 and 237-239 are narration with little predicament. They were summarised anyway, and none of them is in a group.
  • Two descriptions cite a scene one before their group's members (50, 122). This is recorded in events.json.