Files
TheLadder/tools/story-to-pack/probe/jacobs/claude/build.py
T
JesseMarkowitzandClaude Opus 5 e5617b86ba Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.

Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.

Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.

Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.

Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.

The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.

Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
2026-09-20 17:19:42 -04:00

43 lines
2.6 KiB
Python

#!/usr/bin/env python3
"""Assemble records.json and groups.json from batches/ and groups_src.py, and check them."""
import collections, glob, json, pathlib
from groups_src import G, STANDING
here = pathlib.Path(__file__).parent
chunks = json.loads((here.parent / 'chunks.json').read_text())
stories = json.loads((here.parent / 'stories.json').read_text())
recs = [r for f in sorted(glob.glob(str(here / 'batches/s*.json'))) for r in json.loads(open(f).read())]
assert [r['chunk'] for r in recs] == list(range(len(chunks))), 'every scene, once, in order'
assert all(chunks[r['chunk']]['story'] == r['story'] for r in recs)
seen = collections.Counter(m for g in G for m in g['members'])
dup = [m for m, n in seen.items() if n > 1]
assert not dup, f'scenes in two groups: {dup}'
by = {r['chunk']: r for r in recs}
out = []
for g in G:
assert 3 <= len(g['options']) <= 5, g['id']
st = sorted({by[m]['story'] for m in g['members']})
stand = collections.Counter(STANDING[by[m]['story']] for m in g['members'])
out.append({
'id': g['id'], 'name': g['name'], 'predicament': g['predicament'],
'sentence': f"A {g['actor']} wants {g['wants']} from {g['counterpart']}.",
'actor': g['actor'], 'wants': g['wants'], 'counterpart': g['counterpart'],
'stories': st, 'story_titles': [stories[s]['title'] for s in st],
'standing': dict(stand), 'stage': stand.most_common(1)[0][0],
'members': [{'chunk': m, 'story': by[m]['story'], 'predicament': by[m]['predicament']} for m in g['members']],
'options': [{'text': t, 'source': s, 'moves': mv,
'in_group': (int(s.split(':')[1]) in g['members']) if s.startswith('scene:') else None}
for t, s, mv in g['options']],
})
unassigned = [r['chunk'] for r in recs if r['chunk'] not in seen]
(here / 'records.json').write_text(json.dumps(recs, indent=1, ensure_ascii=False))
(here / 'groups.json').write_text(json.dumps({'groups': out, 'unassigned': unassigned}, indent=1, ensure_ascii=False))
multi = sum(len(g['stories']) >= 2 for g in out)
print(f'{len(recs)} scenes, {len(out)} groups ({multi} span 2+ stories), '
f'{sum(seen.values())} assigned, {len(unassigned)} unassigned: {unassigned}')
for g in out:
print(f"{g['id']} {len(g['members']):2} scenes {len(g['stories'])} stories stage {g['stage']:6} {g['name']}")
src = collections.Counter(o['source'].split(':')[0] for g in out for o in g['options'])
print('option sources:', dict(src), '| scene options citing a scene outside their group:',
sum(o['in_group'] is False for g in out for o in g['options']))