The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
49 lines
2.3 KiB
Python
49 lines
2.3 KiB
Python
"""Split the Standard Ebooks single-page HTML of W. W. Jacobs, *The Lady of the Barge*, into stories,
|
|
then cut scenes with the same rule as ../chunk.py (paragraph boundaries, ~320 words, tail >= 80).
|
|
Deterministic: each story is a <section> with an <h2 epub:type="title">. Front and back matter
|
|
(contents, imprint, colophon, uncopyright) are dropped by name."""
|
|
import html, json, pathlib, re
|
|
|
|
SKIP = {'Table of Contents', 'Imprint', 'Colophon', 'Uncopyright'}
|
|
TARGET = 320
|
|
VOLUME = 'The Lady of the Barge (W. W. Jacobs, 1902)'
|
|
|
|
raw = pathlib.Path('corpus/lady-of-the-barge.html').read_text(encoding='utf-8')
|
|
body = raw[raw.find('<body'):]
|
|
parts = re.split(r'<h2[^>]*epub:type="title"[^>]*>(.*?)</h2>', body, flags=re.S)
|
|
|
|
def text_of(fragment):
|
|
paras = []
|
|
for p in re.findall(r'<p[^>]*>(.*?)</p>', fragment, flags=re.S):
|
|
t = html.unescape(re.sub(r'<[^>]+>', '', p))
|
|
t = re.sub(r'\s+', ' ', t).strip()
|
|
if t: paras.append(t)
|
|
return paras
|
|
|
|
stories = []
|
|
for i in range(1, len(parts), 2):
|
|
title = html.unescape(re.sub(r'<[^>]+>', '', parts[i])).strip()
|
|
if title in SKIP: continue
|
|
paras = text_of(parts[i + 1])
|
|
stories.append({'volume': VOLUME, 'title': title, 'words': sum(len(p.split()) for p in paras),
|
|
'text': '\n\n'.join(paras)})
|
|
|
|
chunks = []
|
|
for si, s in enumerate(stories):
|
|
buf, n = [], 0
|
|
for p in s['text'].split('\n\n'):
|
|
w = len(p.split())
|
|
if n + w > TARGET and buf:
|
|
chunks.append({'story': si, 'title': s['title'], 'volume': s['volume'], 'text': ' '.join(buf), 'words': n})
|
|
buf, n = [], 0
|
|
buf.append(p); n += w
|
|
if buf and n >= 80:
|
|
chunks.append({'story': si, 'title': s['title'], 'volume': s['volume'], 'text': ' '.join(buf), 'words': n})
|
|
|
|
pathlib.Path('stories.json').write_text(json.dumps(stories, indent=1, ensure_ascii=False), encoding='utf-8')
|
|
pathlib.Path('chunks.json').write_text(json.dumps(chunks, ensure_ascii=False), encoding='utf-8')
|
|
for si, s in enumerate(stories):
|
|
print(f"{si:2} {s['title']:24} {s['words']:6} words {sum(c['story']==si for c in chunks):3} scenes")
|
|
ws = sorted(c['words'] for c in chunks)
|
|
print(f'stories {len(stories)} scenes {len(chunks)} words/scene min {ws[0]} median {ws[len(ws)//2]} max {ws[-1]} total {sum(ws):,}')
|