The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
43 lines
2.3 KiB
Python
43 lines
2.3 KiB
Python
"""Carry v1 scene summaries over to a re-split corpus, summarising only scenes whose text is new.
|
|
|
|
cd v2 && STP_OLLAMA=http://host:11434 python3 ../carry_summaries.py ../summaries.json summaries.json
|
|
|
|
The v1 split merged three quoted-title stories into their neighbours (SELECTION.md, "Pilot"). Re-splitting
|
|
renumbers every later scene but leaves almost every scene's text identical, so a scene whose text matches
|
|
a v1 scene keeps its v1 summary verbatim, and only the rest go to the model, with summarise.py's model
|
|
and prompt. Output has summarise.py's layout, in the new scene order, with 'carried' marking which is which.
|
|
"""
|
|
import json, pathlib, re, sys, time
|
|
import ollama
|
|
|
|
src, dst = pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2])
|
|
old = json.loads(src.read_text(encoding='utf-8'))
|
|
old_chunks = json.loads((src.parent / 'chunks.json').read_text(encoding='utf-8'))
|
|
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
|
|
|
|
# summarise.py runs on import, so its prompt is read from the file; it must be the one v1 was made with.
|
|
code = (pathlib.Path(__file__).parent / 'summarise.py').read_text(encoding='utf-8')
|
|
SYSTEM = re.search(r'SYSTEM = """(.*?)"""', code, re.S).group(1)
|
|
END = re.compile(r'(?<!\bMr)(?<!\bMrs)(?<!\bMs)(?<!\bDr)(?<!\bSt)(?<!\bNo)[.!?](?:\s|$)')
|
|
if SYSTEM != old['prompt']:
|
|
sys.exit('summarise.py prompt differs from the one v1 summaries were made with')
|
|
|
|
by_text = {old_chunks[x['chunk']]['text']: x for x in old['summaries']}
|
|
out, fresh, t0 = [], 0, time.time()
|
|
for i, c in enumerate(chunks):
|
|
prev = by_text.get(c['text'])
|
|
if prev:
|
|
out.append({'chunk': i, 'story': c['story'], 'summary': prev['summary'],
|
|
'model_output': prev['model_output'], 'carried': prev['chunk']})
|
|
continue
|
|
text, _, _ = ollama.chat(old['model'], SYSTEM, c['text'])
|
|
t = re.sub(r'\s+', ' ', text).strip().strip('"')
|
|
m = END.search(t)
|
|
out.append({'chunk': i, 'story': c['story'], 'summary': t[:m.end()].strip() if m else t,
|
|
'model_output': text, 'carried': None})
|
|
fresh += 1
|
|
print(f' new scene {i} ({c["title"]}): {out[-1]["summary"]}', flush=True)
|
|
|
|
dst.write_text(json.dumps({'model': old['model'], 'prompt': SYSTEM, 'summaries': out}, indent=1), encoding='utf-8')
|
|
print(f'{len(out)} summaries: {len(out) - fresh} carried from {src}, {fresh} new, {time.time() - t0:.0f}s')
|