The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
48 lines
2.4 KiB
Python
48 lines
2.4 KiB
Python
"""How much groups are held together by a topic word rather than a predicament.
|
|
|
|
python3 topic_share.py summary_embeddings.json # baseline: groups of scene summaries
|
|
python3 topic_share.py predicament_embeddings.json # groups of re-described predicaments
|
|
|
|
Clusters the given embeddings at k = 60 (seeds 11, 12, 13). For each candidate group (at least 5
|
|
scenes, at least 4 stories, no story over 40%), finds the job, relationship or place word most
|
|
common in the group's ORIGINAL scene summaries, and the share of its scenes that contain it. The
|
|
first human review read groups that were artists, police, courtship and hotels: that is this
|
|
number being high. Reported as the median over groups, then the median over seeds.
|
|
|
|
No inference.
|
|
"""
|
|
import collections, json, pathlib, re, statistics, sys
|
|
from topic_words import TOPIC_WORDS
|
|
|
|
path = sys.argv[1] if len(sys.argv) > 1 else 'summary_embeddings.json'
|
|
sys.argv = ['measure.py', path]
|
|
exec(pathlib.Path(__file__).with_name('measure.py').read_text(encoding='utf-8').split("print(f'{path.name}")[0])
|
|
# now defined: chunks, V, n, story, kmeans, describe, MIN_SIZE, MIN_STORIES, MAX_DOMINANT
|
|
|
|
original = {x['chunk']: x['summary'] for x in
|
|
json.loads(pathlib.Path('summaries_clean.json').read_text(encoding='utf-8'))['summaries']}
|
|
words = {i: set(re.findall(r"[a-zé]+", original.get(i, '').lower())) & TOPIC_WORDS for i in range(n)}
|
|
|
|
K, SEEDS = 60, (11, 12, 13)
|
|
per_seed, examples = [], []
|
|
for seed in SEEDS:
|
|
shares = []
|
|
for group in kmeans(K, seed):
|
|
size, nst, dom, _ = describe(group)
|
|
if size < MIN_SIZE or nst < MIN_STORIES or dom > MAX_DOMINANT:
|
|
continue
|
|
counts = collections.Counter(w for i in group for w in words[i])
|
|
if counts:
|
|
word, hits = counts.most_common(1)[0]
|
|
else:
|
|
word, hits = '-', 0
|
|
shares.append(hits / size)
|
|
if seed == SEEDS[0]:
|
|
examples.append((hits / size, word, size))
|
|
per_seed.append((len(shares), statistics.median(shares) if shares else 0.0))
|
|
print(f'seed {seed}: {len(shares)} candidate groups, median top topic-word share {per_seed[-1][1]:.0%}', flush=True)
|
|
|
|
print(f'\n{pathlib.Path(path).name}: median over seeds {statistics.median(s for _, s in per_seed):.0%}, '
|
|
f'candidate groups {statistics.median(c for c, _ in per_seed):.0f}')
|
|
print('most topic-bound groups at seed 11:', ', '.join(f'{w} {s:.0%} of {z}' for s, w, z in sorted(examples, reverse=True)[:8]))
|