Research toward building a content pack from a story corpus, kept on its own branch and independent of the game. Records the selection experiments against blind labels, and settles selection as gate G2 followed by a human review: review.py writes REVIEW.md and a review.json form, apply_review.py checks the filled form and writes situations.json for the next stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
37 lines
1.8 KiB
Python
37 lines
1.8 KiB
Python
"""Read what the tight-recurring clusters actually contain (pass-mark criterion 3).
|
|
|
|
python3 sample_clusters.py summary_embeddings.json summaries_clean.json [k] [seed]
|
|
|
|
Uses the same clustering and the same definitions as measure.py, picks three
|
|
tight-recurring clusters at random (seeded, so the pick is reproducible and is
|
|
not chosen by looking), and prints each member's cleaned summary with the story
|
|
it came from. The question for a human reader: is each one recognisably ONE
|
|
kind of situation, or a prose register, or a shared vocabulary?
|
|
"""
|
|
import json, pathlib, random, statistics, sys
|
|
|
|
emb = sys.argv[1] if len(sys.argv) > 1 else 'summary_embeddings.json'
|
|
summ = sys.argv[2] if len(sys.argv) > 2 else 'summaries_clean.json'
|
|
K = int(sys.argv[3]) if len(sys.argv) > 3 else 60
|
|
SEED = int(sys.argv[4]) if len(sys.argv) > 4 else 11
|
|
|
|
sys.argv = ['measure.py', emb]
|
|
src = pathlib.Path('measure.py').read_text(encoding='utf-8')
|
|
exec(src.split("print(f'{path.name}")[0]) # loading, kmeans, describe, constants
|
|
|
|
summaries = json.loads(pathlib.Path(summ).read_text(encoding='utf-8'))['summaries']
|
|
text = {x['chunk']: x['summary'] for x in summaries}
|
|
|
|
clusters = kmeans(K, SEED)
|
|
rows = [(c, describe(c)) for c in clusters]
|
|
sized = [(c, d) for c, d in rows if d[0] >= MIN_SIZE]
|
|
cut = statistics.median(d[3] for _, d in sized)
|
|
tight = [(c, d) for c, d in sized if d[1] >= MIN_STORIES and d[2] <= MAX_DOMINANT and d[3] >= cut]
|
|
print(f'{emb}: k={K} seed={SEED}: {len(tight)} tight-recurring clusters of {len(sized)} sized\n')
|
|
|
|
for n, (c, (size, nst, dom, coh)) in enumerate(random.Random(SEED).sample(tight, min(3, len(tight))), 1):
|
|
print(f'=== sample {n}: size {size}, {nst} stories, dominant {dom:.0%}, coherence {coh:.3f}')
|
|
for i in c:
|
|
print(f" [{chunks[i]['title'][:28]:28}] {text.get(i, '(no summary)')}")
|
|
print()
|