Files
TheLadder/tools/story-to-pack/probe/sample_clusters.py
T
JesseMarkowitzandClaude Opus 5 fa3769d0fe Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
2026-09-15 06:59:35 -04:00

37 lines
1.8 KiB
Python

"""Read what the tight-recurring clusters actually contain (pass-mark criterion 3).
python3 sample_clusters.py summary_embeddings.json summaries_clean.json [k] [seed]
Uses the same clustering and the same definitions as measure.py, picks three
tight-recurring clusters at random (seeded, so the pick is reproducible and is
not chosen by looking), and prints each member's cleaned summary with the story
it came from. The question for a human reader: is each one recognisably ONE
kind of situation, or a prose register, or a shared vocabulary?
"""
import json, pathlib, random, statistics, sys
emb = sys.argv[1] if len(sys.argv) > 1 else 'summary_embeddings.json'
summ = sys.argv[2] if len(sys.argv) > 2 else 'summaries_clean.json'
K = int(sys.argv[3]) if len(sys.argv) > 3 else 60
SEED = int(sys.argv[4]) if len(sys.argv) > 4 else 11
sys.argv = ['measure.py', emb]
src = pathlib.Path('measure.py').read_text(encoding='utf-8')
exec(src.split("print(f'{path.name}")[0]) # loading, kmeans, describe, constants
summaries = json.loads(pathlib.Path(summ).read_text(encoding='utf-8'))['summaries']
text = {x['chunk']: x['summary'] for x in summaries}
clusters = kmeans(K, SEED)
rows = [(c, describe(c)) for c in clusters]
sized = [(c, d) for c, d in rows if d[0] >= MIN_SIZE]
cut = statistics.median(d[3] for _, d in sized)
tight = [(c, d) for c, d in sized if d[1] >= MIN_STORIES and d[2] <= MAX_DOMINANT and d[3] >= cut]
print(f'{emb}: k={K} seed={SEED}: {len(tight)} tight-recurring clusters of {len(sized)} sized\n')
for n, (c, (size, nst, dom, coh)) in enumerate(random.Random(SEED).sample(tight, min(3, len(tight))), 1):
print(f'=== sample {n}: size {size}, {nst} stories, dominant {dom:.0%}, coherence {coh:.3f}')
for i in c:
print(f" [{chunks[i]['title'][:28]:28}] {text.get(i, '(no summary)')}")
print()