Add story-to-pack research and a structured situation review

Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
This commit is contained in:
JesseMarkowitz
2026-09-15 06:59:35 -04:00
co-authored by Claude Opus 5
parent d84ae495f4
commit fa3769d0fe
112 changed files with 79868 additions and 0 deletions
+53
View File
@@ -0,0 +1,53 @@
# story-to-pack — recovered probe scripts
Recovered 2026-09-11 from the Claude Code transcript of session
`d1409289-cea9-4399-a827-2f13479647bc` (2026-09-10). The originals lived in that
session's `/tmp` scratchpad and were lost when the machine rebooted. These are
the last version of each file as the session wrote it. None of them has been
re-run since recovery.
This is research for a tool that builds a The Ladder content pack from a corpus
of stories. It is deliberately independent of `theladder/` — it would live in
that repo's `tools/`, never imported by `src/`, `content/` or `test/`, or on its
own.
| File | What it did |
| --- | --- |
| `probe.mjs`, `probe2.mjs`, `probe3.mjs` | Randomised the Frontier's numbers and ran the conformance checks: random magnitudes pass 3%, random directions 0%, a 4-rule repair loop 42% |
| `split.py` | Splits Gutenberg O. Henry volumes into individual stories |
| `chunk.py` | Cuts stories into ~300-word scenes on paragraph boundaries |
| `embed.py` | Embeds scenes with `nomic-embed-text` on the LAN Ollama host |
| `cluster.py` | k-means over the embeddings; reports how many stories each cluster draws from |
| `sweep.py` | Sweeps k and measures coherence against cross-story spread |
The probes expect to run from inside `theladder/` (they import its engine and
packs). The Python scripts expect the corpus in a `corpus/` directory beside
them.
## Rebuilding the corpus
O. Henry's *Four Million* cycle, Project Gutenberg ids 2776, 1444, 3707, 2141:
```sh
mkdir -p corpus
for id in 2776 1444 3707 2141; do
curl -sL -o corpus/pg$id.txt "https://www.gutenberg.org/cache/epub/$id/pg$id.txt"
done
```
That gave 94 stories, 244,690 words, 838 scenes. Embedding all of them took
about 20 minutes (0.6 chunks/sec).
## Where the research stopped
- Clustering the raw prose failed. Only 3 of 18 candidate clusters recurred
across stories (2% of the corpus). At ~300 words, embedding similarity tracks
which story a scene is from, not what kind of situation it is.
- The proposed fix, untested: summarise each scene into one generic sentence
with no proper names, embed the summary, then cluster.
- It could not be tested because generation on the LAN Ollama host hung. The
model ran on CPU with no GPU offload, and the 16k-context variant stayed
loaded. That is probably the same fault that blocks Interactive Story M01.
- The last open question was whether to use Claude for the generation steps
until that host is fixed. Embeddings stay on Ollama, because Anthropic has no
embeddings API.
+26
View File
@@ -0,0 +1,26 @@
"""Cut stories into scene-sized chunks on paragraph boundaries."""
import json, pathlib, re
TARGET = 320 # words; a scene is a situation with people in it
stories = json.loads(pathlib.Path('stories.json').read_text(encoding='utf-8'))
chunks = []
for si, s in enumerate(stories):
paras = [p.strip() for p in re.split(r'\n\s*\n', s['text']) if p.strip()]
buf, n = [], 0
for p in paras:
w = len(p.split())
if n + w > TARGET and buf:
chunks.append({'story': si, 'title': s['title'], 'volume': s['volume'],
'text': ' '.join(buf), 'words': n})
buf, n = [], 0
buf.append(p); n += w
if buf and n >= 80:
chunks.append({'story': si, 'title': s['title'], 'volume': s['volume'],
'text': ' '.join(buf), 'words': n})
pathlib.Path('chunks.json').write_text(json.dumps(chunks), encoding='utf-8')
ws = sorted(c['words'] for c in chunks)
print(f'chunks: {len(chunks)} from {len(stories)} stories')
print(f'words/chunk min {ws[0]} median {ws[len(ws)//2]} max {ws[-1]}')
print(f'total words: {sum(ws):,}')
+80
View File
@@ -0,0 +1,80 @@
"""Cluster scene embeddings and report what the clusters actually contain.
The question this answers: does a thematically coherent corpus contain
RECURRING situations, or 800 singletons? A pack needs ~15 repeatable
situations per stage; a cluster that draws from many different stories is a
recurring situation, one that draws from a single story is just a scene.
"""
import json, math, pathlib, random, sys
K = int(sys.argv[1]) if len(sys.argv) > 1 else 40
SEED = 11
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
vecs = json.loads(pathlib.Path('embeddings.json').read_text(encoding='utf-8'))
n = min(len(chunks), len(vecs))
chunks, vecs = chunks[:n], vecs[:n]
print(f'clustering {n} chunks into k={K}\n')
def norm(v):
m = math.sqrt(sum(x * x for x in v)) or 1.0
return [x / m for x in v]
vecs = [norm(v) for v in vecs]
dim = len(vecs[0])
def dot(a, b): return sum(x * y for x, y in zip(a, b))
# k-means++ init on cosine distance (vectors are unit length, so dot == cos)
random.seed(SEED)
cent = [vecs[random.randrange(n)]]
while len(cent) < K:
d2 = [min((1 - dot(v, c)) for c in cent) ** 2 for v in vecs]
tot = sum(d2) or 1.0
r = random.random() * tot
acc = 0.0
for i, x in enumerate(d2):
acc += x
if acc >= r:
cent.append(vecs[i]); break
else:
cent.append(vecs[random.randrange(n)])
assign = [0] * n
for it in range(30):
moved = 0
for i, v in enumerate(vecs):
best, bs = 0, -2.0
for k, c in enumerate(cent):
s = dot(v, c)
if s > bs: bs, best = s, k
if assign[i] != best: assign[i] = best; moved += 1
for k in range(K):
mem = [vecs[i] for i in range(n) if assign[i] == k]
if not mem: continue
cent[k] = norm([sum(m[d] for m in mem) / len(mem) for d in range(dim)])
if moved == 0: break
groups = {}
for i, k in enumerate(assign): groups.setdefault(k, []).append(i)
rows = []
for k, idx in groups.items():
stories = {chunks[i]['story'] for i in idx}
coh = sum(dot(vecs[i], cent[k]) for i in idx) / len(idx)
rows.append({'k': k, 'size': len(idx), 'stories': len(stories), 'coh': coh, 'idx': idx})
rows.sort(key=lambda r: (-r['stories'], -r['size']))
print(f"{'size':>5} {'stories':>8} {'coh':>6} sample titles")
print('-' * 92)
for r in rows:
titles = []
for i in r['idx']:
t = chunks[i]['title']
if t not in titles: titles.append(t)
print(f"{r['size']:5} {r['stories']:8} {r['coh']:6.3f} {', '.join(titles[:4])[:66]}")
multi = [r for r in rows if r['stories'] >= 4]
print(f"\nclusters drawing on >=4 different stories: {len(multi)} of {K}")
print(f"clusters that are essentially one story: {sum(1 for r in rows if r['stories'] <= 2)}")
pathlib.Path('clusters.json').write_text(json.dumps(
[{'k': r['k'], 'size': r['size'], 'stories': r['stories'], 'coh': r['coh'], 'idx': r['idx']} for r in rows]))
+34
View File
@@ -0,0 +1,34 @@
"""Embed chunks on the local endpoint, checkpointing as we go."""
import json, pathlib, ssl, time, urllib.request
URL = 'https://inference.lan:8443/v1/embeddings' # placeholder: the real host is never committed
MODEL = 'nomic-embed-text'
BATCH = 24
ctx = ssl.create_default_context(); ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
out = pathlib.Path('embeddings.json')
done = json.loads(out.read_text()) if out.exists() else []
start = len(done)
print(f'{len(chunks)} chunks, resuming at {start}', flush=True)
t0 = time.time()
for i in range(start, len(chunks), BATCH):
batch = [c['text'] for c in chunks[i:i + BATCH]]
body = json.dumps({'model': MODEL, 'input': batch}).encode()
req = urllib.request.Request(URL, data=body, headers={'Content-Type': 'application/json'})
for attempt in range(4):
try:
with urllib.request.urlopen(req, timeout=180, context=ctx) as r:
d = json.load(r)
done.extend(e['embedding'] for e in d['data'])
break
except Exception as e:
if attempt == 3: raise
print(f' retry {attempt+1} at {i}: {e}', flush=True)
time.sleep(3 * (attempt + 1))
out.write_text(json.dumps(done))
el = time.time() - t0
n = len(done)
print(f' {n}/{len(chunks)} {el:.0f}s ({n/max(el,1):.1f}/s)', flush=True)
print(f'done: {len(done)} embeddings in {time.time()-t0:.0f}s')
+59
View File
@@ -0,0 +1,59 @@
// How hard is it to hit the conformance bar by picking numbers at random?
// Keep the Frontier's structure and prose; replace every effect magnitude and
// gate threshold with a random plausible value; see what survives.
import { frontier } from '../../../content/frontier/index.js';
import { validatePack } from '../../../src/engine/validate.js';
import { play, sweep, strategies } from '../../../test/fixtures/strategies.js';
const clone = (o) => JSON.parse(JSON.stringify(o));
const ri = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1));
function randomize(base, { keepSigns }) {
const pack = { ...base, events: clone(base.events) };
for (const e of pack.events) {
if (e.weight !== undefined && e.weight < 5000) e.weight = ri(20, 120);
if (e.cooldown !== undefined) e.cooldown = ri(2, 12);
for (const o of e.options ?? []) {
for (const r of o.requires ?? []) if (typeof r.value === 'number') r.value = ri(5, 60);
for (const f of o.effects ?? []) {
if (f.op !== 'add' || typeof f.value !== 'number') continue;
const mag = f.path === 'money' ? ri(1, 25) : ri(1, 12);
f.value = keepSigns ? Math.sign(f.value) * mag : (Math.random() < 0.5 ? -mag : mag);
}
}
}
return pack;
}
function score(pack) {
if (!validatePack(pack).valid) return { ok: false, why: 'invalid' };
try {
const { firedEvents, offeredOptions } = sweep(pack, { seeds: 4, turns: 200 });
const allE = pack.events.map((e) => e.id);
const allO = pack.events.flatMap((e) => e.options.map((o) => `${e.id}.${o.id}`));
if (allE.some((i) => !firedEvents.has(i))) return { ok: false, why: 'dead events' };
if (allO.some((i) => !offeredOptions.has(i))) return { ok: false, why: 'dead options' };
for (const s of Object.keys(strategies)) {
let red = 0, n = 0; const counts = new Map();
play(pack, { strategy: s, seed: 'probe', turns: 150, visit: ({ state, turn }) => {
n++; if (state.money < 0) red++;
counts.set(turn.event.id, (counts.get(turn.event.id) ?? 0) + 1);
}});
if (red / n >= 0.25) return { ok: false, why: 'stuck below zero' };
const top = Math.max(...counts.values()) / n;
if (top >= 0.25) return { ok: false, why: 'one situation dominates' };
}
return { ok: true };
} catch (err) { return { ok: false, why: 'threw: ' + err.message.slice(0, 40) }; }
}
for (const keepSigns of [true, false]) {
const tally = {}; let pass = 0; const N = 60;
for (let i = 0; i < N; i++) {
const r = score(randomize(frontier, { keepSigns }));
if (r.ok) pass++; else tally[r.why] = (tally[r.why] ?? 0) + 1;
}
console.log(`\nsigns ${keepSigns ? 'preserved' : 'randomised'}: ${pass}/${N} passed`);
for (const [why, n] of Object.entries(tally).sort((a,b)=>b[1]-a[1])) console.log(' ', n, why);
}
+100
View File
@@ -0,0 +1,100 @@
// Same randomised magnitudes, but gates DERIVED from what a pack can actually
// reach rather than picked: for each gated path, simulate the ungated growth
// available and place thresholds inside that range.
import { frontier } from '../../../content/frontier/index.js';
import { validatePack } from '../../../src/engine/validate.js';
import { play, sweep, strategies } from '../../../test/fixtures/strategies.js';
const clone = (o) => JSON.parse(JSON.stringify(o));
const ri = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1));
function randomizeMagnitudes(base) {
const pack = { ...base, events: clone(base.events) };
for (const e of pack.events) {
if (e.weight !== undefined && e.weight < 5000) e.weight = ri(20, 120);
if (e.cooldown !== undefined) e.cooldown = ri(2, 12);
for (const o of e.options ?? []) {
for (const f of o.effects ?? []) {
if (f.op !== 'add' || typeof f.value !== 'number') continue;
const mag = f.path === 'money' ? ri(1, 25) : ri(1, 12);
f.value = Math.sign(f.value) * mag; // direction kept, magnitude random
}
}
}
return pack;
}
/** Ceiling actually reachable on a path, measured by playing for it. */
function reachable(pack, path) {
const gain = (o) => (o.option.effects ?? [])
.filter((e) => e.path === path && e.op === 'add').reduce((s, e) => s + e.value, 0);
let best = 0;
for (let seed = 0; seed < 3; seed++) {
const st = play(pack, { strategy: 'random', seed: `r${seed}`, turns: 120 });
void st;
let peak = 0;
play(pack, { strategy: 'random', seed: `r${seed}`, turns: 120, visit: ({ state, turn }) => {
const v = path.split('.').reduce((o, k) => o?.[k], state);
if (typeof v === 'number' && v > peak) peak = v;
void gain; void turn;
}});
best = Math.max(best, peak);
}
return best;
}
function deriveGates(pack) {
// Two passes: measure with gates removed, then place them in-range.
const stripped = { ...pack, events: clone(pack.events) };
const gatedPaths = new Set();
for (const e of stripped.events) for (const o of e.options ?? []) {
o.requires = (o.requires ?? []).filter((r) => {
const numericStat = typeof r.value === 'number' && !/^(turn|stage)/.test(r.path);
if (numericStat) gatedPaths.add(r.path);
return !numericStat;
});
if (o.requires.length === 0) delete o.requires;
}
const ceiling = {};
for (const p of gatedPaths) ceiling[p] = reachable(stripped, p);
const out = { ...pack, events: clone(pack.events) };
for (const e of out.events) for (const o of e.options ?? []) {
for (const r of o.requires ?? []) {
if (typeof r.value !== 'number' || /^(turn|stage)/.test(r.path)) continue;
const cap = ceiling[r.path] ?? 30;
// Place the gate in the reachable band: a door you can open, not a wall.
r.value = Math.max(1, Math.round(cap * (0.25 + Math.random() * 0.45)));
}
}
return out;
}
function score(pack) {
if (!validatePack(pack).valid) return { ok: false, why: 'invalid' };
try {
const { firedEvents, offeredOptions } = sweep(pack, { seeds: 4, turns: 200 });
const allE = pack.events.map((e) => e.id);
const allO = pack.events.flatMap((e) => e.options.map((o) => `${e.id}.${o.id}`));
if (allE.some((i) => !firedEvents.has(i))) return { ok: false, why: 'dead events' };
if (allO.some((i) => !offeredOptions.has(i))) return { ok: false, why: 'dead options' };
for (const s of Object.keys(strategies)) {
let red = 0, n = 0; const counts = new Map();
play(pack, { strategy: s, seed: 'probe', turns: 150, visit: ({ state, turn }) => {
n++; if (state.money < 0) red++;
counts.set(turn.event.id, (counts.get(turn.event.id) ?? 0) + 1);
}});
if (red / n >= 0.25) return { ok: false, why: 'stuck below zero' };
if (Math.max(...counts.values()) / n >= 0.25) return { ok: false, why: 'one situation dominates' };
}
return { ok: true };
} catch (err) { return { ok: false, why: 'threw: ' + err.message.slice(0, 40) }; }
}
const tally = {}; let pass = 0; const N = 40;
for (let i = 0; i < N; i++) {
const r = score(deriveGates(randomizeMagnitudes(frontier)));
if (r.ok) pass++; else tally[r.why] = (tally[r.why] ?? 0) + 1;
}
console.log(`gates derived from reachable range: ${pass}/${N} passed`);
for (const [why, n] of Object.entries(tally).sort((a,b)=>b[1]-a[1])) console.log(' ', n, why);
+111
View File
@@ -0,0 +1,111 @@
// Does a feedback loop converge? Start from randomised magnitudes, then repair
// whatever the checks complain about, one diagnosis at a time.
import { frontier } from '../../../content/frontier/index.js';
import { validatePack } from '../../../src/engine/validate.js';
import { play, sweep, strategies } from '../../../test/fixtures/strategies.js';
const clone = (o) => JSON.parse(JSON.stringify(o));
const ri = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1));
function seedPack(base) {
const pack = { ...base, events: clone(base.events), upkeep: clone(base.upkeep) };
for (const e of pack.events) {
if (e.weight !== undefined && e.weight < 5000) e.weight = ri(20, 120);
if (e.cooldown !== undefined) e.cooldown = ri(2, 12);
for (const o of e.options ?? []) {
for (const r of o.requires ?? []) {
if (typeof r.value === 'number' && !/^(turn|stage|flags)/.test(r.path)) r.value = ri(5, 60);
}
for (const f of o.effects ?? []) {
if (f.op !== 'add' || typeof f.value !== 'number') continue;
const mag = f.path === 'money' ? ri(1, 25) : ri(1, 12);
f.value = Math.sign(f.value) * mag;
}
}
}
return pack;
}
/** Full diagnosis, not just the first failure — the repair needs all of it. */
function diagnose(pack) {
const problems = [];
if (!validatePack(pack).valid) return [{ kind: 'invalid' }];
const { firedEvents, offeredOptions } = sweep(pack, { seeds: 4, turns: 200 });
for (const e of pack.events) {
if (!firedEvents.has(e.id)) problems.push({ kind: 'deadEvent', id: e.id });
for (const o of e.options ?? []) {
if (!offeredOptions.has(`${e.id}.${o.id}`)) {
problems.push({ kind: 'deadOption', event: e.id, option: o.id });
}
}
}
for (const s of Object.keys(strategies)) {
let red = 0, n = 0; const counts = new Map();
play(pack, { strategy: s, seed: 'probe', turns: 150, visit: ({ state, turn }) => {
n++; if (state.money < 0) red++;
counts.set(turn.event.id, (counts.get(turn.event.id) ?? 0) + 1);
}});
if (red / n >= 0.25) problems.push({ kind: 'broke', strategy: s, share: red / n });
const [id, c] = [...counts.entries()].sort((a, b) => b[1] - a[1])[0];
if (c / n >= 0.25) problems.push({ kind: 'dominates', id, share: c / n });
}
return problems;
}
function repair(pack, problems) {
const next = { ...pack, events: clone(pack.events), upkeep: clone(pack.upkeep) };
const event = (id) => next.events.find((e) => e.id === id);
for (const p of problems) {
if (p.kind === 'deadOption') {
const o = event(p.event)?.options.find((x) => x.id === p.option);
for (const r of o?.requires ?? []) {
if (typeof r.value === 'number' && !/^(turn|stage|flags)/.test(r.path)) {
r.value = Math.max(1, Math.floor(r.value * 0.7)); // lower the door
}
}
}
if (p.kind === 'deadEvent') {
const e = event(p.id);
if (!e) continue;
if (e.cooldown) e.cooldown = Math.max(1, Math.floor(e.cooldown * 0.7));
if (e.weight !== undefined && e.weight < 5000) e.weight = Math.round(e.weight * 1.6);
for (const r of e.requires ?? []) {
if (typeof r.value === 'number' && !/^(turn|stage|flags)/.test(r.path)) {
r.value = Math.max(1, Math.floor(r.value * 0.7));
}
}
}
if (p.kind === 'dominates') {
const e = event(p.id);
if (!e) continue;
if (e.weight !== undefined && e.weight < 5000) e.weight = Math.max(5, Math.round(e.weight * 0.6));
e.cooldown = Math.max(2, (e.cooldown ?? 1) + 2);
}
if (p.kind === 'broke') {
// Lift the floor: the settlement is the pack's own lever on the baseline.
for (const u of next.upkeep) {
if (!/Settlement|pay/i.test(u.label ?? '')) continue;
for (const f of u.effects) if (f.path === 'money') f.value = Math.round(f.value * 1.15) + 1;
}
}
}
return next;
}
let converged = 0; const N = 12; const historyLens = [];
for (let run = 0; run < N; run++) {
let pack = seedPack(frontier);
let iterations = 0;
let problems = diagnose(pack);
while (problems.length > 0 && iterations < 40) {
pack = repair(pack, problems);
problems = diagnose(pack);
iterations++;
}
if (problems.length === 0) { converged++; historyLens.push(iterations); }
}
console.log(`repair loop converged: ${converged}/${N}`);
if (historyLens.length) {
console.log('iterations to clean:', historyLens.sort((a,b)=>a-b).join(', '));
console.log('median:', historyLens[Math.floor(historyLens.length/2)]);
}
+54
View File
@@ -0,0 +1,54 @@
"""Split Gutenberg O. Henry volumes into individual stories.
Deterministic: each volume lists its stories in a CONTENTS block as
ALL-CAPS lines, then repeats each title as a heading in the body. No model
involved -- this is the segmentation step we get for free.
"""
import re, json, sys, pathlib
CAPS = re.compile(r'^[ \t]*([A-Z0-9][A-Z0-9 ,.;:\'\"\-’‘“”!?&()À-ÖØ-Þ]{3,70})[ \t]*$')
def strip_boilerplate(text):
i = text.find('*** START')
j = text.find('*** END')
if i >= 0: text = text[text.find('\n', i) + 1:] if j < 0 else text[text.find('\n', i) + 1:j]
return text
def split_volume(path):
raw = pathlib.Path(path).read_text(encoding='utf-8', errors='replace')
title_m = re.search(r'^Title:\s*(.+)$', raw, re.M)
volume = title_m.group(1).strip() if title_m else pathlib.Path(path).stem
body = strip_boilerplate(raw)
lines = body.split('\n')
caps_idx = [(n, CAPS.match(l).group(1).strip()) for n, l in enumerate(lines) if CAPS.match(l)]
if not caps_idx:
return volume, []
# The contents block is the densest early run of caps lines; every title in
# it appears again later as a heading. Use the *second* occurrence.
counts = {}
for n, t in caps_idx:
counts.setdefault(t, []).append(n)
headings = sorted((ns[-1], t) for t, ns in counts.items() if len(ns) >= 2 and len(t.split()) <= 12)
stories = []
for k, (n, t) in enumerate(headings):
end = headings[k + 1][0] if k + 1 < len(headings) else len(lines)
text = '\n'.join(lines[n + 1:end]).strip()
if len(text.split()) >= 400:
stories.append({'volume': volume, 'title': t, 'words': len(text.split()), 'text': text})
return volume, stories
out = []
for p in sorted(pathlib.Path('corpus').glob('*.txt')):
vol, st = split_volume(p)
print(f'{p.name:14} {vol[:44]:46} {len(st):3} stories')
out.extend(st)
print(f'\ntotal stories: {len(out)}')
if out:
ws = sorted(s['words'] for s in out)
print(f'words/story min {ws[0]} median {ws[len(ws)//2]} max {ws[-1]}')
print(f'median pages ~{ws[len(ws)//2]//250}')
pathlib.Path('stories.json').write_text(json.dumps(out, indent=1), encoding='utf-8')
+65
View File
@@ -0,0 +1,65 @@
"""How many clusters are BOTH situationally sharp and drawn from many stories?
Sharp = high mean cosine to centroid (a real situation, not prose register)
Spread = drawn from >= 4 different stories (recurring, not a single scene)
A pack needs ~15 per stage that are both.
"""
import json, math, pathlib, random
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
raw = json.loads(pathlib.Path('embeddings.json').read_text(encoding='utf-8'))
n = min(len(chunks), len(raw))
def norm(v):
m = math.sqrt(sum(x * x for x in v)) or 1.0
return [x / m for x in v]
V = [norm(v) for v in raw[:n]]
dim = len(V[0])
story = [chunks[i]['story'] for i in range(n)]
def dot(a, b):
s = 0.0
for x, y in zip(a, b): s += x * y
return s
def kmeans(K, seed=11, iters=12):
random.seed(seed)
cent = [V[i] for i in random.sample(range(n), K)]
assign = [-1] * n
for _ in range(iters):
moved = 0
for i, v in enumerate(V):
best, bs = 0, -2.0
for k, c in enumerate(cent):
s = dot(v, c)
if s > bs: bs, best = s, k
if assign[i] != best: assign[i] = best; moved += 1
if moved == 0: break
groups = {}
for i, k in enumerate(assign): groups.setdefault(k, []).append(i)
for k in range(K):
mem = groups.get(k)
if not mem: continue
cent[k] = norm([sum(V[i][d] for i in mem) / len(mem) for d in range(dim)])
groups = {}
for i, k in enumerate(assign): groups.setdefault(k, []).append(i)
out = []
for k, idx in groups.items():
coh = sum(dot(V[i], cent[k]) for i in idx) / len(idx)
out.append({'size': len(idx), 'stories': len({story[i] for i in idx}), 'coh': coh, 'idx': idx})
return out
print(f'{"k":>4} {"clusters":>9} {"sharp>=.83":>11} {"spread>=4":>10} {"BOTH":>6} {"median size":>12}')
print('-' * 60)
best = None
for K in (25, 40, 60, 80, 100):
cl = kmeans(K)
sharp = [c for c in cl if c['coh'] >= 0.83]
spread = [c for c in cl if c['stories'] >= 4]
both = [c for c in cl if c['coh'] >= 0.83 and c['stories'] >= 4]
sizes = sorted(c['size'] for c in cl)
print(f'{K:>4} {len(cl):>9} {len(sharp):>11} {len(spread):>10} {len(both):>6} {sizes[len(sizes)//2]:>12}')
if best is None or len(both) > len(best[1]): best = (K, both, cl)
pathlib.Path('best_k.json').write_text(json.dumps({'k': best[0],
'both': [{'size': c['size'], 'stories': c['stories'], 'coh': c['coh'], 'idx': c['idx']} for c in best[1]]}))
print(f'\nbest k = {best[0]} with {len(best[1])} clusters that are both sharp and recurring')