Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own branch and independent of the game. Records the selection experiments against blind labels, and settles selection as gate G2 followed by a human review: review.py writes REVIEW.md and a review.json form, apply_review.py checks the filled form and writes situations.json for the next stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
This commit is contained in:
co-authored by
Claude Opus 5
parent
d84ae495f4
commit
fa3769d0fe
@@ -0,0 +1,53 @@
|
||||
# story-to-pack — recovered probe scripts
|
||||
|
||||
Recovered 2026-09-11 from the Claude Code transcript of session
|
||||
`d1409289-cea9-4399-a827-2f13479647bc` (2026-09-10). The originals lived in that
|
||||
session's `/tmp` scratchpad and were lost when the machine rebooted. These are
|
||||
the last version of each file as the session wrote it. None of them has been
|
||||
re-run since recovery.
|
||||
|
||||
This is research for a tool that builds a The Ladder content pack from a corpus
|
||||
of stories. It is deliberately independent of `theladder/` — it would live in
|
||||
that repo's `tools/`, never imported by `src/`, `content/` or `test/`, or on its
|
||||
own.
|
||||
|
||||
| File | What it did |
|
||||
| --- | --- |
|
||||
| `probe.mjs`, `probe2.mjs`, `probe3.mjs` | Randomised the Frontier's numbers and ran the conformance checks: random magnitudes pass 3%, random directions 0%, a 4-rule repair loop 42% |
|
||||
| `split.py` | Splits Gutenberg O. Henry volumes into individual stories |
|
||||
| `chunk.py` | Cuts stories into ~300-word scenes on paragraph boundaries |
|
||||
| `embed.py` | Embeds scenes with `nomic-embed-text` on the LAN Ollama host |
|
||||
| `cluster.py` | k-means over the embeddings; reports how many stories each cluster draws from |
|
||||
| `sweep.py` | Sweeps k and measures coherence against cross-story spread |
|
||||
|
||||
The probes expect to run from inside `theladder/` (they import its engine and
|
||||
packs). The Python scripts expect the corpus in a `corpus/` directory beside
|
||||
them.
|
||||
|
||||
## Rebuilding the corpus
|
||||
|
||||
O. Henry's *Four Million* cycle, Project Gutenberg ids 2776, 1444, 3707, 2141:
|
||||
|
||||
```sh
|
||||
mkdir -p corpus
|
||||
for id in 2776 1444 3707 2141; do
|
||||
curl -sL -o corpus/pg$id.txt "https://www.gutenberg.org/cache/epub/$id/pg$id.txt"
|
||||
done
|
||||
```
|
||||
|
||||
That gave 94 stories, 244,690 words, 838 scenes. Embedding all of them took
|
||||
about 20 minutes (0.6 chunks/sec).
|
||||
|
||||
## Where the research stopped
|
||||
|
||||
- Clustering the raw prose failed. Only 3 of 18 candidate clusters recurred
|
||||
across stories (2% of the corpus). At ~300 words, embedding similarity tracks
|
||||
which story a scene is from, not what kind of situation it is.
|
||||
- The proposed fix, untested: summarise each scene into one generic sentence
|
||||
with no proper names, embed the summary, then cluster.
|
||||
- It could not be tested because generation on the LAN Ollama host hung. The
|
||||
model ran on CPU with no GPU offload, and the 16k-context variant stayed
|
||||
loaded. That is probably the same fault that blocks Interactive Story M01.
|
||||
- The last open question was whether to use Claude for the generation steps
|
||||
until that host is fixed. Embeddings stay on Ollama, because Anthropic has no
|
||||
embeddings API.
|
||||
@@ -0,0 +1,26 @@
|
||||
"""Cut stories into scene-sized chunks on paragraph boundaries."""
|
||||
import json, pathlib, re
|
||||
|
||||
TARGET = 320 # words; a scene is a situation with people in it
|
||||
|
||||
stories = json.loads(pathlib.Path('stories.json').read_text(encoding='utf-8'))
|
||||
chunks = []
|
||||
for si, s in enumerate(stories):
|
||||
paras = [p.strip() for p in re.split(r'\n\s*\n', s['text']) if p.strip()]
|
||||
buf, n = [], 0
|
||||
for p in paras:
|
||||
w = len(p.split())
|
||||
if n + w > TARGET and buf:
|
||||
chunks.append({'story': si, 'title': s['title'], 'volume': s['volume'],
|
||||
'text': ' '.join(buf), 'words': n})
|
||||
buf, n = [], 0
|
||||
buf.append(p); n += w
|
||||
if buf and n >= 80:
|
||||
chunks.append({'story': si, 'title': s['title'], 'volume': s['volume'],
|
||||
'text': ' '.join(buf), 'words': n})
|
||||
|
||||
pathlib.Path('chunks.json').write_text(json.dumps(chunks), encoding='utf-8')
|
||||
ws = sorted(c['words'] for c in chunks)
|
||||
print(f'chunks: {len(chunks)} from {len(stories)} stories')
|
||||
print(f'words/chunk min {ws[0]} median {ws[len(ws)//2]} max {ws[-1]}')
|
||||
print(f'total words: {sum(ws):,}')
|
||||
@@ -0,0 +1,80 @@
|
||||
"""Cluster scene embeddings and report what the clusters actually contain.
|
||||
|
||||
The question this answers: does a thematically coherent corpus contain
|
||||
RECURRING situations, or 800 singletons? A pack needs ~15 repeatable
|
||||
situations per stage; a cluster that draws from many different stories is a
|
||||
recurring situation, one that draws from a single story is just a scene.
|
||||
"""
|
||||
import json, math, pathlib, random, sys
|
||||
|
||||
K = int(sys.argv[1]) if len(sys.argv) > 1 else 40
|
||||
SEED = 11
|
||||
|
||||
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
|
||||
vecs = json.loads(pathlib.Path('embeddings.json').read_text(encoding='utf-8'))
|
||||
n = min(len(chunks), len(vecs))
|
||||
chunks, vecs = chunks[:n], vecs[:n]
|
||||
print(f'clustering {n} chunks into k={K}\n')
|
||||
|
||||
def norm(v):
|
||||
m = math.sqrt(sum(x * x for x in v)) or 1.0
|
||||
return [x / m for x in v]
|
||||
vecs = [norm(v) for v in vecs]
|
||||
dim = len(vecs[0])
|
||||
|
||||
def dot(a, b): return sum(x * y for x, y in zip(a, b))
|
||||
|
||||
# k-means++ init on cosine distance (vectors are unit length, so dot == cos)
|
||||
random.seed(SEED)
|
||||
cent = [vecs[random.randrange(n)]]
|
||||
while len(cent) < K:
|
||||
d2 = [min((1 - dot(v, c)) for c in cent) ** 2 for v in vecs]
|
||||
tot = sum(d2) or 1.0
|
||||
r = random.random() * tot
|
||||
acc = 0.0
|
||||
for i, x in enumerate(d2):
|
||||
acc += x
|
||||
if acc >= r:
|
||||
cent.append(vecs[i]); break
|
||||
else:
|
||||
cent.append(vecs[random.randrange(n)])
|
||||
|
||||
assign = [0] * n
|
||||
for it in range(30):
|
||||
moved = 0
|
||||
for i, v in enumerate(vecs):
|
||||
best, bs = 0, -2.0
|
||||
for k, c in enumerate(cent):
|
||||
s = dot(v, c)
|
||||
if s > bs: bs, best = s, k
|
||||
if assign[i] != best: assign[i] = best; moved += 1
|
||||
for k in range(K):
|
||||
mem = [vecs[i] for i in range(n) if assign[i] == k]
|
||||
if not mem: continue
|
||||
cent[k] = norm([sum(m[d] for m in mem) / len(mem) for d in range(dim)])
|
||||
if moved == 0: break
|
||||
|
||||
groups = {}
|
||||
for i, k in enumerate(assign): groups.setdefault(k, []).append(i)
|
||||
|
||||
rows = []
|
||||
for k, idx in groups.items():
|
||||
stories = {chunks[i]['story'] for i in idx}
|
||||
coh = sum(dot(vecs[i], cent[k]) for i in idx) / len(idx)
|
||||
rows.append({'k': k, 'size': len(idx), 'stories': len(stories), 'coh': coh, 'idx': idx})
|
||||
rows.sort(key=lambda r: (-r['stories'], -r['size']))
|
||||
|
||||
print(f"{'size':>5} {'stories':>8} {'coh':>6} sample titles")
|
||||
print('-' * 92)
|
||||
for r in rows:
|
||||
titles = []
|
||||
for i in r['idx']:
|
||||
t = chunks[i]['title']
|
||||
if t not in titles: titles.append(t)
|
||||
print(f"{r['size']:5} {r['stories']:8} {r['coh']:6.3f} {', '.join(titles[:4])[:66]}")
|
||||
|
||||
multi = [r for r in rows if r['stories'] >= 4]
|
||||
print(f"\nclusters drawing on >=4 different stories: {len(multi)} of {K}")
|
||||
print(f"clusters that are essentially one story: {sum(1 for r in rows if r['stories'] <= 2)}")
|
||||
pathlib.Path('clusters.json').write_text(json.dumps(
|
||||
[{'k': r['k'], 'size': r['size'], 'stories': r['stories'], 'coh': r['coh'], 'idx': r['idx']} for r in rows]))
|
||||
@@ -0,0 +1,34 @@
|
||||
"""Embed chunks on the local endpoint, checkpointing as we go."""
|
||||
import json, pathlib, ssl, time, urllib.request
|
||||
|
||||
URL = 'https://inference.lan:8443/v1/embeddings' # placeholder: the real host is never committed
|
||||
MODEL = 'nomic-embed-text'
|
||||
BATCH = 24
|
||||
ctx = ssl.create_default_context(); ctx.check_hostname = False; ctx.verify_mode = ssl.CERT_NONE
|
||||
|
||||
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
|
||||
out = pathlib.Path('embeddings.json')
|
||||
done = json.loads(out.read_text()) if out.exists() else []
|
||||
|
||||
start = len(done)
|
||||
print(f'{len(chunks)} chunks, resuming at {start}', flush=True)
|
||||
t0 = time.time()
|
||||
for i in range(start, len(chunks), BATCH):
|
||||
batch = [c['text'] for c in chunks[i:i + BATCH]]
|
||||
body = json.dumps({'model': MODEL, 'input': batch}).encode()
|
||||
req = urllib.request.Request(URL, data=body, headers={'Content-Type': 'application/json'})
|
||||
for attempt in range(4):
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=180, context=ctx) as r:
|
||||
d = json.load(r)
|
||||
done.extend(e['embedding'] for e in d['data'])
|
||||
break
|
||||
except Exception as e:
|
||||
if attempt == 3: raise
|
||||
print(f' retry {attempt+1} at {i}: {e}', flush=True)
|
||||
time.sleep(3 * (attempt + 1))
|
||||
out.write_text(json.dumps(done))
|
||||
el = time.time() - t0
|
||||
n = len(done)
|
||||
print(f' {n}/{len(chunks)} {el:.0f}s ({n/max(el,1):.1f}/s)', flush=True)
|
||||
print(f'done: {len(done)} embeddings in {time.time()-t0:.0f}s')
|
||||
@@ -0,0 +1,59 @@
|
||||
// How hard is it to hit the conformance bar by picking numbers at random?
|
||||
// Keep the Frontier's structure and prose; replace every effect magnitude and
|
||||
// gate threshold with a random plausible value; see what survives.
|
||||
import { frontier } from '../../../content/frontier/index.js';
|
||||
import { validatePack } from '../../../src/engine/validate.js';
|
||||
import { play, sweep, strategies } from '../../../test/fixtures/strategies.js';
|
||||
|
||||
const clone = (o) => JSON.parse(JSON.stringify(o));
|
||||
const ri = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1));
|
||||
|
||||
function randomize(base, { keepSigns }) {
|
||||
const pack = { ...base, events: clone(base.events) };
|
||||
for (const e of pack.events) {
|
||||
if (e.weight !== undefined && e.weight < 5000) e.weight = ri(20, 120);
|
||||
if (e.cooldown !== undefined) e.cooldown = ri(2, 12);
|
||||
for (const o of e.options ?? []) {
|
||||
for (const r of o.requires ?? []) if (typeof r.value === 'number') r.value = ri(5, 60);
|
||||
for (const f of o.effects ?? []) {
|
||||
if (f.op !== 'add' || typeof f.value !== 'number') continue;
|
||||
const mag = f.path === 'money' ? ri(1, 25) : ri(1, 12);
|
||||
f.value = keepSigns ? Math.sign(f.value) * mag : (Math.random() < 0.5 ? -mag : mag);
|
||||
}
|
||||
}
|
||||
}
|
||||
return pack;
|
||||
}
|
||||
|
||||
function score(pack) {
|
||||
if (!validatePack(pack).valid) return { ok: false, why: 'invalid' };
|
||||
try {
|
||||
const { firedEvents, offeredOptions } = sweep(pack, { seeds: 4, turns: 200 });
|
||||
const allE = pack.events.map((e) => e.id);
|
||||
const allO = pack.events.flatMap((e) => e.options.map((o) => `${e.id}.${o.id}`));
|
||||
if (allE.some((i) => !firedEvents.has(i))) return { ok: false, why: 'dead events' };
|
||||
if (allO.some((i) => !offeredOptions.has(i))) return { ok: false, why: 'dead options' };
|
||||
|
||||
for (const s of Object.keys(strategies)) {
|
||||
let red = 0, n = 0; const counts = new Map();
|
||||
play(pack, { strategy: s, seed: 'probe', turns: 150, visit: ({ state, turn }) => {
|
||||
n++; if (state.money < 0) red++;
|
||||
counts.set(turn.event.id, (counts.get(turn.event.id) ?? 0) + 1);
|
||||
}});
|
||||
if (red / n >= 0.25) return { ok: false, why: 'stuck below zero' };
|
||||
const top = Math.max(...counts.values()) / n;
|
||||
if (top >= 0.25) return { ok: false, why: 'one situation dominates' };
|
||||
}
|
||||
return { ok: true };
|
||||
} catch (err) { return { ok: false, why: 'threw: ' + err.message.slice(0, 40) }; }
|
||||
}
|
||||
|
||||
for (const keepSigns of [true, false]) {
|
||||
const tally = {}; let pass = 0; const N = 60;
|
||||
for (let i = 0; i < N; i++) {
|
||||
const r = score(randomize(frontier, { keepSigns }));
|
||||
if (r.ok) pass++; else tally[r.why] = (tally[r.why] ?? 0) + 1;
|
||||
}
|
||||
console.log(`\nsigns ${keepSigns ? 'preserved' : 'randomised'}: ${pass}/${N} passed`);
|
||||
for (const [why, n] of Object.entries(tally).sort((a,b)=>b[1]-a[1])) console.log(' ', n, why);
|
||||
}
|
||||
@@ -0,0 +1,100 @@
|
||||
// Same randomised magnitudes, but gates DERIVED from what a pack can actually
|
||||
// reach rather than picked: for each gated path, simulate the ungated growth
|
||||
// available and place thresholds inside that range.
|
||||
import { frontier } from '../../../content/frontier/index.js';
|
||||
import { validatePack } from '../../../src/engine/validate.js';
|
||||
import { play, sweep, strategies } from '../../../test/fixtures/strategies.js';
|
||||
|
||||
const clone = (o) => JSON.parse(JSON.stringify(o));
|
||||
const ri = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1));
|
||||
|
||||
function randomizeMagnitudes(base) {
|
||||
const pack = { ...base, events: clone(base.events) };
|
||||
for (const e of pack.events) {
|
||||
if (e.weight !== undefined && e.weight < 5000) e.weight = ri(20, 120);
|
||||
if (e.cooldown !== undefined) e.cooldown = ri(2, 12);
|
||||
for (const o of e.options ?? []) {
|
||||
for (const f of o.effects ?? []) {
|
||||
if (f.op !== 'add' || typeof f.value !== 'number') continue;
|
||||
const mag = f.path === 'money' ? ri(1, 25) : ri(1, 12);
|
||||
f.value = Math.sign(f.value) * mag; // direction kept, magnitude random
|
||||
}
|
||||
}
|
||||
}
|
||||
return pack;
|
||||
}
|
||||
|
||||
/** Ceiling actually reachable on a path, measured by playing for it. */
|
||||
function reachable(pack, path) {
|
||||
const gain = (o) => (o.option.effects ?? [])
|
||||
.filter((e) => e.path === path && e.op === 'add').reduce((s, e) => s + e.value, 0);
|
||||
let best = 0;
|
||||
for (let seed = 0; seed < 3; seed++) {
|
||||
const st = play(pack, { strategy: 'random', seed: `r${seed}`, turns: 120 });
|
||||
void st;
|
||||
let peak = 0;
|
||||
play(pack, { strategy: 'random', seed: `r${seed}`, turns: 120, visit: ({ state, turn }) => {
|
||||
const v = path.split('.').reduce((o, k) => o?.[k], state);
|
||||
if (typeof v === 'number' && v > peak) peak = v;
|
||||
void gain; void turn;
|
||||
}});
|
||||
best = Math.max(best, peak);
|
||||
}
|
||||
return best;
|
||||
}
|
||||
|
||||
function deriveGates(pack) {
|
||||
// Two passes: measure with gates removed, then place them in-range.
|
||||
const stripped = { ...pack, events: clone(pack.events) };
|
||||
const gatedPaths = new Set();
|
||||
for (const e of stripped.events) for (const o of e.options ?? []) {
|
||||
o.requires = (o.requires ?? []).filter((r) => {
|
||||
const numericStat = typeof r.value === 'number' && !/^(turn|stage)/.test(r.path);
|
||||
if (numericStat) gatedPaths.add(r.path);
|
||||
return !numericStat;
|
||||
});
|
||||
if (o.requires.length === 0) delete o.requires;
|
||||
}
|
||||
const ceiling = {};
|
||||
for (const p of gatedPaths) ceiling[p] = reachable(stripped, p);
|
||||
|
||||
const out = { ...pack, events: clone(pack.events) };
|
||||
for (const e of out.events) for (const o of e.options ?? []) {
|
||||
for (const r of o.requires ?? []) {
|
||||
if (typeof r.value !== 'number' || /^(turn|stage)/.test(r.path)) continue;
|
||||
const cap = ceiling[r.path] ?? 30;
|
||||
// Place the gate in the reachable band: a door you can open, not a wall.
|
||||
r.value = Math.max(1, Math.round(cap * (0.25 + Math.random() * 0.45)));
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function score(pack) {
|
||||
if (!validatePack(pack).valid) return { ok: false, why: 'invalid' };
|
||||
try {
|
||||
const { firedEvents, offeredOptions } = sweep(pack, { seeds: 4, turns: 200 });
|
||||
const allE = pack.events.map((e) => e.id);
|
||||
const allO = pack.events.flatMap((e) => e.options.map((o) => `${e.id}.${o.id}`));
|
||||
if (allE.some((i) => !firedEvents.has(i))) return { ok: false, why: 'dead events' };
|
||||
if (allO.some((i) => !offeredOptions.has(i))) return { ok: false, why: 'dead options' };
|
||||
for (const s of Object.keys(strategies)) {
|
||||
let red = 0, n = 0; const counts = new Map();
|
||||
play(pack, { strategy: s, seed: 'probe', turns: 150, visit: ({ state, turn }) => {
|
||||
n++; if (state.money < 0) red++;
|
||||
counts.set(turn.event.id, (counts.get(turn.event.id) ?? 0) + 1);
|
||||
}});
|
||||
if (red / n >= 0.25) return { ok: false, why: 'stuck below zero' };
|
||||
if (Math.max(...counts.values()) / n >= 0.25) return { ok: false, why: 'one situation dominates' };
|
||||
}
|
||||
return { ok: true };
|
||||
} catch (err) { return { ok: false, why: 'threw: ' + err.message.slice(0, 40) }; }
|
||||
}
|
||||
|
||||
const tally = {}; let pass = 0; const N = 40;
|
||||
for (let i = 0; i < N; i++) {
|
||||
const r = score(deriveGates(randomizeMagnitudes(frontier)));
|
||||
if (r.ok) pass++; else tally[r.why] = (tally[r.why] ?? 0) + 1;
|
||||
}
|
||||
console.log(`gates derived from reachable range: ${pass}/${N} passed`);
|
||||
for (const [why, n] of Object.entries(tally).sort((a,b)=>b[1]-a[1])) console.log(' ', n, why);
|
||||
@@ -0,0 +1,111 @@
|
||||
// Does a feedback loop converge? Start from randomised magnitudes, then repair
|
||||
// whatever the checks complain about, one diagnosis at a time.
|
||||
import { frontier } from '../../../content/frontier/index.js';
|
||||
import { validatePack } from '../../../src/engine/validate.js';
|
||||
import { play, sweep, strategies } from '../../../test/fixtures/strategies.js';
|
||||
|
||||
const clone = (o) => JSON.parse(JSON.stringify(o));
|
||||
const ri = (lo, hi) => lo + Math.floor(Math.random() * (hi - lo + 1));
|
||||
|
||||
function seedPack(base) {
|
||||
const pack = { ...base, events: clone(base.events), upkeep: clone(base.upkeep) };
|
||||
for (const e of pack.events) {
|
||||
if (e.weight !== undefined && e.weight < 5000) e.weight = ri(20, 120);
|
||||
if (e.cooldown !== undefined) e.cooldown = ri(2, 12);
|
||||
for (const o of e.options ?? []) {
|
||||
for (const r of o.requires ?? []) {
|
||||
if (typeof r.value === 'number' && !/^(turn|stage|flags)/.test(r.path)) r.value = ri(5, 60);
|
||||
}
|
||||
for (const f of o.effects ?? []) {
|
||||
if (f.op !== 'add' || typeof f.value !== 'number') continue;
|
||||
const mag = f.path === 'money' ? ri(1, 25) : ri(1, 12);
|
||||
f.value = Math.sign(f.value) * mag;
|
||||
}
|
||||
}
|
||||
}
|
||||
return pack;
|
||||
}
|
||||
|
||||
/** Full diagnosis, not just the first failure — the repair needs all of it. */
|
||||
function diagnose(pack) {
|
||||
const problems = [];
|
||||
if (!validatePack(pack).valid) return [{ kind: 'invalid' }];
|
||||
const { firedEvents, offeredOptions } = sweep(pack, { seeds: 4, turns: 200 });
|
||||
for (const e of pack.events) {
|
||||
if (!firedEvents.has(e.id)) problems.push({ kind: 'deadEvent', id: e.id });
|
||||
for (const o of e.options ?? []) {
|
||||
if (!offeredOptions.has(`${e.id}.${o.id}`)) {
|
||||
problems.push({ kind: 'deadOption', event: e.id, option: o.id });
|
||||
}
|
||||
}
|
||||
}
|
||||
for (const s of Object.keys(strategies)) {
|
||||
let red = 0, n = 0; const counts = new Map();
|
||||
play(pack, { strategy: s, seed: 'probe', turns: 150, visit: ({ state, turn }) => {
|
||||
n++; if (state.money < 0) red++;
|
||||
counts.set(turn.event.id, (counts.get(turn.event.id) ?? 0) + 1);
|
||||
}});
|
||||
if (red / n >= 0.25) problems.push({ kind: 'broke', strategy: s, share: red / n });
|
||||
const [id, c] = [...counts.entries()].sort((a, b) => b[1] - a[1])[0];
|
||||
if (c / n >= 0.25) problems.push({ kind: 'dominates', id, share: c / n });
|
||||
}
|
||||
return problems;
|
||||
}
|
||||
|
||||
function repair(pack, problems) {
|
||||
const next = { ...pack, events: clone(pack.events), upkeep: clone(pack.upkeep) };
|
||||
const event = (id) => next.events.find((e) => e.id === id);
|
||||
for (const p of problems) {
|
||||
if (p.kind === 'deadOption') {
|
||||
const o = event(p.event)?.options.find((x) => x.id === p.option);
|
||||
for (const r of o?.requires ?? []) {
|
||||
if (typeof r.value === 'number' && !/^(turn|stage|flags)/.test(r.path)) {
|
||||
r.value = Math.max(1, Math.floor(r.value * 0.7)); // lower the door
|
||||
}
|
||||
}
|
||||
}
|
||||
if (p.kind === 'deadEvent') {
|
||||
const e = event(p.id);
|
||||
if (!e) continue;
|
||||
if (e.cooldown) e.cooldown = Math.max(1, Math.floor(e.cooldown * 0.7));
|
||||
if (e.weight !== undefined && e.weight < 5000) e.weight = Math.round(e.weight * 1.6);
|
||||
for (const r of e.requires ?? []) {
|
||||
if (typeof r.value === 'number' && !/^(turn|stage|flags)/.test(r.path)) {
|
||||
r.value = Math.max(1, Math.floor(r.value * 0.7));
|
||||
}
|
||||
}
|
||||
}
|
||||
if (p.kind === 'dominates') {
|
||||
const e = event(p.id);
|
||||
if (!e) continue;
|
||||
if (e.weight !== undefined && e.weight < 5000) e.weight = Math.max(5, Math.round(e.weight * 0.6));
|
||||
e.cooldown = Math.max(2, (e.cooldown ?? 1) + 2);
|
||||
}
|
||||
if (p.kind === 'broke') {
|
||||
// Lift the floor: the settlement is the pack's own lever on the baseline.
|
||||
for (const u of next.upkeep) {
|
||||
if (!/Settlement|pay/i.test(u.label ?? '')) continue;
|
||||
for (const f of u.effects) if (f.path === 'money') f.value = Math.round(f.value * 1.15) + 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
return next;
|
||||
}
|
||||
|
||||
let converged = 0; const N = 12; const historyLens = [];
|
||||
for (let run = 0; run < N; run++) {
|
||||
let pack = seedPack(frontier);
|
||||
let iterations = 0;
|
||||
let problems = diagnose(pack);
|
||||
while (problems.length > 0 && iterations < 40) {
|
||||
pack = repair(pack, problems);
|
||||
problems = diagnose(pack);
|
||||
iterations++;
|
||||
}
|
||||
if (problems.length === 0) { converged++; historyLens.push(iterations); }
|
||||
}
|
||||
console.log(`repair loop converged: ${converged}/${N}`);
|
||||
if (historyLens.length) {
|
||||
console.log('iterations to clean:', historyLens.sort((a,b)=>a-b).join(', '));
|
||||
console.log('median:', historyLens[Math.floor(historyLens.length/2)]);
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
"""Split Gutenberg O. Henry volumes into individual stories.
|
||||
|
||||
Deterministic: each volume lists its stories in a CONTENTS block as
|
||||
ALL-CAPS lines, then repeats each title as a heading in the body. No model
|
||||
involved -- this is the segmentation step we get for free.
|
||||
"""
|
||||
import re, json, sys, pathlib
|
||||
|
||||
CAPS = re.compile(r'^[ \t]*([A-Z0-9][A-Z0-9 ,.;:\'\"\-’‘“”!?&()À-ÖØ-Þ]{3,70})[ \t]*$')
|
||||
|
||||
def strip_boilerplate(text):
|
||||
i = text.find('*** START')
|
||||
j = text.find('*** END')
|
||||
if i >= 0: text = text[text.find('\n', i) + 1:] if j < 0 else text[text.find('\n', i) + 1:j]
|
||||
return text
|
||||
|
||||
def split_volume(path):
|
||||
raw = pathlib.Path(path).read_text(encoding='utf-8', errors='replace')
|
||||
title_m = re.search(r'^Title:\s*(.+)$', raw, re.M)
|
||||
volume = title_m.group(1).strip() if title_m else pathlib.Path(path).stem
|
||||
body = strip_boilerplate(raw)
|
||||
|
||||
lines = body.split('\n')
|
||||
caps_idx = [(n, CAPS.match(l).group(1).strip()) for n, l in enumerate(lines) if CAPS.match(l)]
|
||||
if not caps_idx:
|
||||
return volume, []
|
||||
|
||||
# The contents block is the densest early run of caps lines; every title in
|
||||
# it appears again later as a heading. Use the *second* occurrence.
|
||||
counts = {}
|
||||
for n, t in caps_idx:
|
||||
counts.setdefault(t, []).append(n)
|
||||
headings = sorted((ns[-1], t) for t, ns in counts.items() if len(ns) >= 2 and len(t.split()) <= 12)
|
||||
|
||||
stories = []
|
||||
for k, (n, t) in enumerate(headings):
|
||||
end = headings[k + 1][0] if k + 1 < len(headings) else len(lines)
|
||||
text = '\n'.join(lines[n + 1:end]).strip()
|
||||
if len(text.split()) >= 400:
|
||||
stories.append({'volume': volume, 'title': t, 'words': len(text.split()), 'text': text})
|
||||
return volume, stories
|
||||
|
||||
out = []
|
||||
for p in sorted(pathlib.Path('corpus').glob('*.txt')):
|
||||
vol, st = split_volume(p)
|
||||
print(f'{p.name:14} {vol[:44]:46} {len(st):3} stories')
|
||||
out.extend(st)
|
||||
|
||||
print(f'\ntotal stories: {len(out)}')
|
||||
if out:
|
||||
ws = sorted(s['words'] for s in out)
|
||||
print(f'words/story min {ws[0]} median {ws[len(ws)//2]} max {ws[-1]}')
|
||||
print(f'median pages ~{ws[len(ws)//2]//250}')
|
||||
pathlib.Path('stories.json').write_text(json.dumps(out, indent=1), encoding='utf-8')
|
||||
@@ -0,0 +1,65 @@
|
||||
"""How many clusters are BOTH situationally sharp and drawn from many stories?
|
||||
|
||||
Sharp = high mean cosine to centroid (a real situation, not prose register)
|
||||
Spread = drawn from >= 4 different stories (recurring, not a single scene)
|
||||
A pack needs ~15 per stage that are both.
|
||||
"""
|
||||
import json, math, pathlib, random
|
||||
|
||||
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
|
||||
raw = json.loads(pathlib.Path('embeddings.json').read_text(encoding='utf-8'))
|
||||
n = min(len(chunks), len(raw))
|
||||
|
||||
def norm(v):
|
||||
m = math.sqrt(sum(x * x for x in v)) or 1.0
|
||||
return [x / m for x in v]
|
||||
V = [norm(v) for v in raw[:n]]
|
||||
dim = len(V[0])
|
||||
story = [chunks[i]['story'] for i in range(n)]
|
||||
|
||||
def dot(a, b):
|
||||
s = 0.0
|
||||
for x, y in zip(a, b): s += x * y
|
||||
return s
|
||||
|
||||
def kmeans(K, seed=11, iters=12):
|
||||
random.seed(seed)
|
||||
cent = [V[i] for i in random.sample(range(n), K)]
|
||||
assign = [-1] * n
|
||||
for _ in range(iters):
|
||||
moved = 0
|
||||
for i, v in enumerate(V):
|
||||
best, bs = 0, -2.0
|
||||
for k, c in enumerate(cent):
|
||||
s = dot(v, c)
|
||||
if s > bs: bs, best = s, k
|
||||
if assign[i] != best: assign[i] = best; moved += 1
|
||||
if moved == 0: break
|
||||
groups = {}
|
||||
for i, k in enumerate(assign): groups.setdefault(k, []).append(i)
|
||||
for k in range(K):
|
||||
mem = groups.get(k)
|
||||
if not mem: continue
|
||||
cent[k] = norm([sum(V[i][d] for i in mem) / len(mem) for d in range(dim)])
|
||||
groups = {}
|
||||
for i, k in enumerate(assign): groups.setdefault(k, []).append(i)
|
||||
out = []
|
||||
for k, idx in groups.items():
|
||||
coh = sum(dot(V[i], cent[k]) for i in idx) / len(idx)
|
||||
out.append({'size': len(idx), 'stories': len({story[i] for i in idx}), 'coh': coh, 'idx': idx})
|
||||
return out
|
||||
|
||||
print(f'{"k":>4} {"clusters":>9} {"sharp>=.83":>11} {"spread>=4":>10} {"BOTH":>6} {"median size":>12}')
|
||||
print('-' * 60)
|
||||
best = None
|
||||
for K in (25, 40, 60, 80, 100):
|
||||
cl = kmeans(K)
|
||||
sharp = [c for c in cl if c['coh'] >= 0.83]
|
||||
spread = [c for c in cl if c['stories'] >= 4]
|
||||
both = [c for c in cl if c['coh'] >= 0.83 and c['stories'] >= 4]
|
||||
sizes = sorted(c['size'] for c in cl)
|
||||
print(f'{K:>4} {len(cl):>9} {len(sharp):>11} {len(spread):>10} {len(both):>6} {sizes[len(sizes)//2]:>12}')
|
||||
if best is None or len(both) > len(best[1]): best = (K, both, cl)
|
||||
pathlib.Path('best_k.json').write_text(json.dumps({'k': best[0],
|
||||
'both': [{'size': c['size'], 'stories': c['stories'], 'coh': c['coh'], 'idx': c['idx']} for c in best[1]]}))
|
||||
print(f'\nbest k = {best[0]} with {len(best[1])} clusters that are both sharp and recurring')
|
||||
Reference in New Issue
Block a user