Research toward building a content pack from a story corpus, kept on its own branch and independent of the game. Records the selection experiments against blind labels, and settles selection as gate G2 followed by a human review: review.py writes REVIEW.md and a review.json form, apply_review.py checks the filled form and writes situations.json for the next stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
276 lines
19 KiB
Markdown
276 lines
19 KiB
Markdown
# story-to-pack — design note
|
||
|
||
A tool that builds a playable The Ladder content pack from a corpus of stories.
|
||
|
||
**Status (2026-09-14): direction agreed, nothing built.** The decisions below were made in
|
||
conversation on 2026-09-10 and had until now existed only in that session's transcript. The
|
||
architecture rests on one assumption that has so far **failed** its first test (see
|
||
[Assumptions](#assumptions-under-test)); nothing past the probes should be built until it passes.
|
||
|
||
This is a design note, not a specification. It records what was decided and why, so that the next
|
||
step is argued from it rather than re-derived.
|
||
|
||
---
|
||
|
||
## What it is for
|
||
|
||
Feed it a body of fiction set in one world — say 20–30 short stories — plus a handful of structural
|
||
numbers, and get back a content pack the game can load and that its own conformance checks judge
|
||
viable. Then play it, hand-edit it, and retune it from play logs.
|
||
|
||
The Ladder's founding commitment (Planning/01 §9) is that a new setting is purely data. This tool
|
||
tests how far that goes: whether "data" can come from a corpus rather than a designer.
|
||
|
||
## Decisions
|
||
|
||
| Area | Decision |
|
||
| --- | --- |
|
||
| **Inputs** | A thematically coherent corpus (one world), 20–30 stories of 2–15 pages. Plus: number of stages, turns per stage, number of promotions, number of characters, and a one-line player role written by the designer. Early development uses much less text. |
|
||
| **Corpus** | Saved permanently beside the pack. Read once, at fixed depth. **Provenance recorded:** which chunks fed which event, even though nothing reads it yet. |
|
||
| **Artifact** | **One.** The generated pack is the source of truth from the moment it is written. The designer and the tool both edit that same file in place. Seed text is a seed, not a source: there is no regenerate-and-merge. |
|
||
| **Pipeline** | Chunk → embed → cluster → select recurring situations → extract prose → assign state paths and directions → search numbers → verify → repair. |
|
||
| **The model's job** | Selection and classification. **It never chooses a number.** Its generative tasks are entity substitution and, only where needed, filling out options. |
|
||
| **Inference** | Once, up front, checkpointed and cached. The numeric search and every repair iteration are pure computation and never call a model. This is what makes a small local model viable. |
|
||
| **Stat vocabulary** | A fixed mechanical spine (currency, debt, standing, two skills, flags, notable, per-character relationship stats). The corpus supplies the *labels*, not the structure. |
|
||
| **Cast** | Recurring **roles** harvested across the corpus (the authority, the rival, the peer), not people from any one story. Characters are minted for the roles and entity names are substituted in extracted prose. |
|
||
| **Stages** | Assigned automatically by the protagonist's standing in each scene, binned and balanced across stages. Overridable by hand. |
|
||
| **Event count** | Derived from turns per stage, with a sufficiency report *before* inference is spent: "this corpus yields N situations, the pack needs M". |
|
||
| **Options** | 3–5 per event: 1–3 fixed archetypes ("say nothing"), 1–3 drawn from other scenes in the cluster, and 1–3 generated only to fill gaps. Deduplicated **mechanically** — two options with the same effects are one option. |
|
||
| **Description variants** | An event can carry several descriptions, sharing one general option set. This needs an engine change (see Open questions). |
|
||
| **Viability** | Fully automatic — a human does not have to play a pack for it to count as viable. A **graded score**, not pass/fail, so extra time improves it. The search **always keeps the best-scoring candidate**, never the most recent. When it cannot reach viable, it reports precisely which checks it could not satisfy. |
|
||
| **Time and quality** | Start fast and iterate; later, overnight runs for higher quality. Time buys **numeric search depth and prose fidelity** — not corpus reach. |
|
||
| **Retuning** | Deterministic, driven by the play log `npm run analyze` already reads. Not model-mediated. |
|
||
| **Extension** | Pulling more corpus into an existing pack is **out of scope** for now. Provenance is recorded so it stays possible. |
|
||
| **Home** | `theladder/tools/`, never imported by `src/`, `content/` or `test/`. Its own dependencies stay out of the game. |
|
||
| **Hardware** | Must be able to run on a low-powered local inference host, taking longer if it has to. |
|
||
|
||
## Why the hard part is numbers, not prose
|
||
|
||
Counted from the two hand-written packs:
|
||
|
||
| | Corporate Ladder | The Frontier |
|
||
| --- | --- | --- |
|
||
| Prose strings | 219 | 211 |
|
||
| Numeric decisions | 347 | 427 |
|
||
| Stable ids | 125 | 120 |
|
||
|
||
A pack is roughly 35% prose and 65% numbers, and every defect the project has hit — the debt trap,
|
||
the rent retune, the negotiation cliff — lived in the numbers. The corpus supplies the prose; the
|
||
tool has to invent the rest, and has to be able to show it is not a trap.
|
||
|
||
Randomising The Frontier's numbers while keeping its prose and structure: random magnitudes pass the
|
||
conformance checks 3% of the time, random directions 0%, and a crude four-rule repair loop 42% (5 of
|
||
12 runs converged). The search space is not hopeless, but it is not a broad basin either — hence a
|
||
real search against a graded score.
|
||
|
||
## Why recurrence is the central problem
|
||
|
||
Hand-written packs are almost entirely **recurring** situations: 29 of 31 events in Corporate Ladder
|
||
and 28 of 29 in The Frontier can fire more than once, and a 41-turn playtest saw 12 distinct
|
||
situations an average of 3.4 times each. Stories are the opposite: every scene happens once.
|
||
|
||
So selection is not "pick 30 good scenes". It is **find the situations this corpus keeps returning
|
||
to**. Twenty stories in one world should contain eight variations on "a superior asks you to cover
|
||
something up"; clustering those is what yields one recurring event with several descriptions. One
|
||
long novel is the hard case and is not the target.
|
||
|
||
## Boundaries
|
||
|
||
- **Planning/01 §10 is not violated by a hosted model here.** §10 says the *game* must never require
|
||
a hosted inference service. The generator is a build-time tool whose output is plain data with no
|
||
runtime dependency; a shipped `.html` still plays offline. Record this in theladder's DECISIONS
|
||
when the tool lands there.
|
||
- **There are two provider slots, not one.** Embeddings and generation are configured separately.
|
||
Anthropic offers generation but no embeddings API, so any mix must keep embeddings local.
|
||
- **The inference host is shared.** Other work sends requests to the same server. Clients run one
|
||
request at a time with a pause between requests, set a short `keep_alive`, and read the host from
|
||
the environment (`STP_OLLAMA`) — never from a tracked file. Every run is agreed with its owner
|
||
before it starts.
|
||
|
||
## Assumptions under test
|
||
|
||
The architecture rests on three assumptions. All three are cheap to test, and the first could
|
||
invalidate everything else.
|
||
|
||
### 1. A coherent corpus clusters into recurring situations — FAILED in naive form
|
||
|
||
**Test (2026-09-10):** O. Henry's *Four Million* cycle — four Gutenberg volumes, 94 stories, 244,690
|
||
words — cut into 838 scenes of ~300 words, embedded with `nomic-embed-text`, clustered with k-means.
|
||
|
||
**Result:** at the best k, 18 clusters were both tight and drawn from four or more stories. But
|
||
reading them showed most were 60–79% **one** story plus stragglers. Measured by concentration, only
|
||
**3 of 18** were genuinely recurring, covering **2% of the corpus**. At scene length, embedding
|
||
similarity tracks *which story a scene is from* — names, setting, narrator voice — not *what kind of
|
||
situation it is*.
|
||
|
||
**Fix under test:** summarise each scene into one generic sentence with no proper names, embed the
|
||
summary, and cluster those instead. This probe is in `probe/`.
|
||
|
||
**Baseline re-measured (2026-09-14)** on the rebuilt corpus — prose embeddings made on a GPU host,
|
||
same model, same 838 chunks. Seed 11 at k = 100 reproduces the 2026-09-10 figures exactly (18 sharp
|
||
clusters spanning ≥ 4 stories, 3 genuinely recurring, 2% coverage). Seeds 12 and 13 give 7 and 7,
|
||
covering 8% and 10%: the old 2% was one unlucky initialisation, and the true baseline is poor rather
|
||
than hopeless.
|
||
|
||
**Pass mark, fixed before any summary existed.** Judged at **k = 60 and k = 100**, medians over
|
||
three seeds, summary embeddings against prose embeddings as reported by `probe/measure.py`:
|
||
|
||
1. Median dominant-story share of clusters of size ≥ 5 falls to **40% or below** at both k.
|
||
2. **Tight-recurring coverage** — recurring clusters (≥ 4 stories, dominant ≤ 40%) whose coherence is
|
||
at or above that file's median cluster — reaches **at least 10% of the corpus** at both k. (A
|
||
relative bar was written first and dropped once the baseline came in at 2% and 0%, where "double"
|
||
means nothing. 10% of 838 scenes is ~84 — enough raw material for a stage's worth of situations.)
|
||
3. Reading three sampled tight-recurring clusters, each is recognisably **one kind of situation**,
|
||
not a prose register ("men walking through the city") or a vocabulary cluster.
|
||
|
||
The prose baseline numbers the first two are judged against are in `probe/measure-prose.log`.
|
||
|
||
**Result (2026-09-14): criteria 1 and 2 pass decisively; criterion 3 passes only in part.**
|
||
|
||
| k | Median dominant share, prose → summaries | Tight-recurring coverage, prose → summaries |
|
||
| --- | --- | --- |
|
||
| 60 | 57% → **22%** | 2% → **26%** |
|
||
| 100 | 67% → **28%** | 0% → **19%** |
|
||
|
||
Pipeline as run: `qwen2.5:3b-instruct` summaries with the v2 prompt (no example roles), names removed
|
||
by `strip_names.py`, embedded with `nomic-embed-text`. Logs: `probe/measure-summary.log`,
|
||
`probe/sample-k60.log`, `probe/sample-k100.log`.
|
||
|
||
Reading the seeded samples (three at each k; one cluster appeared at both, so five distinct):
|
||
|
||
- **Clearly one situation (2):** a spouse scheming against or confronting a spouse over infidelity
|
||
(15 scenes, 10 stories); a woman made the centre of attention while men compete for her (5 scenes,
|
||
5 stories).
|
||
- **Broadly one situation (2):** a secret identity or past revealed in a romantic encounter, often over
|
||
dinner (15 scenes, 13 stories); hosts and guests jockeying for social rank at a dinner (6, 6).
|
||
- **Not a situation (1):** fortune-tellers, adventurers, a found tyre, a graft inquiry — held together
|
||
by "seeks", "fortune" and "mysterious" and dominated by three stories. It is the same cluster at
|
||
both k.
|
||
|
||
The rule as written requires every sample to read as one situation, so **strictly this is not a full
|
||
pass**, and it is recorded that way rather than rounded up.
|
||
|
||
**Follow-up: the selection step is designed and tested in [`SELECTION.md`](SELECTION.md).** A
|
||
two-signal gate (cross-story cohesion, and no single word dominating) accepted 10 clusters across
|
||
development and test, all of them situations; it rejects the vocabulary clusters this result found.
|
||
But it keeps only 5 of ~40 candidates — far short of what a pack needs — so recall is the open problem.
|
||
Three recall-recovery methods were then tested on a fresh partition and **none met its pre-registered
|
||
bar**: re-clustering found nothing new, consensus clustering found only sharp situations (3 of 3) but
|
||
too few, and prune-then-grow found 10 distinct situations with too much noise. The finding that
|
||
matters is that the gate is size-dependent in both directions, so it has to be made size-invariant
|
||
before recall can be recovered.
|
||
A second gate, G2, built for that (cohesion against a same-size random baseline, a significance test
|
||
for shared words), **failed validation on fresh data**: it accepts three to four times as many
|
||
clusters as the first gate at about the same precision (0.79), but still rejects every cluster under
|
||
10 members and still lets growth run to the cap. A fixed threshold on a z-score reverses the size bias
|
||
rather than removing it.
|
||
A third gate, G3 (a percentile against same-size random sets, plus a mandatory per-member floor),
|
||
**failed all four validation criteria and did worse than G2**: 17 accepted across fresh partitions,
|
||
only 8 of them situations. The random-set baseline turned out to be the wrong comparison — any
|
||
k-means cluster beats random scenes — and the per-member floor is noisiest, and so most lenient,
|
||
exactly where clusters are smallest.
|
||
A fourth gate, G4, had the inference model read each cluster and say which summaries share one
|
||
situation. It **failed with both `qwen2.5:3b-instruct` and `qwen3:14b`** (primary precision 0.48 and
|
||
0.52 against G2's 0.73): the models answer with abstractions ("power dynamics", "someone wants
|
||
something from someone") that most summaries fit. Observed but not pre-registered: the 14B model AND
|
||
G2 accepted 10 clusters, all situations — precise, narrow, and untested on fresh data. The substantive conclusion: summarising
|
||
before embedding **does** break the story-identity effect that sank the naive design, and most of
|
||
what clusters is recognisably situational — the recurrence premise is alive, not proven. The design
|
||
needs a **selection step that rejects vocabulary clusters**, not only a clustering step.
|
||
|
||
Known defects in this run, none large enough to change the verdict:
|
||
|
||
- `first_sentence` cut 7 of 838 summaries that begin with an initial ("E. Rushmore Coglan…") to
|
||
"E." — single-letter initials and "Gen." are not handled.
|
||
- Heavy use of "someone" is itself a shared token across summaries with many characters. Whether it
|
||
pulls clusters together has not been measured; a replacement by role or a neutral mix would test it.
|
||
- A few names that are also common words survive stripping ("Silver", "Cherry", "Grace", "the Kid").
|
||
|
||
**Why k ≥ 60, and why a tightness filter.** The first version of this pass mark said "median dominant
|
||
≤ 40% at k ≥ 40" — and the prose baseline, the approach known to fail, passed it at k = 40 (36%).
|
||
Two things were wrong. Small k forces clusters to merge stories whatever the embeddings encode: on
|
||
synthetic vectors carrying nothing but story identity, `measure.py` reports 36% at k = 25, 53% at k
|
||
= 40 and 100% at k ≥ 60, while vectors with planted cross-story situations report 11–17% at every k
|
||
≥ 40. And without a tightness filter, the loose clusters the 2026-09-10 reading identified as prose
|
||
register count as recurring. The tightness cutoff is relative to each file's own median because an
|
||
absolute cosine threshold does not transfer between 300-word scenes and one-sentence summaries.
|
||
|
||
If it fails, the recurrence premise needs rethinking before any design goes further.
|
||
|
||
### 2. A small model can classify against a closed path vocabulary
|
||
|
||
Roughly 100+ calls per pack deciding which state paths an option moves and in which direction.
|
||
Malformed or wrong output at a meaningful rate changes the design (batching, retries, a bigger
|
||
model). **Untested.**
|
||
|
||
### 3. A small model can substitute entity names without wrecking prose
|
||
|
||
The main generative task, run on every description. **Untested.**
|
||
|
||
The summarisation probe in assumption 1 incidentally gives the first evidence on 2 and 3: it asks a
|
||
small model for constrained, name-free output and the results can be read.
|
||
|
||
**First evidence (2026-09-14, `qwen2.5:3b-instruct`, two 40-scene pilots):**
|
||
|
||
- **It copies examples from the prompt into the output.** A prompt that illustrated roles with "a
|
||
clerk", "a landlady", "a rich man", "a city park" got those phrases back 18 times — in scenes
|
||
containing **none** of them, every one invented. For this probe that is worse than noise: shared
|
||
invented phrases make unrelated scenes cluster together and fake exactly the recurrence being
|
||
measured. Removing the examples brought it to 1. **Rule: prompts for a small model name no example
|
||
of the output vocabulary.**
|
||
- **It does not follow "never use a name."** 25 of 40 summaries kept names with the first prompt,
|
||
and 27 of 40 with a stricter one plus an automatic retry — the retry fired 35 times, nearly doubled
|
||
the run, and fixed almost nothing. Names are therefore removed **deterministically** afterwards
|
||
(`probe/strip_names.py`): a lexicon built from the corpus, of words seen capitalised mid-sentence
|
||
and almost never in lowercase. It misses names that are also common words ("Bill", "the Kid",
|
||
"New" in "New York"). That bears directly on assumption 3: **entity substitution should be done by
|
||
code over a lexicon, with the model's output treated as untrusted.**
|
||
- **Throughput is not the constraint.** On a GTX 1080 Ti, a scene summary takes ~0.6 s of model time
|
||
(about 1.1 s including a deliberate pause between requests), and embedding runs at ~16 texts/s.
|
||
A 838-scene corpus summarises in roughly a quarter of an hour.
|
||
|
||
## Open questions
|
||
|
||
- **Description variants need an engine change.** Events carry one description today. Scope it
|
||
before building, and record it in theladder's DECISIONS — it is the first engine change content
|
||
generation would force.
|
||
- **The test strategies hardcode the stat spine.** `theladder/test/fixtures/strategies.js` reads the
|
||
literal path `money` and the prefix `relationships.`. A pack that named its currency `coin` would
|
||
make four of six archetypes stop discriminating while the sweep still ran and reported. The fixed
|
||
spine satisfies this by construction; it is also a coupling that should be made explicit.
|
||
- **Options versus variants.** A cluster-derived option is tied to the scene it came from and may
|
||
read wrongly under the event's other descriptions. The more variants an event has, the more its
|
||
options must come from the fixed archetypes.
|
||
- **More options, harder search.** Moving from ~3.1 to ~4 options per event adds roughly 90 effect
|
||
magnitudes to a search that already converged only 5 times in 12 with crude repair.
|
||
- **Prompts may not transfer between models.** Prompts developed against a capable model are
|
||
typically under-specified for a 3B one. The small-model question stays open until it is tested on
|
||
a small model, however well a larger one does.
|
||
- **Story splitting is per-volume fiddly.** Segmentation needs no model, but each source's format
|
||
needs its own pattern. A production tool needs a report when a volume does not split cleanly.
|
||
The probe's `split.py` yields 94 stories from 98 contents entries: it silently drops titles
|
||
containing quotation marks or an em dash ("LITTLE SPECK IN GARNERED FRUIT", "THE THING'S THE PLAY",
|
||
"WHAT YOU WANT", "THE GUILTY PARTY — AN EAST SIDE TRAGEDY"). The 2026-09-10 run dropped the same
|
||
ones, so it is left alone to keep the before/after comparison exact.
|
||
|
||
## Layout
|
||
|
||
```
|
||
story-to-pack/
|
||
DESIGN.md this note
|
||
recovered/ the 2026-09-10 scripts, as recovered from the transcript; not re-run
|
||
probe/ the working probe for assumption 1
|
||
corpus/ Gutenberg source texts (rebuildable; see recovered/README.md)
|
||
split.py volumes -> stories.json
|
||
chunk.py stories -> chunks.json (~300-word scenes)
|
||
ollama.py the only module that talks to an inference host
|
||
embed.py texts -> vectors, checkpointed
|
||
summarise.py scenes -> one generic sentence each, checkpointed
|
||
strip_names.py summaries -> summaries with names replaced, from a corpus-built lexicon
|
||
measure.py clustering concentration report for any embedding file
|
||
sample_clusters.py seeded sample of tight-recurring clusters, for reading (criterion 3)
|
||
summaries-v1-pilot.json, summaries-v2-pilot-retry.json the two 40-scene pilots, kept as evidence
|
||
*.log run and measurement output
|
||
gpu-logs/ the inference host's power/link/kernel logs for the 2026-09-14 run
|
||
```
|