# story-to-pack — design note A tool that builds a playable The Ladder content pack from a corpus of stories. **Status (2026-09-14): direction agreed, nothing built.** The decisions below were made in conversation on 2026-09-10 and had until now existed only in that session's transcript. The architecture rests on one assumption that has so far **failed** its first test (see [Assumptions](#assumptions-under-test)); nothing past the probes should be built until it passes. This is a design note, not a specification. It records what was decided and why, so that the next step is argued from it rather than re-derived. --- ## What it is for Feed it a body of fiction set in one world — say 20–30 short stories — plus a handful of structural numbers, and get back a content pack the game can load and that its own conformance checks judge viable. Then play it, hand-edit it, and retune it from play logs. The Ladder's founding commitment (Planning/01 §9) is that a new setting is purely data. This tool tests how far that goes: whether "data" can come from a corpus rather than a designer. ## Decisions | Area | Decision | | --- | --- | | **Inputs** | A thematically coherent corpus (one world), 20–30 stories of 2–15 pages. Plus: number of stages, turns per stage, number of promotions, number of characters, and a one-line player role written by the designer. Early development uses much less text. | | **Corpus** | Saved permanently beside the pack. Read once, at fixed depth. **Provenance recorded:** which chunks fed which event, even though nothing reads it yet. | | **Artifact** | **One.** The generated pack is the source of truth from the moment it is written. The designer and the tool both edit that same file in place. Seed text is a seed, not a source: there is no regenerate-and-merge. | | **Pipeline** | Chunk → embed → cluster → select recurring situations → extract prose → assign state paths and directions → search numbers → verify → repair. | | **The model's job** | Selection and classification. **It never chooses a number.** Its generative tasks are entity substitution and, only where needed, filling out options. | | **Inference** | Once, up front, checkpointed and cached. The numeric search and every repair iteration are pure computation and never call a model. This is what makes a small local model viable. | | **Stat vocabulary** | A fixed mechanical spine (currency, debt, standing, two skills, flags, notable, per-character relationship stats). The corpus supplies the *labels*, not the structure. | | **Cast** | Recurring **roles** harvested across the corpus (the authority, the rival, the peer), not people from any one story. Characters are minted for the roles and entity names are substituted in extracted prose. | | **Stages** | Assigned automatically by the protagonist's standing in each scene, binned and balanced across stages. Overridable by hand. | | **Event count** | Derived from turns per stage, with a sufficiency report *before* inference is spent: "this corpus yields N situations, the pack needs M". | | **Options** | 3–5 per event: 1–3 fixed archetypes ("say nothing"), 1–3 drawn from other scenes in the cluster, and 1–3 generated only to fill gaps. Deduplicated **mechanically** — two options with the same effects are one option. | | **Description variants** | An event can carry several descriptions, sharing one general option set. This needs an engine change (see Open questions). | | **Viability** | Fully automatic — a human does not have to play a pack for it to count as viable. A **graded score**, not pass/fail, so extra time improves it. The search **always keeps the best-scoring candidate**, never the most recent. When it cannot reach viable, it reports precisely which checks it could not satisfy. | | **Time and quality** | Start fast and iterate; later, overnight runs for higher quality. Time buys **numeric search depth and prose fidelity** — not corpus reach. | | **Retuning** | Deterministic, driven by the play log `npm run analyze` already reads. Not model-mediated. | | **Extension** | Pulling more corpus into an existing pack is **out of scope** for now. Provenance is recorded so it stays possible. | | **Home** | `theladder/tools/`, never imported by `src/`, `content/` or `test/`. Its own dependencies stay out of the game. | | **Hardware** | Must be able to run on a low-powered local inference host, taking longer if it has to. | ## Why the hard part is numbers, not prose Counted from the two hand-written packs: | | Corporate Ladder | The Frontier | | --- | --- | --- | | Prose strings | 219 | 211 | | Numeric decisions | 347 | 427 | | Stable ids | 125 | 120 | A pack is roughly 35% prose and 65% numbers, and every defect the project has hit — the debt trap, the rent retune, the negotiation cliff — lived in the numbers. The corpus supplies the prose; the tool has to invent the rest, and has to be able to show it is not a trap. Randomising The Frontier's numbers while keeping its prose and structure: random magnitudes pass the conformance checks 3% of the time, random directions 0%, and a crude four-rule repair loop 42% (5 of 12 runs converged). The search space is not hopeless, but it is not a broad basin either — hence a real search against a graded score. ## Why recurrence is the central problem Hand-written packs are almost entirely **recurring** situations: 29 of 31 events in Corporate Ladder and 28 of 29 in The Frontier can fire more than once, and a 41-turn playtest saw 12 distinct situations an average of 3.4 times each. Stories are the opposite: every scene happens once. So selection is not "pick 30 good scenes". It is **find the situations this corpus keeps returning to**. Twenty stories in one world should contain eight variations on "a superior asks you to cover something up"; clustering those is what yields one recurring event with several descriptions. One long novel is the hard case and is not the target. ## Boundaries - **Planning/01 §10 is not violated by a hosted model here.** §10 says the *game* must never require a hosted inference service. The generator is a build-time tool whose output is plain data with no runtime dependency; a shipped `.html` still plays offline. Record this in theladder's DECISIONS when the tool lands there. - **There are two provider slots, not one.** Embeddings and generation are configured separately. Anthropic offers generation but no embeddings API, so any mix must keep embeddings local. - **The inference host is shared.** Other work sends requests to the same server. Clients run one request at a time with a pause between requests, set a short `keep_alive`, and read the host from the environment (`STP_OLLAMA`) — never from a tracked file. Every run is agreed with its owner before it starts. ## Assumptions under test The architecture rests on three assumptions. All three are cheap to test, and the first could invalidate everything else. ### 1. A coherent corpus clusters into recurring situations — FAILED in naive form **Test (2026-09-10):** O. Henry's *Four Million* cycle — four Gutenberg volumes, 94 stories, 244,690 words — cut into 838 scenes of ~300 words, embedded with `nomic-embed-text`, clustered with k-means. **Result:** at the best k, 18 clusters were both tight and drawn from four or more stories. But reading them showed most were 60–79% **one** story plus stragglers. Measured by concentration, only **3 of 18** were genuinely recurring, covering **2% of the corpus**. At scene length, embedding similarity tracks *which story a scene is from* — names, setting, narrator voice — not *what kind of situation it is*. **Fix under test:** summarise each scene into one generic sentence with no proper names, embed the summary, and cluster those instead. This probe is in `probe/`. **Baseline re-measured (2026-09-14)** on the rebuilt corpus — prose embeddings made on a GPU host, same model, same 838 chunks. Seed 11 at k = 100 reproduces the 2026-09-10 figures exactly (18 sharp clusters spanning ≥ 4 stories, 3 genuinely recurring, 2% coverage). Seeds 12 and 13 give 7 and 7, covering 8% and 10%: the old 2% was one unlucky initialisation, and the true baseline is poor rather than hopeless. **Pass mark, fixed before any summary existed.** Judged at **k = 60 and k = 100**, medians over three seeds, summary embeddings against prose embeddings as reported by `probe/measure.py`: 1. Median dominant-story share of clusters of size ≥ 5 falls to **40% or below** at both k. 2. **Tight-recurring coverage** — recurring clusters (≥ 4 stories, dominant ≤ 40%) whose coherence is at or above that file's median cluster — reaches **at least 10% of the corpus** at both k. (A relative bar was written first and dropped once the baseline came in at 2% and 0%, where "double" means nothing. 10% of 838 scenes is ~84 — enough raw material for a stage's worth of situations.) 3. Reading three sampled tight-recurring clusters, each is recognisably **one kind of situation**, not a prose register ("men walking through the city") or a vocabulary cluster. The prose baseline numbers the first two are judged against are in `probe/measure-prose.log`. **Result (2026-09-14): criteria 1 and 2 pass decisively; criterion 3 passes only in part.** | k | Median dominant share, prose → summaries | Tight-recurring coverage, prose → summaries | | --- | --- | --- | | 60 | 57% → **22%** | 2% → **26%** | | 100 | 67% → **28%** | 0% → **19%** | Pipeline as run: `qwen2.5:3b-instruct` summaries with the v2 prompt (no example roles), names removed by `strip_names.py`, embedded with `nomic-embed-text`. Logs: `probe/measure-summary.log`, `probe/sample-k60.log`, `probe/sample-k100.log`. Reading the seeded samples (three at each k; one cluster appeared at both, so five distinct): - **Clearly one situation (2):** a spouse scheming against or confronting a spouse over infidelity (15 scenes, 10 stories); a woman made the centre of attention while men compete for her (5 scenes, 5 stories). - **Broadly one situation (2):** a secret identity or past revealed in a romantic encounter, often over dinner (15 scenes, 13 stories); hosts and guests jockeying for social rank at a dinner (6, 6). - **Not a situation (1):** fortune-tellers, adventurers, a found tyre, a graft inquiry — held together by "seeks", "fortune" and "mysterious" and dominated by three stories. It is the same cluster at both k. The rule as written requires every sample to read as one situation, so **strictly this is not a full pass**, and it is recorded that way rather than rounded up. **Follow-up: the selection step is designed and tested in [`SELECTION.md`](SELECTION.md).** A two-signal gate (cross-story cohesion, and no single word dominating) accepted 10 clusters across development and test, all of them situations; it rejects the vocabulary clusters this result found. But it keeps only 5 of ~40 candidates — far short of what a pack needs — so recall is the open problem. Three recall-recovery methods were then tested on a fresh partition and **none met its pre-registered bar**: re-clustering found nothing new, consensus clustering found only sharp situations (3 of 3) but too few, and prune-then-grow found 10 distinct situations with too much noise. The finding that matters is that the gate is size-dependent in both directions, so it has to be made size-invariant before recall can be recovered. A second gate, G2, built for that (cohesion against a same-size random baseline, a significance test for shared words), **failed validation on fresh data**: it accepts three to four times as many clusters as the first gate at about the same precision (0.79), but still rejects every cluster under 10 members and still lets growth run to the cap. A fixed threshold on a z-score reverses the size bias rather than removing it. A third gate, G3 (a percentile against same-size random sets, plus a mandatory per-member floor), **failed all four validation criteria and did worse than G2**: 17 accepted across fresh partitions, only 8 of them situations. The random-set baseline turned out to be the wrong comparison — any k-means cluster beats random scenes — and the per-member floor is noisiest, and so most lenient, exactly where clusters are smallest. A fourth gate, G4, had the inference model read each cluster and say which summaries share one situation. It **failed with both `qwen2.5:3b-instruct` and `qwen3:14b`** (primary precision 0.48 and 0.52 against G2's 0.73): the models answer with abstractions ("power dynamics", "someone wants something from someone") that most summaries fit. Observed but not pre-registered: the 14B model AND G2 accepted 10 clusters, all situations — precise, narrow, and untested on fresh data. The substantive conclusion: summarising before embedding **does** break the story-identity effect that sank the naive design, and most of what clusters is recognisably situational — the recurrence premise is alive, not proven. The design needs a **selection step that rejects vocabulary clusters**, not only a clustering step. Known defects in this run, none large enough to change the verdict: - `first_sentence` cut 7 of 838 summaries that begin with an initial ("E. Rushmore Coglan…") to "E." — single-letter initials and "Gen." are not handled. - Heavy use of "someone" is itself a shared token across summaries with many characters. Whether it pulls clusters together has not been measured; a replacement by role or a neutral mix would test it. - A few names that are also common words survive stripping ("Silver", "Cherry", "Grace", "the Kid"). **Why k ≥ 60, and why a tightness filter.** The first version of this pass mark said "median dominant ≤ 40% at k ≥ 40" — and the prose baseline, the approach known to fail, passed it at k = 40 (36%). Two things were wrong. Small k forces clusters to merge stories whatever the embeddings encode: on synthetic vectors carrying nothing but story identity, `measure.py` reports 36% at k = 25, 53% at k = 40 and 100% at k ≥ 60, while vectors with planted cross-story situations report 11–17% at every k ≥ 40. And without a tightness filter, the loose clusters the 2026-09-10 reading identified as prose register count as recurring. The tightness cutoff is relative to each file's own median because an absolute cosine threshold does not transfer between 300-word scenes and one-sentence summaries. If it fails, the recurrence premise needs rethinking before any design goes further. ### 2. A small model can classify against a closed path vocabulary Roughly 100+ calls per pack deciding which state paths an option moves and in which direction. Malformed or wrong output at a meaningful rate changes the design (batching, retries, a bigger model). **Untested.** ### 3. A small model can substitute entity names without wrecking prose The main generative task, run on every description. **Untested.** The summarisation probe in assumption 1 incidentally gives the first evidence on 2 and 3: it asks a small model for constrained, name-free output and the results can be read. **First evidence (2026-09-14, `qwen2.5:3b-instruct`, two 40-scene pilots):** - **It copies examples from the prompt into the output.** A prompt that illustrated roles with "a clerk", "a landlady", "a rich man", "a city park" got those phrases back 18 times — in scenes containing **none** of them, every one invented. For this probe that is worse than noise: shared invented phrases make unrelated scenes cluster together and fake exactly the recurrence being measured. Removing the examples brought it to 1. **Rule: prompts for a small model name no example of the output vocabulary.** - **It does not follow "never use a name."** 25 of 40 summaries kept names with the first prompt, and 27 of 40 with a stricter one plus an automatic retry — the retry fired 35 times, nearly doubled the run, and fixed almost nothing. Names are therefore removed **deterministically** afterwards (`probe/strip_names.py`): a lexicon built from the corpus, of words seen capitalised mid-sentence and almost never in lowercase. It misses names that are also common words ("Bill", "the Kid", "New" in "New York"). That bears directly on assumption 3: **entity substitution should be done by code over a lexicon, with the model's output treated as untrusted.** - **Throughput is not the constraint.** On a GTX 1080 Ti, a scene summary takes ~0.6 s of model time (about 1.1 s including a deliberate pause between requests), and embedding runs at ~16 texts/s. A 838-scene corpus summarises in roughly a quarter of an hour. ## Open questions - **Description variants need an engine change.** Events carry one description today. Scope it before building, and record it in theladder's DECISIONS — it is the first engine change content generation would force. - **The test strategies hardcode the stat spine.** `theladder/test/fixtures/strategies.js` reads the literal path `money` and the prefix `relationships.`. A pack that named its currency `coin` would make four of six archetypes stop discriminating while the sweep still ran and reported. The fixed spine satisfies this by construction; it is also a coupling that should be made explicit. - **Options versus variants.** A cluster-derived option is tied to the scene it came from and may read wrongly under the event's other descriptions. The more variants an event has, the more its options must come from the fixed archetypes. - **More options, harder search.** Moving from ~3.1 to ~4 options per event adds roughly 90 effect magnitudes to a search that already converged only 5 times in 12 with crude repair. - **Prompts may not transfer between models.** Prompts developed against a capable model are typically under-specified for a 3B one. The small-model question stays open until it is tested on a small model, however well a larger one does. - **Story splitting is per-volume fiddly.** Segmentation needs no model, but each source's format needs its own pattern. A production tool needs a report when a volume does not split cleanly. The probe's `split.py` yields 94 stories from 98 contents entries: it silently drops titles containing quotation marks or an em dash ("LITTLE SPECK IN GARNERED FRUIT", "THE THING'S THE PLAY", "WHAT YOU WANT", "THE GUILTY PARTY — AN EAST SIDE TRAGEDY"). The 2026-09-10 run dropped the same ones, so it is left alone to keep the before/after comparison exact. ## Layout ``` story-to-pack/ DESIGN.md this note recovered/ the 2026-09-10 scripts, as recovered from the transcript; not re-run probe/ the working probe for assumption 1 corpus/ Gutenberg source texts (rebuildable; see recovered/README.md) split.py volumes -> stories.json chunk.py stories -> chunks.json (~300-word scenes) ollama.py the only module that talks to an inference host embed.py texts -> vectors, checkpointed summarise.py scenes -> one generic sentence each, checkpointed strip_names.py summaries -> summaries with names replaced, from a corpus-built lexicon measure.py clustering concentration report for any embedding file sample_clusters.py seeded sample of tight-recurring clusters, for reading (criterion 3) summaries-v1-pilot.json, summaries-v2-pilot-retry.json the two 40-scene pilots, kept as evidence *.log run and measurement output gpu-logs/ the inference host's power/link/kernel logs for the 2026-09-14 run ```