# story-to-pack — recovered probe scripts Recovered 2026-09-11 from the Claude Code transcript of session `d1409289-cea9-4399-a827-2f13479647bc` (2026-09-10). The originals lived in that session's `/tmp` scratchpad and were lost when the machine rebooted. These are the last version of each file as the session wrote it. None of them has been re-run since recovery. This is research for a tool that builds a The Ladder content pack from a corpus of stories. It is deliberately independent of `theladder/` — it would live in that repo's `tools/`, never imported by `src/`, `content/` or `test/`, or on its own. | File | What it did | | --- | --- | | `probe.mjs`, `probe2.mjs`, `probe3.mjs` | Randomised the Frontier's numbers and ran the conformance checks: random magnitudes pass 3%, random directions 0%, a 4-rule repair loop 42% | | `split.py` | Splits Gutenberg O. Henry volumes into individual stories | | `chunk.py` | Cuts stories into ~300-word scenes on paragraph boundaries | | `embed.py` | Embeds scenes with `nomic-embed-text` on the LAN Ollama host | | `cluster.py` | k-means over the embeddings; reports how many stories each cluster draws from | | `sweep.py` | Sweeps k and measures coherence against cross-story spread | The probes expect to run from inside `theladder/` (they import its engine and packs). The Python scripts expect the corpus in a `corpus/` directory beside them. ## Rebuilding the corpus O. Henry's *Four Million* cycle, Project Gutenberg ids 2776, 1444, 3707, 2141: ```sh mkdir -p corpus for id in 2776 1444 3707 2141; do curl -sL -o corpus/pg$id.txt "https://www.gutenberg.org/cache/epub/$id/pg$id.txt" done ``` That gave 94 stories, 244,690 words, 838 scenes. Embedding all of them took about 20 minutes (0.6 chunks/sec). ## Where the research stopped - Clustering the raw prose failed. Only 3 of 18 candidate clusters recurred across stories (2% of the corpus). At ~300 words, embedding similarity tracks which story a scene is from, not what kind of situation it is. - The proposed fix, untested: summarise each scene into one generic sentence with no proper names, embed the summary, then cluster. - It could not be tested because generation on the LAN Ollama host hung. The model ran on CPU with no GPU offload, and the 16k-context variant stayed loaded. That is probably the same fault that blocks Interactive Story M01. - The last open question was whether to use Claude for the generation steps until that host is fixed. Embeddings stay on Ollama, because Anthropic has no embeddings API.