Files
TheLadder/tools/story-to-pack/recovered
JesseMarkowitzandClaude Opus 5 fa3769d0fe Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
2026-09-15 06:59:35 -04:00
..

story-to-pack — recovered probe scripts

Recovered 2026-09-11 from the Claude Code transcript of session d1409289-cea9-4399-a827-2f13479647bc (2026-09-10). The originals lived in that session's /tmp scratchpad and were lost when the machine rebooted. These are the last version of each file as the session wrote it. None of them has been re-run since recovery.

This is research for a tool that builds a The Ladder content pack from a corpus of stories. It is deliberately independent of theladder/ — it would live in that repo's tools/, never imported by src/, content/ or test/, or on its own.

File What it did
probe.mjs, probe2.mjs, probe3.mjs Randomised the Frontier's numbers and ran the conformance checks: random magnitudes pass 3%, random directions 0%, a 4-rule repair loop 42%
split.py Splits Gutenberg O. Henry volumes into individual stories
chunk.py Cuts stories into ~300-word scenes on paragraph boundaries
embed.py Embeds scenes with nomic-embed-text on the LAN Ollama host
cluster.py k-means over the embeddings; reports how many stories each cluster draws from
sweep.py Sweeps k and measures coherence against cross-story spread

The probes expect to run from inside theladder/ (they import its engine and packs). The Python scripts expect the corpus in a corpus/ directory beside them.

Rebuilding the corpus

O. Henry's Four Million cycle, Project Gutenberg ids 2776, 1444, 3707, 2141:

mkdir -p corpus
for id in 2776 1444 3707 2141; do
  curl -sL -o corpus/pg$id.txt "https://www.gutenberg.org/cache/epub/$id/pg$id.txt"
done

That gave 94 stories, 244,690 words, 838 scenes. Embedding all of them took about 20 minutes (0.6 chunks/sec).

Where the research stopped

  • Clustering the raw prose failed. Only 3 of 18 candidate clusters recurred across stories (2% of the corpus). At ~300 words, embedding similarity tracks which story a scene is from, not what kind of situation it is.
  • The proposed fix, untested: summarise each scene into one generic sentence with no proper names, embed the summary, then cluster.
  • It could not be tested because generation on the LAN Ollama host hung. The model ran on CPU with no GPU offload, and the 16k-context variant stayed loaded. That is probably the same fault that blocks Interactive Story M01.
  • The last open question was whether to use Claude for the generation steps until that host is fixed. Embeddings stay on Ollama, because Anthropic has no embeddings API.