Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own branch and independent of the game. Records the selection experiments against blind labels, and settles selection as gate G2 followed by a human review: review.py writes REVIEW.md and a review.json form, apply_review.py checks the filled form and writes situations.json for the next stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
This commit is contained in:
co-authored by
Claude Opus 5
parent
d84ae495f4
commit
fa3769d0fe
@@ -0,0 +1,53 @@
|
||||
# story-to-pack — recovered probe scripts
|
||||
|
||||
Recovered 2026-09-11 from the Claude Code transcript of session
|
||||
`d1409289-cea9-4399-a827-2f13479647bc` (2026-09-10). The originals lived in that
|
||||
session's `/tmp` scratchpad and were lost when the machine rebooted. These are
|
||||
the last version of each file as the session wrote it. None of them has been
|
||||
re-run since recovery.
|
||||
|
||||
This is research for a tool that builds a The Ladder content pack from a corpus
|
||||
of stories. It is deliberately independent of `theladder/` — it would live in
|
||||
that repo's `tools/`, never imported by `src/`, `content/` or `test/`, or on its
|
||||
own.
|
||||
|
||||
| File | What it did |
|
||||
| --- | --- |
|
||||
| `probe.mjs`, `probe2.mjs`, `probe3.mjs` | Randomised the Frontier's numbers and ran the conformance checks: random magnitudes pass 3%, random directions 0%, a 4-rule repair loop 42% |
|
||||
| `split.py` | Splits Gutenberg O. Henry volumes into individual stories |
|
||||
| `chunk.py` | Cuts stories into ~300-word scenes on paragraph boundaries |
|
||||
| `embed.py` | Embeds scenes with `nomic-embed-text` on the LAN Ollama host |
|
||||
| `cluster.py` | k-means over the embeddings; reports how many stories each cluster draws from |
|
||||
| `sweep.py` | Sweeps k and measures coherence against cross-story spread |
|
||||
|
||||
The probes expect to run from inside `theladder/` (they import its engine and
|
||||
packs). The Python scripts expect the corpus in a `corpus/` directory beside
|
||||
them.
|
||||
|
||||
## Rebuilding the corpus
|
||||
|
||||
O. Henry's *Four Million* cycle, Project Gutenberg ids 2776, 1444, 3707, 2141:
|
||||
|
||||
```sh
|
||||
mkdir -p corpus
|
||||
for id in 2776 1444 3707 2141; do
|
||||
curl -sL -o corpus/pg$id.txt "https://www.gutenberg.org/cache/epub/$id/pg$id.txt"
|
||||
done
|
||||
```
|
||||
|
||||
That gave 94 stories, 244,690 words, 838 scenes. Embedding all of them took
|
||||
about 20 minutes (0.6 chunks/sec).
|
||||
|
||||
## Where the research stopped
|
||||
|
||||
- Clustering the raw prose failed. Only 3 of 18 candidate clusters recurred
|
||||
across stories (2% of the corpus). At ~300 words, embedding similarity tracks
|
||||
which story a scene is from, not what kind of situation it is.
|
||||
- The proposed fix, untested: summarise each scene into one generic sentence
|
||||
with no proper names, embed the summary, then cluster.
|
||||
- It could not be tested because generation on the LAN Ollama host hung. The
|
||||
model ran on CPU with no GPU offload, and the 16k-context variant stayed
|
||||
loaded. That is probably the same fault that blocks Interactive Story M01.
|
||||
- The last open question was whether to use Claude for the generation steps
|
||||
until that host is fixed. Embeddings stay on Ollama, because Anthropic has no
|
||||
embeddings API.
|
||||
Reference in New Issue
Block a user