Files
TheLadder/tools/story-to-pack/probe/catalogue
JesseMarkowitzandClaude Opus 5.5 3a4758172d story-to-pack: add the corpora and the catalogue v2/v3
Fourteen corpora split and classified: aesop, bierce, chekhov, holmes,
keefe, lawson, lorimer, maupassant, nobody, plaintales, poe, torchy,
wallingford and winesburg, each with its splitter and the hand-written
groups and pages; catalogue v2 and v3; and the shared splitters
gutenberg_chunks.py, se_split.py and se_build.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rwKTmug58sEsJ72AuWsEi
2026-10-07 06:05:51 -04:00
..

Situation catalogue: classify, don't cluster

Built 2026-09-17 from the three hand references: O. Henry (../v2/claude/), Wharton (../wharton/claude/) and Jacobs (../jacobs/claude/). This is the step that replaces embed-and-cluster in the proposed pipeline: a model assigns each scene's predicament one catalogue entry, or none.

Files

File What
catalogue_src.py The 36 entries (definition, includes, excludes, tone, recurs) and SOURCES, which maps every hand group in the three references to exactly one entry. Running it checks the mapping and writes the rest. No inference.
catalogue.json The entries, with up to three worked examples per corpus, each tagged with its corpus
labels-<corpus>.json Every scene's catalogue label from the hand grouping, or none. This is the answer key for scoring
recurrence.log One recurrence rule applied to all three: at least 5 scenes, at least max(3, 4% of the stories) stories, and no story over 50%
classify.py Classifies one scene per request through ../ollama.py. Leave-one-corpus-out: the examples never come from the corpus being classified. --k 0 offers the whole catalogue
score.py Scores classifier output against labels-*.json. --selftest checks the scorer on hand-vs-hand and on random labels
emb/ Cached nomic-embed-text vectors for the catalogue and the O. Henry 14B predicaments
out/ Classifier output

Results so far

1. One rule for all three corpora (recurrence.log). O. Henry keeps 16 recurring situations and Wharton keeps 16. Jacobs falls from its hand-counted 19 to 6, because its groups had used a two-story minimum. With 12 stories, three stories a group is the least that means "recurs", and most of its groups do not reach it. Only rival, strings and newcomer are found in all three corpora; ten more are in two. A 20-30 story corpus is a real floor.

2. Clustering fails even on the best text. My 839 hand-written O. Henry predicaments were embedded exactly as embed.py embedded the 14B ones (../v2/claude/predicament_embeddings-claude.json), and the same k = 60 k-means was run on them (../v2/claude/compare-claude-embeddings.log). The groups agree with my own hand grouping at ARI 0.05 (0.05-0.05 over four seeds), no better than the 14B text's 0.07. The fault is the embed-and-cluster step, not the predicament text. Better rewriting will not rescue clustering.

3. Embeddings cannot even shortlist. With the 14B predicaments against the catalogue definitions, the hand label is the nearest entry for 13% of situated scenes, is in the top 8 for 50%, and is in the top 20 for 83%. So classify.py now offers the whole catalogue (--k 0), with the catalogue in the system prompt so the server can reuse the cached prefix.

4. Classifier pilot: 40 scenes only, too small to judge (pilot40.log; qwen3:14b, shortlist k = 8, 14B predicaments). The model called 34 of 40 scenes situated where the hand labels have 14. Where both called a scene situated, its label was right in 3 of 13 (23%). Kappa was 0.09. The prompt has since been changed: the whole catalogue, a note to expect about half none, and a fit rating from 1 to 3 that score.py --minfit= can threshold. The changed prompt has not been measured.

4b. The changed prompt, on 30 scenes before the host hung (out/ohenry-q14-dev1.partial.json, scored in pilot40.log): it tells situated scenes from none better (precision 67%, recall 100%, kappa 0.23), but where both call a scene situated, its label is right in only 2 of 14. Reading the confusions, some of its labels are defensible. I labelled a whole story arc with one situation (every scene of A Lickpenny Lover is station), while the model judges each scene alone (pretence, unspoken). And the 14B predicaments ("A man struggles to reconcile his pride with...") are too vague to classify from. So the next development round should (a) also give the classifier the scene summary, or the predicaments either side, and (b) score at the level a pack uses: which situations recur, with what members, rather than scene-exact agreement.

Not done: the inference host's GPU fell off the bus

At 22:19:07 on 2026-09-17, the inference host's GPU fell off the PCIe bus (kernel: Xid 79, then Xid 154, GPU Reset Required). It happened in the middle of a classify.py development request whose prompt was about 2,450 tokens (the whole catalogue). Over the ten minutes before, calls had slowed from about 2 s a scene to about 20 s. This is the second such fault; the first was on 2026-09-13. No power or link loggers were running, because SSH was refused while the key was out of the agent, so there is no evidence either way on power. After the fault, Ollama carried on on the CPU. That is why nomic-embed-text still answered, and finding 2 was computed on CPU embeddings, which does not change the method. All client processes were stopped. The host needs a reboot (reset required). The NVIDIA driver asks for nvidia-bug-report.sh to be run as root before the module is unloaded.

The prompt design matters for this host: offering the whole catalogue makes every call a long prefill. Either keep the catalogue as a cached prefix (the current classify.py, which puts it in the system prompt; untested on the GPU), or shortlist with something better than embeddings, or run with a lower power limit and the loggers on.

Third fault, with the loggers running: 2026-09-18 01:00:03

After a reboot, with the power limit set to 200 W (from 280 W) and all three loggers running, a 20-scene pilot (classify.py ohenry q14 --limit 20 --tag=-dev1, catalogue as the cached system prompt) lost the GPU about 15 s in. The logs are in gpu-logs/*-2026-09-18-0056.*. From the link/power CSV:

Time Power Utilisation Notes
00:59:54 309.31 W 49% enforced limit 200 W; throttle reason 0x4 (SW power cap)
00:59:55 to 00:59:59 about 178 to 202 W 100% sustained prefill
01:00:02 174.96 W last good sample
01:00:03 "GPU is lost"; dmon shows - from here

The temperature never went above 39 °C, so the cause is not thermal. The link was PCIe gen 3 x4. The card drew 309 W for a sample despite a 200 W cap, and then dropped after about ten seconds of load at the cap. That points to power delivery to the card or the OCuLink dock, not to the prompt or the software. A lower cap did not prevent it. No more classification runs on that host until its hardware has been looked at. The pilot output is not a result.

gpu-logs/nvidia-bug-report-2026-09-18.log.gz is the driver's report, taken after the fault. The kernel log in it shows no AER or PCIe error messages at any fault. That is not evidence of a clean link: at boot the kernel logs _OSC: platform does not support [SHPCHotplug AER LTR DPC], so the firmware never gives the OS AER control and link errors would not be reported. There was an Xid 32 (corrupted push-buffer stream) from Ollama at 00:54:48, five minutes before this drop, and an Xid 13 graphics exception on 2026-09-13, before the first drop.

Next, once the host is healthy

The plan below is fixed before any of the runs, so they cannot be tuned after the fact.

  • Development slice: O. Henry scenes 0-119. Tune the prompt and the --minfit threshold only here. (classify.py ohenry q14 --limit 120 --tag=-dev1)
  • Test: O. Henry scenes 120-838, run once (--start 120 --tag=-test). Score with score.py out/ohenry-q14-test.json --range=120:839. Then run the same prompt on the hand-written predicaments (ohenry claude) to separate rewrite quality from classification quality. Then run wharton claude and jacobs claude, which have no local-model predicaments yet, to see how the catalogue transfers.
  • What counts as success is precision per situation, not the number of situations that pass: --selftest shows random labels also "pass" all 16. The bar is set before the test run: for the situations both hand and model call recurring, a median model precision of at least 60%.
  • At about 2 s a scene, O. Henry takes about 30 minutes a text set. That is a long run, so start the loggers first.