Files
TheLadder/tools/story-to-pack/probe/catalogue
JesseMarkowitzandClaude Opus 5 e5617b86ba Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.

Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.

Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.

Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.

Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.

The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.

Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
2026-09-20 17:19:42 -04:00
..

Situation catalogue: classify, don't cluster

Built 2026-09-17 from the three hand references: O. Henry (../v2/claude/), Wharton (../wharton/claude/) and Jacobs (../jacobs/claude/). This is the step that replaces embed-and-cluster in the proposed pipeline: a model assigns each scene's predicament one catalogue entry, or none.

Files

File What
catalogue_src.py The 36 entries (definition, includes, excludes, tone, recurs) and SOURCES, which maps every hand group in the three references to exactly one entry. Running it checks the mapping and writes the rest. No inference.
catalogue.json The entries, with up to three worked examples per corpus, each tagged with its corpus
labels-<corpus>.json Every scene's catalogue label from the hand grouping, or none. This is the answer key for scoring
recurrence.log One recurrence rule applied to all three: at least 5 scenes, at least max(3, 4% of the stories) stories, and no story over 50%
classify.py Classifies one scene per request through ../ollama.py. Leave-one-corpus-out: the examples never come from the corpus being classified. --k 0 offers the whole catalogue
score.py Scores classifier output against labels-*.json. --selftest checks the scorer on hand-vs-hand and on random labels
emb/ Cached nomic-embed-text vectors for the catalogue and the O. Henry 14B predicaments
out/ Classifier output

Results so far

1. One rule for all three corpora (recurrence.log). O. Henry keeps 16 recurring situations and Wharton keeps 16. Jacobs falls from its hand-counted 19 to 6, because its groups had used a two-story minimum. With 12 stories, three stories a group is the least that means "recurs", and most of its groups do not reach it. Only rival, strings and newcomer are found in all three corpora; ten more are in two. A 20-30 story corpus is a real floor.

2. Clustering fails even on the best text. My 839 hand-written O. Henry predicaments were embedded exactly as embed.py embedded the 14B ones (../v2/claude/predicament_embeddings-claude.json), and the same k = 60 k-means was run on them (../v2/claude/compare-claude-embeddings.log). The groups agree with my own hand grouping at ARI 0.05 (0.05-0.05 over four seeds), no better than the 14B text's 0.07. The fault is the embed-and-cluster step, not the predicament text. Better rewriting will not rescue clustering.

3. Embeddings cannot even shortlist. With the 14B predicaments against the catalogue definitions, the hand label is the nearest entry for 13% of situated scenes, is in the top 8 for 50%, and is in the top 20 for 83%. So classify.py now offers the whole catalogue (--k 0), with the catalogue in the system prompt so the server can reuse the cached prefix.

4. Classifier pilot: 40 scenes only, too small to judge (pilot40.log; qwen3:14b, shortlist k = 8, 14B predicaments). The model called 34 of 40 scenes situated where the hand labels have 14. Where both called a scene situated, its label was right in 3 of 13 (23%). Kappa was 0.09. The prompt has since been changed: the whole catalogue, a note to expect about half none, and a fit rating from 1 to 3 that score.py --minfit= can threshold. The changed prompt has not been measured.

4b. The changed prompt, on 30 scenes before the host hung (out/ohenry-q14-dev1.partial.json, scored in pilot40.log): it tells situated scenes from none better (precision 67%, recall 100%, kappa 0.23), but where both call a scene situated, its label is right in only 2 of 14. Reading the confusions, some of its labels are defensible. I labelled a whole story arc with one situation (every scene of A Lickpenny Lover is station), while the model judges each scene alone (pretence, unspoken). And the 14B predicaments ("A man struggles to reconcile his pride with...") are too vague to classify from. So the next development round should (a) also give the classifier the scene summary, or the predicaments either side, and (b) score at the level a pack uses: which situations recur, with what members, rather than scene-exact agreement.

Not done: the inference host's GPU fell off the bus

At 22:19:07 on 2026-09-17, the inference host's GPU fell off the PCIe bus (kernel: Xid 79, then Xid 154, GPU Reset Required). It happened in the middle of a classify.py development request whose prompt was about 2,450 tokens (the whole catalogue). Over the ten minutes before, calls had slowed from about 2 s a scene to about 20 s. This is the second such fault; the first was on 2026-09-13. No power or link loggers were running, because SSH was refused while the key was out of the agent, so there is no evidence either way on power. After the fault, Ollama carried on on the CPU. That is why nomic-embed-text still answered, and finding 2 was computed on CPU embeddings, which does not change the method. All client processes were stopped. The host needs a reboot (reset required). The NVIDIA driver asks for nvidia-bug-report.sh to be run as root before the module is unloaded.

The prompt design matters for this host: offering the whole catalogue makes every call a long prefill. Either keep the catalogue as a cached prefix (the current classify.py, which puts it in the system prompt; untested on the GPU), or shortlist with something better than embeddings, or run with a lower power limit and the loggers on.

Third fault, with the loggers running: 2026-09-18 01:00:03

After a reboot, with the power limit set to 200 W (from 280 W) and all three loggers running, a 20-scene pilot (classify.py ohenry q14 --limit 20 --tag=-dev1, catalogue as the cached system prompt) lost the GPU about 15 s in. The logs are in gpu-logs/*-2026-09-18-0056.*. From the link/power CSV:

Time Power Utilisation Notes
00:59:54 309.31 W 49% enforced limit 200 W; throttle reason 0x4 (SW power cap)
00:59:55 to 00:59:59 about 178 to 202 W 100% sustained prefill
01:00:02 174.96 W last good sample
01:00:03 "GPU is lost"; dmon shows - from here

The temperature never went above 39 °C, so the cause is not thermal. The link was PCIe gen 3 x4. The card drew 309 W for a sample despite a 200 W cap, and then dropped after about ten seconds of load at the cap. That points to power delivery to the card or the OCuLink dock, not to the prompt or the software. A lower cap did not prevent it. No more classification runs on that host until its hardware has been looked at. The pilot output is not a result.

gpu-logs/nvidia-bug-report-2026-09-18.log.gz is the driver's report, taken after the fault. The kernel log in it shows no AER or PCIe error messages at any fault. That is not evidence of a clean link: at boot the kernel logs _OSC: platform does not support [SHPCHotplug AER LTR DPC], so the firmware never gives the OS AER control and link errors would not be reported. There was an Xid 32 (corrupted push-buffer stream) from Ollama at 00:54:48, five minutes before this drop, and an Xid 13 graphics exception on 2026-09-13, before the first drop.

Next, once the host is healthy

The plan below is fixed before any of the runs, so they cannot be tuned after the fact.

  • Development slice: O. Henry scenes 0-119. Tune the prompt and the --minfit threshold only here. (classify.py ohenry q14 --limit 120 --tag=-dev1)
  • Test: O. Henry scenes 120-838, run once (--start 120 --tag=-test). Score with score.py out/ohenry-q14-test.json --range=120:839. Then run the same prompt on the hand-written predicaments (ohenry claude) to separate rewrite quality from classification quality. Then run wharton claude and jacobs claude, which have no local-model predicaments yet, to see how the catalogue transfers.
  • What counts as success is precision per situation, not the number of situations that pass: --selftest shows random labels also "pass" all 16. The bar is set before the test run: for the situations both hand and model call recurring, a median model precision of at least 60%.
  • At about 2 s a scene, O. Henry takes about 30 minutes a text set. That is a long run, so start the loggers first.