The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
Situation catalogue: classify, don't cluster
Built 2026-09-17 from the three hand references: O. Henry (../v2/claude/), Wharton (../wharton/claude/)
and Jacobs (../jacobs/claude/). This is the step that replaces embed-and-cluster in the proposed pipeline:
a model assigns each scene's predicament one catalogue entry, or none.
Files
| File | What |
|---|---|
catalogue_src.py |
The 36 entries (definition, includes, excludes, tone, recurs) and SOURCES, which maps every hand group in the three references to exactly one entry. Running it checks the mapping and writes the rest. No inference. |
catalogue.json |
The entries, with up to three worked examples per corpus, each tagged with its corpus |
labels-<corpus>.json |
Every scene's catalogue label from the hand grouping, or none. This is the answer key for scoring |
recurrence.log |
One recurrence rule applied to all three: at least 5 scenes, at least max(3, 4% of the stories) stories, and no story over 50% |
classify.py |
Classifies one scene per request through ../ollama.py. Leave-one-corpus-out: the examples never come from the corpus being classified. --k 0 offers the whole catalogue |
score.py |
Scores classifier output against labels-*.json. --selftest checks the scorer on hand-vs-hand and on random labels |
emb/ |
Cached nomic-embed-text vectors for the catalogue and the O. Henry 14B predicaments |
out/ |
Classifier output |
Results so far
1. One rule for all three corpora (recurrence.log). O. Henry keeps 16 recurring situations and
Wharton keeps 16. Jacobs falls from its hand-counted 19 to 6, because its groups had used a two-story
minimum. With 12 stories, three stories a group is the least that means "recurs", and most of its
groups do not reach it. Only rival, strings and newcomer are found in all three corpora; ten more
are in two. A 20-30 story corpus is a real floor.
2. Clustering fails even on the best text. My 839 hand-written O. Henry predicaments were
embedded exactly as embed.py embedded the 14B ones (../v2/claude/predicament_embeddings-claude.json),
and the same k = 60 k-means was run on them (../v2/claude/compare-claude-embeddings.log). The groups
agree with my own hand grouping at ARI 0.05 (0.05-0.05 over four seeds), no better than the 14B text's
0.07. The fault is the embed-and-cluster step, not the predicament text. Better rewriting will not
rescue clustering.
3. Embeddings cannot even shortlist. With the 14B predicaments against the catalogue definitions, the
hand label is the nearest entry for 13% of situated scenes, is in the top 8 for 50%, and is in the top 20
for 83%. So classify.py now offers the whole catalogue (--k 0), with the catalogue in the system prompt
so the server can reuse the cached prefix.
4. Classifier pilot: 40 scenes only, too small to judge (pilot40.log; qwen3:14b, shortlist k = 8,
14B predicaments). The model called 34 of 40 scenes situated where the hand labels have 14. Where both
called a scene situated, its label was right in 3 of 13 (23%). Kappa was 0.09. The prompt has since been
changed: the whole catalogue, a note to expect about half none, and a fit rating from 1 to 3 that
score.py --minfit= can threshold. The changed prompt has not been measured.
4b. The changed prompt, on 30 scenes before the host hung (out/ohenry-q14-dev1.partial.json, scored
in pilot40.log): it tells situated scenes from none better (precision 67%, recall 100%, kappa 0.23), but
where both call a scene situated, its label is right in only 2 of 14. Reading the confusions, some of its
labels are defensible. I labelled a whole story arc with one situation (every scene of A Lickpenny Lover is
station), while the model judges each scene alone (pretence, unspoken). And the 14B predicaments
("A man struggles to reconcile his pride with...") are too vague to classify from. So the next development
round should (a) also give the classifier the scene summary, or the predicaments either side, and (b) score
at the level a pack uses: which situations recur, with what members, rather than scene-exact agreement.
Not done: the inference host's GPU fell off the bus
At 22:19:07 on 2026-09-17, the inference host's GPU fell off the PCIe bus (kernel: Xid 79, then Xid 154, GPU Reset Required). It happened in the middle of a classify.py development request whose prompt was about 2,450 tokens (the
whole catalogue). Over the ten minutes before, calls had slowed from about 2 s a scene to about 20 s. This is the
second such fault; the first was on 2026-09-13. No power or link loggers were running, because SSH was refused
while the key was out of the agent, so there is no evidence either way on power. After the fault, Ollama carried on
on the CPU. That is why nomic-embed-text still answered, and finding 2 was computed on CPU embeddings, which does
not change the method. All client processes were stopped. The host needs a reboot (reset required). The NVIDIA
driver asks for nvidia-bug-report.sh to be run as root before the module is unloaded.
The prompt design matters for this host: offering the whole catalogue makes every call a long prefill. Either keep
the catalogue as a cached prefix (the current classify.py, which puts it in the system prompt; untested on the
GPU), or shortlist with something better than embeddings, or run with a lower power limit and the loggers on.
Third fault, with the loggers running: 2026-09-18 01:00:03
After a reboot, with the power limit set to 200 W (from 280 W) and all three loggers running, a 20-scene pilot
(classify.py ohenry q14 --limit 20 --tag=-dev1, catalogue as the cached system prompt) lost the GPU about 15 s in.
The logs are in gpu-logs/*-2026-09-18-0056.*. From the link/power CSV:
| Time | Power | Utilisation | Notes |
|---|---|---|---|
| 00:59:54 | 309.31 W | 49% | enforced limit 200 W; throttle reason 0x4 (SW power cap) |
| 00:59:55 to 00:59:59 | about 178 to 202 W | 100% | sustained prefill |
| 01:00:02 | 174.96 W | last good sample | |
| 01:00:03 | "GPU is lost"; dmon shows - from here |
The temperature never went above 39 °C, so the cause is not thermal. The link was PCIe gen 3 x4. The card drew 309 W for a sample despite a 200 W cap, and then dropped after about ten seconds of load at the cap. That points to power delivery to the card or the OCuLink dock, not to the prompt or the software. A lower cap did not prevent it. No more classification runs on that host until its hardware has been looked at. The pilot output is not a result.
gpu-logs/nvidia-bug-report-2026-09-18.log.gz is the driver's report, taken after the fault. The kernel log in it
shows no AER or PCIe error messages at any fault. That is not evidence of a clean link: at boot the kernel logs _OSC: platform does not support [SHPCHotplug AER LTR DPC], so the firmware never gives the OS AER control and link errors would not be reported. There
was an Xid 32 (corrupted push-buffer stream) from Ollama at 00:54:48, five minutes before this drop, and an Xid 13
graphics exception on 2026-09-13, before the first drop.
Next, once the host is healthy
The plan below is fixed before any of the runs, so they cannot be tuned after the fact.
- Development slice: O. Henry scenes 0-119. Tune the prompt and the
--minfitthreshold only here. (classify.py ohenry q14 --limit 120 --tag=-dev1) - Test: O. Henry scenes 120-838, run once (
--start 120 --tag=-test). Score withscore.py out/ohenry-q14-test.json --range=120:839. Then run the same prompt on the hand-written predicaments (ohenry claude) to separate rewrite quality from classification quality. Then runwharton claudeandjacobs claude, which have no local-model predicaments yet, to see how the catalogue transfers. - What counts as success is precision per situation, not the number of situations that pass:
--selftestshows random labels also "pass" all 16. The bar is set before the test run: for the situations both hand and model call recurring, a median model precision of at least 60%. - At about 2 s a scene, O. Henry takes about 30 minutes a text set. That is a long run, so start the loggers first.