Fourteen corpora split and classified: aesop, bierce, chekhov, holmes, keefe, lawson, lorimer, maupassant, nobody, plaintales, poe, torchy, wallingford and winesburg, each with its splitter and the hand-written groups and pages; catalogue v2 and v3; and the shared splitters gutenberg_chunks.py, se_split.py and se_build.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019rwKTmug58sEsJ72AuWsEi
Situation catalogue: classify, don't cluster
Built 2026-09-17 from the three hand references: O. Henry (../v2/claude/), Wharton (../wharton/claude/)
and Jacobs (../jacobs/claude/). This is the step that replaces embed-and-cluster in the proposed pipeline:
a model assigns each scene's predicament one catalogue entry, or none.
Files
| File | What |
|---|---|
catalogue_src.py |
The 36 entries (definition, includes, excludes, tone, recurs) and SOURCES, which maps every hand group in the three references to exactly one entry. Running it checks the mapping and writes the rest. No inference. |
catalogue.json |
The entries, with up to three worked examples per corpus, each tagged with its corpus |
labels-<corpus>.json |
Every scene's catalogue label from the hand grouping, or none. This is the answer key for scoring |
recurrence.log |
One recurrence rule applied to all three: at least 5 scenes, at least max(3, 4% of the stories) stories, and no story over 50% |
classify.py |
Classifies one scene per request through ../ollama.py. Leave-one-corpus-out: the examples never come from the corpus being classified. --k 0 offers the whole catalogue |
score.py |
Scores classifier output against labels-*.json. --selftest checks the scorer on hand-vs-hand and on random labels |
emb/ |
Cached nomic-embed-text vectors for the catalogue and the O. Henry 14B predicaments |
out/ |
Classifier output |
Results so far
1. One rule for all three corpora (recurrence.log). O. Henry keeps 16 recurring situations and
Wharton keeps 16. Jacobs falls from its hand-counted 19 to 6, because its groups had used a two-story
minimum. With 12 stories, three stories a group is the least that means "recurs", and most of its
groups do not reach it. Only rival, strings and newcomer are found in all three corpora; ten more
are in two. A 20-30 story corpus is a real floor.
2. Clustering fails even on the best text. My 839 hand-written O. Henry predicaments were
embedded exactly as embed.py embedded the 14B ones (../v2/claude/predicament_embeddings-claude.json),
and the same k = 60 k-means was run on them (../v2/claude/compare-claude-embeddings.log). The groups
agree with my own hand grouping at ARI 0.05 (0.05-0.05 over four seeds), no better than the 14B text's
0.07. The fault is the embed-and-cluster step, not the predicament text. Better rewriting will not
rescue clustering.
3. Embeddings cannot even shortlist. With the 14B predicaments against the catalogue definitions, the
hand label is the nearest entry for 13% of situated scenes, is in the top 8 for 50%, and is in the top 20
for 83%. So classify.py now offers the whole catalogue (--k 0), with the catalogue in the system prompt
so the server can reuse the cached prefix.
4. Classifier pilot: 40 scenes only, too small to judge (pilot40.log; qwen3:14b, shortlist k = 8,
14B predicaments). The model called 34 of 40 scenes situated where the hand labels have 14. Where both
called a scene situated, its label was right in 3 of 13 (23%). Kappa was 0.09. The prompt has since been
changed: the whole catalogue, a note to expect about half none, and a fit rating from 1 to 3 that
score.py --minfit= can threshold. The changed prompt has not been measured.
4b. The changed prompt, on 30 scenes before the host hung (out/ohenry-q14-dev1.partial.json, scored
in pilot40.log): it tells situated scenes from none better (precision 67%, recall 100%, kappa 0.23), but
where both call a scene situated, its label is right in only 2 of 14. Reading the confusions, some of its
labels are defensible. I labelled a whole story arc with one situation (every scene of A Lickpenny Lover is
station), while the model judges each scene alone (pretence, unspoken). And the 14B predicaments
("A man struggles to reconcile his pride with...") are too vague to classify from. So the next development
round should (a) also give the classifier the scene summary, or the predicaments either side, and (b) score
at the level a pack uses: which situations recur, with what members, rather than scene-exact agreement.
Not done: the inference host's GPU fell off the bus
At 22:19:07 on 2026-09-17, the inference host's GPU fell off the PCIe bus (kernel: Xid 79, then Xid 154, GPU Reset Required). It happened in the middle of a classify.py development request whose prompt was about 2,450 tokens (the
whole catalogue). Over the ten minutes before, calls had slowed from about 2 s a scene to about 20 s. This is the
second such fault; the first was on 2026-09-13. No power or link loggers were running, because SSH was refused
while the key was out of the agent, so there is no evidence either way on power. After the fault, Ollama carried on
on the CPU. That is why nomic-embed-text still answered, and finding 2 was computed on CPU embeddings, which does
not change the method. All client processes were stopped. The host needs a reboot (reset required). The NVIDIA
driver asks for nvidia-bug-report.sh to be run as root before the module is unloaded.
The prompt design matters for this host: offering the whole catalogue makes every call a long prefill. Either keep
the catalogue as a cached prefix (the current classify.py, which puts it in the system prompt; untested on the
GPU), or shortlist with something better than embeddings, or run with a lower power limit and the loggers on.
Third fault, with the loggers running: 2026-09-18 01:00:03
After a reboot, with the power limit set to 200 W (from 280 W) and all three loggers running, a 20-scene pilot
(classify.py ohenry q14 --limit 20 --tag=-dev1, catalogue as the cached system prompt) lost the GPU about 15 s in.
The logs are in gpu-logs/*-2026-09-18-0056.*. From the link/power CSV:
| Time | Power | Utilisation | Notes |
|---|---|---|---|
| 00:59:54 | 309.31 W | 49% | enforced limit 200 W; throttle reason 0x4 (SW power cap) |
| 00:59:55 to 00:59:59 | about 178 to 202 W | 100% | sustained prefill |
| 01:00:02 | 174.96 W | last good sample | |
| 01:00:03 | "GPU is lost"; dmon shows - from here |
The temperature never went above 39 °C, so the cause is not thermal. The link was PCIe gen 3 x4. The card drew 309 W for a sample despite a 200 W cap, and then dropped after about ten seconds of load at the cap. That points to power delivery to the card or the OCuLink dock, not to the prompt or the software. A lower cap did not prevent it. No more classification runs on that host until its hardware has been looked at. The pilot output is not a result.
gpu-logs/nvidia-bug-report-2026-09-18.log.gz is the driver's report, taken after the fault. The kernel log in it
shows no AER or PCIe error messages at any fault. That is not evidence of a clean link: at boot the kernel logs _OSC: platform does not support [SHPCHotplug AER LTR DPC], so the firmware never gives the OS AER control and link errors would not be reported. There
was an Xid 32 (corrupted push-buffer stream) from Ollama at 00:54:48, five minutes before this drop, and an Xid 13
graphics exception on 2026-09-13, before the first drop.
Next, once the host is healthy
The plan below is fixed before any of the runs, so they cannot be tuned after the fact.
- Development slice: O. Henry scenes 0-119. Tune the prompt and the
--minfitthreshold only here. (classify.py ohenry q14 --limit 120 --tag=-dev1) - Test: O. Henry scenes 120-838, run once (
--start 120 --tag=-test). Score withscore.py out/ohenry-q14-test.json --range=120:839. Then run the same prompt on the hand-written predicaments (ohenry claude) to separate rewrite quality from classification quality. Then runwharton claudeandjacobs claude, which have no local-model predicaments yet, to see how the catalogue transfers. - What counts as success is precision per situation, not the number of situations that pass:
--selftestshows random labels also "pass" all 16. The bar is set before the test run: for the situations both hand and model call recurring, a median model precision of at least 60%. - At about 2 s a scene, O. Henry takes about 30 minutes a text set. That is a long run, so start the loggers first.