The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
54 KiB
story-to-pack — selecting situations
Decision (2026-09-15): selection is gate G2 followed by a structured human review.
probe/review.pywritesREVIEW.html, a page where the reviewer keeps, drops or merges each group, unticks scenes and names the situation; the page saves a review file, andprobe/apply_review.pychecks it and writessituations.jsonfor the next stage. See the README.The first version asked the reviewer to read a long Markdown file and edit JSON by hand. The first real attempt stopped before a single entry: what to decide was unclear, there was too much to read at once, hand-edited JSON broke easily, and the wording was hard to write. The page answers each of those. The record below is how the G2 decision was reached: four gate designs and two model-judged gates, each tested against blind labels.
Superseded in part (2026-09-15), see "Predicaments" at the end: reading the page, the reviewer found the groups were topics, not situations.
The step after clustering: from the clusters of summarised scenes, choose the ones that are genuine
recurring situations — the raw material of a recurring event — and reject the rest. Companion to
DESIGN.md; this is the step its assumption-1 result called for.
Status (2026-09-14): designed and pre-registered below, before any feature was computed or any label was read against a feature. Results are appended at the end, and nothing above the Results heading is edited after the test runs, except to fix a typo.
The problem, from the assumption-1 result
Summarise-then-embed broke the story-identity effect: clusters now draw from many stories. But of five
sampled tight-recurring clusters, one was a vocabulary cluster — fortune-tellers, adventurers, a
found tyre, a graft inquiry — held together by "seeks", "fortune", "mysterious", and mostly from three
stories. The concentration and tightness filters in measure.py let it through. A selector has to
reject clusters like that one and keep clusters like "a spouse schemes against or confronts a spouse
over infidelity".
Constraints
- No inference. Selection uses only what is already on disk: the summary embeddings, the prose embeddings, the summary text, and which story each scene came from. It is pure computation, like the numeric search, so it can be re-run and tuned for free. (It also keeps the inference host free for other work.)
- Deterministic. Seeded clustering; the same inputs give the same selection.
- Pure Python, like the rest of the probe.
What a selected situation is
A set of scenes, from at least 4 different stories, with no story contributing more than 40%, that a reader can name as one social transaction: who wants what from whom. It becomes one recurring event; its member scenes become that event's description variants and option material.
Rejected: story-dominated clusters (already mechanical), vocabulary clusters (shared words, no shared transaction), register clusters (shared tone or setting, e.g. "men walking through the city"), and near-duplicates of an already-selected situation.
Candidate signals
All computed per cluster, from existing vectors and text. Each carries a pre-registered prediction of which way it moves for a genuine situation.
| # | Signal | Definition | Prediction for a situation |
|---|---|---|---|
| S1 | Cross-story cohesion | Mean summary-embedding cosine over member pairs from different stories | Higher |
| S2 | Cross/within ratio | S1 divided by the mean cosine over member pairs from the same story (1.0 if there are none) | Closer to 1 — a situation is as similar across stories as within one; a cluster held together by a few stories' scenes is not |
| S3 | Effective stories | Inverse Simpson index of the member story counts, divided by cluster size | Higher — members spread evenly rather than a few stories plus stragglers |
| S4 | Prose agreement | Mean prose-embedding cosine over cross-story member pairs, minus the corpus-wide mean cross-story prose cosine | Higher — a second, independent view: scenes of the same situation should read a little alike even in the original text |
| S5 | Stability | Mean pairwise co-association of members across 12 k-means runs (k from 50 to 100, several seeds): the share of runs in which two members land together | Higher |
| S6 | Word dependence | Share of members containing the cluster's single most over-represented content word (log-odds against the whole summary set, stopwords and "someone" excluded) | Higher for a vocabulary cluster — one word doing the work |
S6 is predicted to indicate the failure, the others to indicate success.
Protocol
-
Candidates. Cluster the summary embeddings with k = 60. A candidate is a cluster with ≥ 5 members, ≥ 4 distinct stories and dominant share ≤ 40% (
measure.py's "recurring"; the relative tightness filter is not applied — the selector replaces it). -
Two candidate sets from two partitions. Development: seed 11. Test: seed 12.
-
Blind labels. Each set is dumped to a sheet showing, per cluster, only its member summaries in shuffled order — no story titles, no scores, no sizes beyond the lines themselves. Each cluster is labelled by reading:
- 2 — one situation: one nameable social transaction that most members (roughly 70%+) instantiate. The label records its name.
- 1 — broad: a shared setting or theme, but the transaction varies across members.
- 0 — not a situation: vocabulary, register, or nothing nameable.
The test sheet is labelled before any signal is computed for the test set.
-
Choose the rule on development only. Compute S1–S6 for the development set. Candidate rules are single thresholds on one signal, or an AND of two. The chosen rule is the one with the highest precision (share of accepted clusters labelled 2), subject to accepting at least half of the development set's label-2 clusters; ties go to the simpler rule. Freeze it here, in writing, before step 5.
Detail fixed after the labels were written and before any signal was computed: thresholds are placed midway between adjacent observed development values, never on one. Ties beyond simplicity break on more label-2 clusters accepted, then on precision for label ≥ 1. Two reference points are reported but cannot be chosen: accepting every candidate, and
measure.py's tightness filter (coherence at or above the partition's median sized cluster), called B0. -
Test once. Apply the frozen rule to the test set. Report precision (label 2, and label ≥ 1), recall of label-2 clusters, and the number accepted. No re-tuning after seeing test results; any later change is a new rule and needs new data to be tested on.
-
Near-duplicates. Among accepted clusters, merge any pair whose centroids have cosine ≥ a threshold chosen on development by reading the labels' names. Report how many merges happened on the test set and whether each merged pair's names agree.
Known limits of this test
- The labels are one reader's judgement — the model running the probe. They are written to a file so the designer can spot-check or overrule them, and the verdict should be read with that in mind.
- Development and test come from the same corpus. Two partitions of the same 838 scenes share situations, so the test measures stability across clustering, not generalisation to a new corpus. A genuinely held-out test needs a second corpus, which needs embedding, which needs the inference host.
- k is fixed at 60 for the test. The selector's behaviour at other k is not tested.
Frozen rule (written before any test signal was computed)
Chosen on development (40 candidates, 5 labelled 2; must accept ≥ 3), by probe/select_rule.py choose:
Accept a candidate when S1 ≥ 0.6119 AND S6 ≤ 0.1603. Cross-story cohesion at least 0.6119, and no single over-represented word in more than 16% of members.
| On development | Accepted | Precision (2) | Precision (≥ 1) | Recall (2) |
|---|---|---|---|---|
| Accept every candidate | 40 | 0.12 | 0.62 | 1.00 |
| B0, tightness filter | 17 | 0.24 | 0.76 | 0.80 |
| S1 alone (best single) | 9 | 0.44 | 0.89 | 0.80 |
| Frozen rule | 5 | 0.60 | 1.00 | 0.60 |
Accepted on development: D04 (2), D09 (1), D13 (2), D16 (2), D33 (1). Rejected label-2: D15 (S1 0.603),
D30 (S6 0.20). Full leaderboard: probe/choose-dev.log; signals: probe/signals-dev.json.
Observations recorded now, so they cannot be shaped by the test:
- This is very likely overfit. Five positives, fine-grained thresholds, and an AND of two signals is exactly the setting in which a rule memorises its development set. The test is what decides.
- S1 does most of the work. It is the best single signal by a wide margin, and it alone rejects the vocabulary cluster that motivated this step (D18, S1 0.605).
- S6 did not behave as predicted. Its most over-represented word is usually a story's word ("juggler", "furs", "eviction"), not generic vocabulary, and D18's is low (0.11). It catches D39 ("young", 1.00) — a vocabulary cluster of a different kind. Its contribution to the frozen rule is a single cut, and may not carry over.
- S2, S3 and S4 were weak on development, and S5 middling. None of the predicted directions was strongly wrong, but only S1 separated well.
Near-duplicates (step 6): centroid cosine cannot do it, so no automatic merge is applied. Among the five clusters the rule accepts on development, no two share a situation, so there was nothing there to calibrate on. Across all 26 development candidates labelled ≥ 1, the one clear duplicate — D13 and D30, both "a run-in with the police" — has centroid cosine 0.9026, while five pairs of plainly different situations score higher, up to 0.9326 ("money and marriage" against "schemes, cons and theft"). D05/D23 (artists) sit at 0.9005. Any threshold that merges the duplicate merges different situations first. Summary embeddings place situations about money, romance and conflict in one dense region, so closeness between centroids is not identity of situation.
Decision, fixed before the test: the test reports, by label name, whether any two accepted clusters are the same situation, as a measurement of how often duplicates survive selection. A merge method is an open problem; member overlap across partitions (co-association between two clusters' members) is the candidate to try next, and is untested.
Results
Test run once, 2026-09-14, on the seed-12 partition (41 candidates, 5 labelled 2), frozen rule
unchanged. Logs: probe/signals-test.json, probe/apply-test.log.
| On test | Accepted | Precision (2) | Precision (≥ 1) | Recall (2) |
|---|---|---|---|---|
| Accept every candidate | 41 | 0.12 | 0.51 | 1.00 |
| B0, tightness filter | 18 | 0.17 | 0.44 | 0.60 |
| Frozen rule | 5 | 0.40 | 1.00 | 0.40 |
Accepted: T35 (2) coercion over money; T39 (2) spouse against spouse over infidelity; T05 (1) schemes and cons over money; T07 (1) courtship; T25 (1) trying to impress or win someone's favour. Rejected label-2: T16 newcomer seeks acceptance (S1 0.598), T17 proposal or engagement (S1 0.596), T30 run-in with the police (S1 0.604, S6 0.23).
Verdict
The selector is precise and narrow. Across both sets it accepted 10 clusters, and all 10 are situations or broad situational themes; none is a vocabulary, register or grab-bag cluster. Accepting everything gives about half that. The tightness filter used in the assumption-1 test is worse than accepting everything on the test set (0.44 against 0.51) — it keeps vocabulary clusters, which are tight.
It does not find most of the situations. Recall of label-2 clusters fell from 0.60 to 0.40, and precision for label 2 from 0.60 to 0.40 — the expected overfitting. The rejected label-2 clusters sit just under the S1 cut (0.596–0.604), so the threshold is doing real work at a place where situations and non-situations overlap. Five accepted clusters out of 41 is far short of the ~15 situations per stage a pack needs, so as it stands selection would make the sufficiency report fail on this corpus.
S6 earned its place on the test set, in the way it failed to on development. Two test clusters are held together by a single word — T10 "caliph" (S6 0.70) and T18 "young" (1.00) — and both have high cross-story cohesion (S1 0.628, 0.626). S1 alone would accept them; S6 rejects them. (Observation made after the test, not a re-tuning: S1 ≥ 0.6119 alone would accept 10 test clusters, 2 labelled 2 and 6 labelled ≥ 1.) Its original motivation, D18, was caught by S1 instead. Both failure modes exist, and each signal catches one.
Duplicates among accepted clusters: no two carry the same label name. T07 (courtship) and T25 (trying to impress or win someone's favour) overlap in substance; a reader could reasonably call them one situation. Automatic merging remains unsolved (see step 6).
What this means for the design
- Two-signal selection is a sound first gate, not a complete step. It should run as a filter that is allowed to reject — its precision is what matters for pack quality — followed by a second stage that recovers recall.
- Recall is the next problem, and clustering is the likely cause, not selection. The rejected situations (police run-ins, proposals, social climbing) are recognisable but diluted by off-situation members at k = 60. Candidates to test, all without inference: re-cluster the members of rejected clusters at finer k and re-apply the gate; select from the union of several partitions, using co-association to keep only members that consistently travel together; or grow an accepted cluster's core by adding nearest cross-story members while S1 stays above the cut.
- Merging duplicates needs member overlap, not centroid distance.
- These labels are one reader's. Every number above rests on them. They are in
probe/labels-dev.jsonandprobe/labels-test.json, with the sheets they were read from, and are worth a designer's spot-check — especially the 1-versus-0 boundary, which is the least certain. - The test is within one corpus. Generalisation to a different corpus is untested and needs embeddings, which needs the inference host.
Recovering recall (pre-registered 2026-09-14, before any recovery method was run)
The gate is precise and keeps too few. Three ways to find more situations, compared on fresh data, because the test set above is spent. Still no inference.
Fixed before running
- The gate does not change.
S1 ≥ 0.6119 AND S6 ≤ 0.1603, plus the candidate conditions (≥ 5 members, ≥ 4 stories, dominant ≤ 40%). "Passes the gate" below means all of these. - A property of the gate, noted in advance: S6 counts a word only if at least two members share it, so a cluster of fewer than 13 members fails whenever any content word repeats (2/12 = 0.167). In effect the gate requires ~13 members. The methods are built to produce clusters of that size; anything smaller is expected to fail, and that is the gate's behaviour, not a method's.
- Fresh partition: k = 60, seed 13 — neither development (11) nor test (12).
- Baseline (B): the gate applied to seed 13's candidates, as before.
- M1 — re-cluster the rejected. Every scene not in a B-accepted cluster is pooled and re-clustered with k-means, k = round(pool size / 20), seed 13. Result: B's accepted clusters plus pool clusters that pass the gate.
- M2 — consensus clustering. Each scene is represented by its co-association row over the 12 cached runs (k 50–100, seeds 21 and 22 — no overlap with 11, 12 or 13): how often it landed with every other scene. Those rows are clustered with k-means, k = 60, seed 13. Result: clusters that pass the gate.
- M3 — prune, then grow. For each seed-13 candidate: if it fails the gate, repeatedly remove the member with the lowest mean similarity to members from other stories, until it passes or falls to 13 members (then it is dropped). Then grow every passing cluster: add the non-member scene nearest its centroid while the result still passes; stop at the first addition that would fail, or at 30 members. Scenes may end up in more than one M3 cluster; the overlap is reported.
Evaluation
- Every cluster accepted by any method is pooled (identical member sets once), given a shuffled id, and dumped to a blind sheet showing only member summaries. Which method produced a cluster is written to a separate map file that is not read until labelling is finished.
- Labels use the same 2 / 1 / 0 scale. Names are reused verbatim from the development and test label files wherever the situation is the same, so duplicate situations can be counted by name.
- Per method: clusters accepted; precision for label ≥ 1 and for label 2; distinct situations (distinct names among label ≥ 1); scene coverage; overlap between accepted clusters.
- Success, per method: at least twice B's accepted clusters, and precision (≥ 1) of at least 0.80, and at least twice B's distinct situations. For scale: a two-stage pack wants ~30.
Known limits
The reader has now read every summary several times; labels are blind to method and signals, not to content. The methods and the gate are evaluated on one corpus.
Results (2026-09-14)
Seed 13 gave 42 candidates. 21 distinct accepted clusters were labelled blind (probe/labels-recover.json,
3 labelled 2, 13 labelled 1, 5 labelled 0), then scored by probe/evaluate_recover.py:
| Method | Accepted | Precision (≥ 1) | Precision (2) | Distinct situations | Scene coverage | Success |
|---|---|---|---|---|---|---|
| B, gate on candidates | 4 | 1.00 | 0.00 | 4 | 13% | — |
| M1, re-cluster the rejected | 4 | 1.00 | 0.00 | 4 | 13% | fail — 0 new |
| M2, consensus clustering | 3 | 1.00 | 1.00 | 3 | 5% | fail — too few |
| M3, prune then grow | 16 | 0.69 | 0.00 | 10 | 45% | fail — precision |
No method meets the pre-registered bar.
- M1 adds nothing. Re-clustering the 729 scenes outside B's clusters at k = 36 produced no cluster that passes the gate. The gate's selectivity does not come from dilution by one bad partition.
- M2 is the only method that produced sharp situations. All three of its clusters are labelled 2 — a wealthy suitor weighed against love, a spouse against a spouse over infidelity, coercion by threat — at 14–17 members, against B's 13–38. Consensus over many partitions isolates the scenes that genuinely travel together. It produces too few clusters to be the selection step on its own.
- M3 finds the most, and the gate cannot stop it diluting. Pruning recovered 12 rejected candidates and M3 reached 10 distinct situations (B: 4), but 5 of its 16 clusters are not situations. Growth was stopped by the 30-member cap in 10 of 16 clusters, not by the gate. That exposes the gate's second size effect: S6 is a share, so it gets easier to satisfy as a cluster grows, and S1 moves slowly. The gate calibrated on k = 60 clusters does not constrain a cluster that is being built up one scene at a time.
- Duplicates survive. Across methods, "a spouse against a spouse" appears three times (R006, R015) and courtship twice within M3 (R008, R010); 68 scenes sit in more than one M3 cluster.
- The seed-13 baseline itself accepted no label-2 cluster (0 of 4), where development and test each had 2–3. The baseline's precision for label 2 is unstable across partitions.
What this means
The gate, as frozen, is size-dependent in both directions: it structurally rejects clusters under ~13 members and barely constrains clusters above ~25. Any method that changes cluster size — pruning, growing, finer re-clustering — changes what the gate measures. Every recovery method was, in effect, testing the gate at sizes it was not calibrated for.
Observation only, not pre-registered and not a result: B and M2 together would give 7 clusters, 6 distinct situations, all labelled ≥ 1, and 3 labelled 2.
Directions this points to, none tested:
- Make the gate size-invariant before anything else. Replace S6's share with a count or a significance test against the corpus, and judge S1 relative to what a random cluster of the same size scores. Then re-validate — which needs fresh labels on fresh partitions.
- Use consensus (M2) to form clusters, not k-means. It was the only source of sharp situations. Running more co-association partitions (more seeds, more k) and extracting cores at several agreement levels may raise its yield without losing precision.
- Grow only with a size-invariant stopping rule. M3's recall is real; its stopping rule is not.
A size-invariant gate, G2 (pre-registered 2026-09-14, before any G2 signal was computed)
Why each old signal depended on size
- S6 is a share with a floor of 2 members: 2/size. Under ~13 members any repeated word fails the cut; above ~25 almost nothing does. A share is not evidence unless the count is improbable.
- S1's expectation does not depend on size — it is a mean over pairs — but its noise does: small clusters reach a high S1 by chance, which is part of why k-means' small clusters look tight.
- Nothing looked at individual members. Adding one off-situation scene to 30 moves S1 by roughly a thirtieth of the difference, so growth was never stopped.
G2 signals
| Signal | Definition | Why it is size-invariant |
|---|---|---|
| S1 | Unchanged: mean summary cosine over cross-story member pairs | Unbiased at any size |
| Z1 | (S1 − μₘ) / σₘ, where μₘ and σₘ come from random member sets of the same size m (≥ 4 stories, dominant ≤ 40%): 300 samples for m ≤ 40, 120 above, sizes 5–150, seeded | Requires more evidence where chance does more |
| MMIN | The lowest, over members, of a member's mean cosine to members of other stories | A per-member quantity; one diluting member fails it at any size |
| W | Among words shared by ≥ 2 members, those whose count is significant by a hypergeometric upper-tail test at Bonferroni 0.01 over the corpus vocabulary; W is the largest share among them (0 if none) | Counts only improbable sharing, then measures how dominant it is |
Word extraction, stopwords and the corpus are identical to signals.py. Pairwise similarities are
precomputed once. Candidate conditions (≥ 5 members, ≥ 4 stories, dominant ≤ 40%) still apply.
Choosing thresholds
- Tuning set: every labelled cluster so far — development (40), test (41) and the recovery pool (21): 102 clusters, 63 labelled ≥ 1, sizes 5–39. All have been seen; none is used for validation.
- Rule form: S1 ≥ a AND Z1 ≥ b AND MMIN ≥ c AND W < d, where each term may be switched off. Grid: a and c at the deciles of the tuning set's values or off; b ∈ {off, 1, 2, 3, 4, 5}; d ∈ {off, 0.25, 0.30, 0.35, 0.40, 0.50, 0.60}.
- Objective: the most tuning clusters labelled ≥ 1 accepted, subject to precision (≥ 1) of at least 0.85; ties on more label-2 accepted, then fewer active terms. The objective changes from the first gate (which maximised label-2 precision) because the failure now is recall, and label 1 — a situational theme — is usable material for a pack.
- The tuning report also gives, per size band (≤ 12, 13–25, ≥ 26), how many clusters the old and new gates accept and at what precision.
Validation on fresh data (nothing below has been computed)
-
V1, small clusters: k = 100, seed 15; 30 candidates sampled at random (seed 15).
-
V2, large clusters: k = 40, seed 14; 30 candidates sampled at random (seed 14).
-
V3, growth: k = 60, seed 16; every candidate G2 accepts is grown by adding the non-member nearest its centroid while the result still passes G2, up to 45 members. The stop reason — gate or cap — is recorded.
-
All V1 and V2 samples and all V3 grown clusters are pooled, deduplicated, shuffled and labelled blind, with the set and method in a separate map file.
-
Success, all three required:
- G2's precision (≥ 1) is at least 0.80 on V1 and on V2 separately, and G2 accepts at least 3 clusters in each.
- In V1, G2 accepts more small clusters than the old gate.
- In V3, most growths are stopped by the gate, not the cap, and grown clusters have precision (≥ 1) of at least 0.80.
The old gate is scored on V1 and V2 alongside, for comparison.
Frozen G2 (written after tuning, before any validation cluster was built)
probe/tune_gate2.py, log probe/tune-gate2.log. (Correction: the tuning set has 62 clusters
labelled ≥ 1, not 63.)
Accept a candidate when S1 ≥ 0.6037 AND Z1 ≥ 5 AND W < 0.60. MMIN is off.
| Tuning set, 102 clusters | Accepted | Label ≥ 1 | Precision (≥ 1) | Label 2 |
|---|---|---|---|---|
| Accept all | 102 | 62 | 0.61 | 13 |
| Old gate | 31 | 26 | 0.84 | 8 |
| G2 | 40 | 34 | 0.85 | 9 |
| Size band | Clusters (≥ 1) | Old gate: accepted / P1+ | G2: accepted / P1+ |
|---|---|---|---|
| ≤ 12 | 20 (7) | 0 / — | 1 / 1.00 |
| 13–25 | 52 (34) | 13 / 0.92 | 18 / 0.89 |
| ≥ 26 | 30 (21) | 18 / 0.78 | 21 / 0.81 |
Observations and predictions, recorded before validation:
- G2 is not size-invariant at the small end, and the objective chose that. Z1 ≥ 5 is a fixed bar on a statistic whose value, for a fixed true cohesion, grows with size (σₘ shrinks as m grows). It replaces S6's structural rejection of small clusters with a milder one: 1 small cluster accepted instead of 0. Prediction: criterion 2 passes narrowly or fails, and V1 yields fewer than 3 accepted clusters, failing criterion 1.
- MMIN, the signal designed to stop dilution, was not selected, and W's cut (0.60) is loose enough to act only on near-total single-word clusters. Growing a cluster raises Z1 and barely moves S1. Prediction: most V3 growths run to the cap, failing criterion 3.
- The gain over the old gate is real but modest, and it is in the mid and large bands: 8 more clusters labelled ≥ 1 accepted at unchanged precision.
- The thresholds are not changed in light of these predictions. The validation tests the rule the pre-registered procedure produced.
Validation results (2026-09-14)
73 fresh clusters labelled blind (probe/labels-validate.json: 6 labelled 2, 33 labelled 1, 34
labelled 0), scored by probe/validate_gate2.py score (log probe/validate-score.log).
| Set | Sampled (≥ 1) | Old gate: accepted / ≥ 1 / P1+ | G2: accepted / ≥ 1 / P1+ |
|---|---|---|---|
| V1, k = 100 | 30 (13) | 3 / 2 / 0.67 | 10 / 8 / 0.80 (2 labelled 2) |
| V2, k = 40 | 30 (16) | 1 / 1 / 1.00 | 4 / 3 / 0.75 |
| V3, grown | 13 | — | 13 grown: 11 stopped by the cap, 2 by the gate; 10 ≥ 1, P1+ 0.77 |
| Criterion | Result |
|---|---|
| 1. P1+ ≥ 0.80 and ≥ 3 accepted, on V1 and V2 | fail — V1 passes (0.80, 10), V2 does not (0.75, 4) |
| 2. G2 accepts more V1 clusters than the old gate | pass — 10 against 3 |
| 3. V3 mostly stopped by the gate, P1+ ≥ 0.80 | fail — 2 of 13, 0.77 |
| Overall | fail |
By size, across V1 and V2 together:
| Size | Old gate: accepted (≥ 1) | G2: accepted (≥ 1) |
|---|---|---|
| ≤ 9 | 0 | 0 — including two clusters labelled 2 (5 and 6 members) |
| 10–12 | 0 | 3 (3) |
| 13–25 | 3 (2) | 7 (5) |
| ≥ 26 | 1 (1) | 4 (3) |
| All | 4 (3), P1+ 0.75 | 14 (11), P1+ 0.79 |
Verdict on G2
G2 is a real improvement on the old gate, and it is not size-invariant. It accepts three to four times as many clusters at slightly better precision, and it opens the 10–25 member range the old gate all but closed. Both pre-registered predictions of failure held:
- The small end is still closed below 10 members. Z1 ≥ 5 is unreachable for a cluster of 5–9 however cohesive it is. The two sharpest small clusters on the sheet — a fare dispute (5 members, label 2) and a spouse against a spouse (6 members, label 2) — were both rejected.
- Growth is still unconstrained. Growing a cluster raises Z1 and barely moves S1, so both gates accept a 45-member cluster as readily as a 15-member one. The one per-member signal built to stop this, MMIN, was not selected by the tuning objective.
- Large clusters are still mostly rejected (4 of 30 in V2), including clear themes such as the saloon, dining rooms and illness.
What was learned about size-invariance itself: a fixed threshold on a z-score is not size-invariant; it reverses the old bias rather than removing it. A size-invariant cohesion test needs a threshold that scales with the null — for example a fixed percentile of the same-size random distribution combined with an effect-size floor on S1 — and a stopping rule for growth has to be a per-member test that the tuning cannot switch off. Both are untested proposals, and testing them needs another fresh partition and another round of blind labels, because every cluster labelled so far has now informed a design decision.
G3: a percentile of the same-size null, and a mandatory member floor (pre-registered 2026-09-14, before any G3 signal was computed)
Changes from G2
- Z1 is replaced by PCT: the fraction of same-size random member sets (≥ 4 stories, dominant ≤ 40%) whose S1 is below the cluster's. 1,000 samples per size for 5–60 members, 200 for 61–150, seeded. A percentile threshold asks "is this cluster as unusual as the top x% of random sets of its own size", which a z-score with a fixed bar does not.
- S1 floor is mandatory. At large sizes the null is narrow and PCT saturates at 1.0 for any cluster slightly above average, so an effect-size floor has to carry the weight there.
- MMIN is mandatory: every member's mean cosine to members of other stories must reach the floor. It is the growth stop; tuning chooses its value but cannot switch it off.
- W is unchanged from G2 and remains optional.
Tuning
- Tuning set: all 175 labelled clusters — development 40, test 41, recovery 21, G2 validation 73.
- Rule form: S1 ≥ a AND PCT ≥ p AND MMIN ≥ c AND W < d. a and c at the tuning set's deciles (no "off"); p ∈ {0.90, 0.95, 0.975, 0.99, 0.999}; d ∈ {off, 0.30, 0.40, 0.50, 0.60}.
- Objective unchanged from G2: most clusters labelled ≥ 1 accepted, subject to precision (≥ 1) of at least 0.85; ties on more label-2 accepted, then fewer active terms.
Risk named in advance: MMIN is a minimum over per-member means. Small clusters have noisier member means but fewer of them; large clusters have steadier means but more chances to contain one low member. Its net size effect is not predictable, and it could re-introduce a size bias of its own.
Validation on fresh partitions (seeds never used before)
- V0, very small: k = 150, seed 20; 25 candidates sampled (seed 20).
- V1, small: k = 100, seed 17; 20 sampled (seed 17).
- V2, large: k = 40, seed 18; 20 sampled (seed 18).
- V3, growth: k = 60, seed 19; every G3-accepted candidate grown toward its nearest scene while it passes G3, cap 45, stop reason recorded.
- Pooled, deduplicated, shuffled and labelled blind as before. The old gate and G2 are scored on the same clusters.
Success, all four required:
- G3 precision (≥ 1) ≥ 0.80 and ≥ 3 accepted, on V1 and on V2 separately.
- Small clusters: among V0 and V1 clusters of ≤ 9 members, G3 accepts at least 2, with precision (≥ 1) of at least 0.67, and more than G2 accepts.
- Growth: most V3 growths stopped by the gate, and grown clusters have precision (≥ 1) ≥ 0.80.
- G3 accepts at least as many clusters labelled ≥ 1 across V0–V2 as G2, at precision no lower than G2's.
Frozen G3 (written after tuning, before any validation cluster was built)
probe/tune_gate3.py, log probe/tune-gate3.log.
Accept a candidate when S1 ≥ 0.6175 AND PCT ≥ 0.90 AND MMIN ≥ 0.5296 AND W < 0.60.
| Tuning set, 175 clusters (101 ≥ 1, 19 label 2) | Accepted | Label ≥ 1 | Precision (≥ 1) | Label 2 |
|---|---|---|---|---|
| Old gate | 47 | 38 | 0.81 | 9 |
| G2 | 67 | 55 | 0.82 | 12 |
| G3 | 47 | 40 | 0.85 | 13 |
| Size band | Clusters (≥ 1) | Old: accepted/≥ 1 | G2 | G3 |
|---|---|---|---|---|
| ≤ 9 | 26 (5) | 0/0 | 0/0 | 4/3 |
| 10–12 | 21 (11) | 0/0 | 4/4 | 4/4 |
| 13–25 | 75 (47) | 17/15 | 26/22 | 17/17 |
| ≥ 26 | 53 (38) | 30/23 | 37/29 | 22/16 |
Observations and predictions, recorded before validation:
- G3 is the first gate to accept any cluster under 10 members (4 on tuning, 3 of them ≥ 1). Prediction: criterion 2 passes.
- The percentile threshold settled at its loosest value, 0.90, and the S1 floor rose from G2's 0.6037 to 0.6175. The floor, not the percentile, is doing most of the work at mid and large sizes, and G3 accepts 15 fewer large clusters than G2 on tuning. Prediction: criterion 4 fails — G3 finds fewer clusters labelled ≥ 1 than G2 (40 against 55 on tuning), though at higher precision. Criterion 1 is at risk on V2 for the same reason.
- MMIN at 0.53 is the growth stop. Whether a nearest-to-centroid scene falls below it is not predictable from tuning, which contains no grown-under-G3 clusters. No prediction for criterion 3.
- The trade is visible already: G3 exchanges recall at the large end for precision everywhere and for admission at the small end. Whether that is the right trade for a pack is a design question the validation cannot answer.
Validation results (2026-09-14)
75 fresh clusters labelled blind (probe/labels-validate3.json: 4 labelled 2, 24 labelled 1, 47
labelled 0), scored by probe/validate_gate3.py score (log probe/validate3-score.log).
| Set | Sampled (≥ 1) | Old: acc / ≥ 1 / P1+ | G2: acc / ≥ 1 / P1+ | G3: acc / ≥ 1 / P1+ |
|---|---|---|---|---|
| V0, k = 150 | 25 (5) | 2 / 1 / 0.50 | 8 / 4 / 0.50 | 11 / 4 / 0.36 |
| V1, k = 100 | 20 (5) | 1 / 0 / 0.00 | 1 / 1 / 1.00 | 4 / 2 / 0.50 |
| V2, k = 40 | 20 (12) | 0 / 0 / — | 4 / 4 / 1.00 | 2 / 2 / 1.00 |
| ≤ 9 members, V0 + V1 | 27 (4) | 1 / 0 / 0.00 | 3 / 1 / 0.33 | 11 / 3 / 0.27 |
| V0–V2 pooled | 65 (22) | 3 / 1 / 0.33 | 13 / 9 / 0.69 | 17 / 8 / 0.47 |
| V3, grown under G3 | 10 | — | — | 7 stopped by the cap, 3 by the gate; P1+ 0.60 |
All four criteria fail, and G3 is worse than G2 on fresh data: more clusters accepted, fewer of them situations. G3's gain on the tuning set (P1+ 0.85, 13 label-2) did not survive a new partition.
Why G3 failed — measured, not inferred
- PCT carries no information. Across all 175 tuning clusters, 173 reach PCT ≥ 0.90 and 141 reach 1.0 — including 72 of the 74 clusters labelled 0. The random-set null is the wrong null. k-means groups each scene with its nearest neighbours, so every cluster it produces is far tighter than a random set of scenes of the same size, whether or not it is a situation. A cohesion test has to ask whether a cluster is tighter than other clusters of its size, not than random scenes.
- MMIN barely separates labels and favours small clusters. Median MMIN is 0.536 for label 0, 0.543 for label 1 and 0.564 for label 2; by size it is 0.547 for ≤ 9 members, 0.537 for 10–25 and 0.552 for ≥ 26. A minimum over few, noisy per-member means is not reliably lower for small clusters, so the floor admitted 11 small clusters of which 8 are not situations. This is the risk recorded before tuning.
- Growth is still not stopped by the gate (7 of 10 to the cap). A nearest-to-centroid scene typically clears a member floor calibrated on whole clusters.
Where this leaves selection
Across three gates and five blind-labelled evaluations, G2 remains the best: pooled precision (≥ 1) of 0.79 on its own validation and 0.69 here, with the fewest non-situations admitted. None of the three is size-invariant, and none stops growth.
Two conclusions look robust enough to build on:
- Cohesion alone does not identify a situation. Clusters labelled 0, 1 and 2 overlap heavily on every cohesion statistic tried (S1, Z1, PCT, MMIN). What separates the best-labelled clusters is visible to a reader — one transaction — and is not a geometric property of summary embeddings at the resolution tried.
- The labelling budget is the binding constraint. Each gate needs fresh partitions and 70–80 new blind labels to test honestly, and every labelled cluster so far has informed a design. Continuing to iterate gate designs on this corpus without a different kind of evidence is unlikely to converge.
Candidate directions, none tested: a null built from k-means clusters of the same size on shuffled-story data rather than random scenes; a selection step that uses the inference model to name each cluster's transaction and checks agreement across members (needs the GPU); or accepting G2 as the gate and moving the problem downstream to event construction, where a human pass over ~40 candidate clusters is cheap compared with the labelling this probe has already required.
G4: the model reads the cluster (pre-registered 2026-09-14, before any model call)
The finding above is that what makes a cluster a situation is visible to a reader and not to the geometry. So let a reader decide — the inference model — and measure it against the blind labels.
Method
- For each cluster, up to 15 member summaries (a seeded sample when larger), shuffled and numbered, go to the model in one call with the prompt below. It returns JSON: the single social transaction most of them clearly show, or NONE, and the numbers of the summaries that clearly show it.
- FIT = (number of valid summary numbers listed) / (number shown); 0 when the answer is NONE or the reply cannot be parsed. Parse failures are counted and reported.
- Gate: accept when FIT ≥ θ, plus the usual candidate conditions (≥ 5 members, ≥ 4 stories, dominant ≤ 40%), which every labelled cluster already meets.
- Deterministic: temperature 0, 2,048-token context,
format: json. One request at a time, as before. - The prompt contains no example of a situation — the 3B model copied the examples it was given in the summarisation pilot.
Prompt (system), fixed:
You read short summaries of scenes from different stories and decide whether they share one kind of social situation.
A social situation is a specific transaction between people: who wants what from whom, and what is at stake. Describe it in your own words.
Rules:
- Find the single situation that the largest number of summaries clearly show.
- It is normal for many summaries not to fit. Include a summary only if it clearly shows that situation.
- A shared place, mood, profession or word is not a situation.
- If no situation is clearly shown by at least three summaries, the situation is NONE.
Reply with JSON only: {"situation": "<at most 10 words, or NONE>", "fits": []}
The user message is the numbered list. For qwen3:14b it ends with /no_think, and any think block
is stripped before parsing.
Models
- Primary:
qwen2.5:3b-instruct— the design requires a small local model; the verdict rests on it. - Secondary:
qwen3:14b— reported, to show whether capability is what limits the result.
Data — no new labels
The model is never fitted to labels; the only fitted quantity is θ.
- θ is chosen on development (40 clusters) only, with the objective used for G2 and G3: most clusters labelled ≥ 1 accepted, subject to precision (≥ 1) ≥ 0.85; ties on more label-2 accepted, then the higher θ. Grid: θ ∈ {0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0}.
- Primary evaluation: the G2 and G3 validation sets, 148 clusters, which tuned neither G2 nor G4, so G2 can be compared fairly on them. Secondary: test and recovery, 62 clusters.
- Identical member sets appearing in more than one file are judged once.
- Pilot, allowed once: the first 5 development clusters, to check the JSON parses. If more than 1 of 5 fails to parse, one prompt revision is allowed and recorded here before any other call.
Success, on the 148-cluster primary set, with the 3B model
- Precision (≥ 1) at least 0.80.
- More clusters labelled ≥ 1 accepted than G2 accepts on the same set.
- Among accepted clusters of ≤ 9 members, precision (≥ 1) at least 0.67, if there are at least 3.
Also reported: G4 AND G2; FIT distribution by label; the situations named for accepted clusters against the label names; and the 14B results against the same bar.
Pilot (21:23): 5 of 5 replies parsed with the 3B model, so the prompt is unchanged. About 1.6 s per cluster. Noted without drawing a conclusion from five: the situations named are often generic ("Confrontation and hidden motives" for D04, a label-2 spouse-against-spouse cluster, FIT 0.47), and a label-1 argument cluster scored FIT 0.87.
Results, qwen2.5:3b-instruct (21:24–21:29, 245 distinct clusters in 286 s, 1 parse failure)
Log probe/evaluate-judge-3b.log. θ chosen on development: 0.6 (9 accepted, P1+ 0.89).
| Primary set, 148 clusters (67 ≥ 1) | Accepted | ≥ 1 | Label 2 | P1+ | Recall |
|---|---|---|---|---|---|
| Accept all | 148 | 67 | 10 | 0.45 | 1.00 |
| G2 | 49 | 36 | 5 | 0.73 | 0.54 |
| G4, FIT ≥ 0.6 | 50 | 24 | 4 | 0.48 | 0.36 |
| G4 AND G2 | 18 | 11 | 1 | 0.61 | 0.16 |
On the secondary set (62) G4 reaches P1+ 0.71 against G2's 0.82. All three criteria fail (P1+ 0.48; 24 ≥ 1 against G2's 36; 13 accepted small clusters at P1+ 0.23).
FIT does not separate the labels: mean 0.39 for label 0, 0.49 for label 1 — and 0.42 for label 2, lower than for label 1. Precision at the primary set (0.48) is barely above accepting everything (0.45).
Why, visible in the replies: the 3B model names situations at a level of abstraction where almost anything fits — "confrontation", "power dynamics in social hierarchies", "tension between parties", "coercion and manipulation" — and then lists most summaries as fitting it. A 5-member cluster labelled 0 scored FIT 1.00 as "power struggle/conflict". When it is specific it is often right ("marital conflict" for a spouse-against-spouse cluster, "con or scam", "argue over payment"), but specificity is not what FIT rewards. The rule against shared place, mood and profession was followed; a rule against abstraction was never written, and the model supplies abstraction freely.
14B pilot and the one allowed revision
The qwen3:14b pilot failed to parse 5 of 5: every reply's content was empty. The model spent its
whole 200-token budget in a separate thinking channel (Ollama returns it outside content), and the
/no_think suffix did not stop it. Revision, the one allowed by the pilot rule, affecting only the
14B run: requests set Ollama's think: false and allow 300 output tokens. The prompt text,
sampling, θ procedure and success criteria are unchanged. The five empty results were deleted before
re-running the pilot.
Re-run pilot: 5 of 5 parsed, about 2.4 s per cluster. The situations named are more specific than the 3B model's ("Wealthy person pressures someone into marriage", "Struggling artist facing societal or personal challenges"), though an abstract one still scored FIT 0.80 ("conflict between individuals with differing perspectives").
Results, qwen3:14b (21:31–21:45, 245 clusters in 832 s, 0 parse failures)
Log probe/evaluate-judge-14b.log. θ chosen on development: 0.6 (8 accepted, P1+ 0.88).
| Primary set, 148 clusters (67 ≥ 1) | Accepted | ≥ 1 | Label 2 | P1+ | Recall |
|---|---|---|---|---|---|
| Accept all | 148 | 67 | 10 | 0.45 | 1.00 |
| G2 | 49 | 36 | 5 | 0.73 | 0.54 |
| G4 (14B), FIT ≥ 0.6 | 42 | 22 | 4 | 0.52 | 0.33 |
| G4 (14B) AND G2 | 10 | 10 | 1 | 1.00 | 0.15 |
By size, G4 (14B) on the primary set: ≤ 9 members 17 accepted, P1+ 0.29; 10–25, 14 accepted, P1+ 0.50; ≥ 26, 11 accepted, P1+ 0.91. Secondary set (62): G4 P1+ 0.60 against G2's 0.82.
All three criteria fail with the 14B model as well (P1+ 0.52; 22 ≥ 1 against G2's 36; 17 small clusters accepted at P1+ 0.29).
FIT now rises with the label — mean 0.40, 0.47, 0.52 for labels 0, 1, 2 — where the 3B model's did not, so capability helps. But the separation is small next to the spread, and the 14B model still names abstractions that fit almost anything ("Someone wants something from someone else", "Power dynamics between individuals", "Someone holds power over someone else"), most often for small clusters.
GPU host during both runs: no faults logged; peak 284.9 W (1 s), mean 163 W under load, 65 °C, PCIe
Gen 3 x4 throughout. Logs in probe/gpu-logs/*2122*.
Verdict on G4
Neither model's judgement is a better gate than G2, and the pre-registered verdict (on the 3B model) is a fail. The failure is specific and repeatable: asked for "the single situation most summaries share", a model answers at whatever level of abstraction makes the most summaries fit, and FIT rewards exactly that. Small clusters make it worse, because with 5–9 summaries an abstraction that covers four or five of them is always available.
Two observations, recorded as observations because neither was a pre-registered criterion:
- G4 (14B) AND G2 accepted 10 clusters on the primary set, and all 10 are labelled ≥ 1. The two gates fail in different ways — G2 admits cohesive clusters that are not situations, the model admits abstract "situations" that are not cohesive — and their intersection is precise. It is also narrow: recall 0.15.
- For clusters of 26 or more, the 14B model's precision is 0.91. The abstraction problem is a small-cluster problem: 15 shown summaries rarely all fit one vague phrase and get listed.
What would test the obvious next step, none of it done: a prompt that asks for the situation in terms of roles and a concrete want and rejects answers without both; scoring the specificity of the named situation (for example, requiring the named situation to be closer to the member summaries than to the corpus average in embedding space); and G2 AND G4 evaluated as a pre-registered gate on a fresh partition, since every labelled cluster now informs this observation.
Predicaments (recorded 2026-09-15, before any model call)
What the first real review found
The designer opened REVIEW.html for run k60-s24 and read four groups closely:
- Group 1 is artists. It holds at least three different situations: an artist in a new place using the art to get on, being paid for the work against what the work is worth, and plain money trouble.
- Group 2 is police and court officers. Inside it: being coerced, with the police watching or arriving; and a crime that is stopped or found out.
- Group 6 is flirting and infatuation.
- Extra group 1 is room and board: hotels, dinners and the like.
None of these is "one person wanting something specific from another". Checked across all 20 groups, most are held together by who is in the scenes or where they happen: artist in 44% of group 1's scenes, policeman in 54% of group 3's, uncle and lawyer in group 4, marriage and wedding in group 8, restaurant and customers in group 12, hotel in 41% of extra group 1.
Why every earlier test missed it. A topic group is as cohesive and as spread across stories as a situation group, so cohesion, story spread, word dominance and the model's own judgement all passed it. Summaries keep jobs and places, and the embeddings group on them. It took a person reading the groups.
Decisions (designer, 2026-09-15)
- The unit is a predicament, not a two-person transaction: what the main person is up against and what they want out of it, as The Ladder's existing events are (the rent bouncing, the mail cart). The sub-themes the designer found inside groups 1 and 2 are this kind of thing.
- Re-describe every scene as a predicament, then regroup, rather than splitting topic groups by hand.
Method, fixed before any call
probe/redescribe.pygives the model each scene's own prose (not its summary) and the prompt in that script: one sentence, at most 20 words, no names, jobs, places, businesses or story-specific objects, saying what the person is up against and what they want. No example of the output vocabulary.- Names are then removed deterministically (
strip_names.py), the sentences embedded (embed.py), and the embeddings clustered at k = 60. - Pilot first: the first 40 scenes, with both
qwen2.5:3b-instructandqwen3:14b, read by the designer side by side before any full run. The full run uses the model the designer prefers.
How it is judged
- Topic share (
probe/topic_share.py): for each candidate group, the share of its scenes whose original summary contains the group's most common job, relationship or place word; median over groups and seeds 11–13 at k = 60. The baseline on the summary embeddings is measured before any call and recorded below. Pass: the median falls to half the baseline or less, with at least 30 candidate groups remaining at k = 60 (enough material for a two-stage pack). - Leakage: how many predicament sentences still contain a job, relationship or place word. Reported, not gated.
- The real test is the designer reading the regrouped top groups in the review page: most should read as one predicament. Numbers alone passed topic groups before.
Baseline (measured 2026-09-15, before any model call)
A first measurement counted words for a person with no role at all ("man", "woman", "girl"), and scored
one group 90% topic-bound for saying "man". Those words (GENERIC_PEOPLE in probe/topic_words.py)
were taken out of the measure, and the baseline re-measured, still before any predicament was generated.
python3 topic_share.py summary_embeddings.json (probe/topic-share-baseline.log):
| Seed | Candidate groups | Median top topic-word share |
|---|---|---|
| 11 | 40 | 16% |
| 12 | 41 | 20% |
| 13 | 42 | 22% |
| Median | 41 | 20% |
Most topic-bound at seed 11: wife 47% of 15, dinner 47% of 15, policeman 40% of 15 and 38% of 16, brother 38% of 8, artist 36% of 11. Pass mark for the predicament groups: median 10% or less, with at least 30 candidate groups.
Pilot (2026-09-15)
First 40 scenes, one request at a time on the GPU host (3B 39 s, 14B 111 s; no GPU faults; logs in
probe/gpu-logs/*-2026-09-15-1803*). Read side by side in probe/pilot-compare.html (pilot_page.py).
| qwen2.5:3b-instruct | qwen3:14b | |
|---|---|---|
| Has a job, relationship or place word | 15 of 40 | 13 of 40 (mostly wife, husband) |
| Mean words (limit 20) | 14 | 25; 34 of 40 over |
| Most common opening | "The Kid seeks…" ×3 | "A man struggles…" ×6 |
3B mostly retells the plot, keeping names, objects and places. 14B writes something close to a predicament but ignores the length limit and repeats a few openings. Designer's read: scenes 1–10, 14B dramatically better in every one; stopped there. The full run uses qwen3:14b.
Found while reading: split.py skips a title written in quotation marks, so three stories were
merged into the story before them. "Little Speck in Garnered Fruit" is inside "Dougherty's
Eye-Opener", "The Guilty Party" inside "A Harlem Tragedy", and "What You Want" inside "The Duel". The
corpus has 97 stories, not 94. Separately, strip_names.py treats York, Manhattan and Broadway
as personal names, which gives "New someone" in 29 of 838 summaries.
Corrections before the full run (2026-09-15, designer's go-ahead)
The v1 data and every G1–G4 record stay as they were, in probe/. The corrected corpus is built in
probe/v2/, running the same scripts from that directory.
- Split.
split.pyaccepts a quoted title and matches a contents entry by the part before a dash (“THE GUILTY PARTY”—AN EAST SIDE TRAGEDY). Result: 97 stories, 839 scenes. Exactly the three missing stories were added, nothing was lost, and only their three host stories changed. - Summaries carried over. 829 of the 839 scenes have text identical to a v1 scene and keep that
v1 summary verbatim (
carry_summaries.py). The other 10 are summarised with the v1 model and prompt. - Name stripper. New York places (New York, Manhattan, Broadway, Harlem, the Bowery, and a few others) become "the city" before names are looked for, and days and holidays are never names. Months stay ambiguous ("May" is a person in scene 5) and are left alone.
- The baseline is re-measured on v2 (the v2 summaries, re-embedded), because both fixes change what is embedded. It is recorded below before any predicament is read.
- No trimming to 20 words. Tried on the 14B pilot: cutting at the last clause break within 20 words removed what the person is up against in most cut sentences ("…with a promise of future money" without "but his dishonesty sparks a violent confrontation"), and 10 of 40 had no clean break. The full 14B sentence is embedded.
Result (2026-09-15)
Full run: 839 scenes re-described by qwen3:14b, one request at a time, 37 minutes on the GPU host
(peak 270 W, mean 136 W, 65 °C, no faults; logs in probe/gpu-logs/*-2026-09-15-1831*). Sentences
average 25 words. 219 of 839 still contain a job, relationship or place word, most often a relationship
(wife, husband) — reported, not gated.
topic_share.py, k = 60, seeds 11–13, both measured on the v2 corpus:
| Candidate groups | Median top topic-word share | |
|---|---|---|
v2 summaries (baseline, topic-share-baseline-v2.log) |
45 | 20% |
Predicaments (topic-share-predicaments.log) |
42 | 13% |
The pre-registered pass mark is not met. It asked for half the baseline or less (10%); the measure fell by about a third, from 20% to 13%. The group-count mark is met (42 ≥ 30). The most topic-bound groups are also much weaker than the baseline's: the worst predicament group is room at 40% of 5 scenes, against the baseline's bride at 62% of 8 and policeman at 60% of 20.
Not a pass and not a wash, so the decision goes to the reading test that this whole step exists for:
groups_page.py builds v2/groups-predicaments-k60-s24.html, which lists each group's predicament
sentences with each scene's original summary beneath, most cohesive group first, and asks the designer
to mark each group one predicament, two or three mixed, or not a predicament. The openings
repeat ("A man must" 45 times, "A man struggles" 37), so the risk to watch for is a group held together
by that phrasing rather than by a predicament.