Files
TheLadder/tools/story-to-pack/SELECTION.md
T
JesseMarkowitzandClaude Opus 5 fa3769d0fe Add story-to-pack research and a structured situation review
Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
2026-09-15 06:59:35 -04:00

745 lines
46 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# story-to-pack — selecting situations
> **Decision (2026-09-15): selection is gate G2 followed by a structured human review.** `probe/review.py`
> writes a review package (`REVIEW.md`, `review.json`, `candidates.json`); the reviewer fills in
> `review.json`; `probe/apply_review.py` checks it and writes `situations.json` for the next stage. See
> the README. The record below is how that decision was reached: four gate designs and two
> model-judged gates, each tested against blind labels.
The step after clustering: from the clusters of summarised scenes, choose the ones that are genuine
**recurring situations** — the raw material of a recurring event — and reject the rest. Companion to
`DESIGN.md`; this is the step its assumption-1 result called for.
**Status (2026-09-14): designed and pre-registered below, before any feature was computed or any
label was read against a feature.** Results are appended at the end, and nothing above the Results
heading is edited after the test runs, except to fix a typo.
---
## The problem, from the assumption-1 result
Summarise-then-embed broke the story-identity effect: clusters now draw from many stories. But of five
sampled tight-recurring clusters, one was a **vocabulary cluster** — fortune-tellers, adventurers, a
found tyre, a graft inquiry — held together by "seeks", "fortune", "mysterious", and mostly from three
stories. The concentration and tightness filters in `measure.py` let it through. A selector has to
reject clusters like that one and keep clusters like "a spouse schemes against or confronts a spouse
over infidelity".
## Constraints
- **No inference.** Selection uses only what is already on disk: the summary embeddings, the prose
embeddings, the summary text, and which story each scene came from. It is pure computation, like the
numeric search, so it can be re-run and tuned for free. (It also keeps the inference host free for
other work.)
- **Deterministic.** Seeded clustering; the same inputs give the same selection.
- **Pure Python**, like the rest of the probe.
## What a selected situation is
A set of scenes, from **at least 4 different stories**, with no story contributing more than 40%,
that a reader can name as **one social transaction**: who wants what from whom. It becomes one
recurring event; its member scenes become that event's description variants and option material.
Rejected: story-dominated clusters (already mechanical), **vocabulary clusters** (shared words, no
shared transaction), **register clusters** (shared tone or setting, e.g. "men walking through the
city"), and **near-duplicates** of an already-selected situation.
## Candidate signals
All computed per cluster, from existing vectors and text. Each carries a pre-registered prediction of
which way it moves for a genuine situation.
| # | Signal | Definition | Prediction for a situation |
| --- | --- | --- | --- |
| S1 | **Cross-story cohesion** | Mean summary-embedding cosine over member pairs from *different* stories | Higher |
| S2 | **Cross/within ratio** | S1 divided by the mean cosine over member pairs from the *same* story (1.0 if there are none) | Closer to 1 — a situation is as similar across stories as within one; a cluster held together by a few stories' scenes is not |
| S3 | **Effective stories** | Inverse Simpson index of the member story counts, divided by cluster size | Higher — members spread evenly rather than a few stories plus stragglers |
| S4 | **Prose agreement** | Mean *prose*-embedding cosine over cross-story member pairs, minus the corpus-wide mean cross-story prose cosine | Higher — a second, independent view: scenes of the same situation should read a little alike even in the original text |
| S5 | **Stability** | Mean pairwise co-association of members across 12 k-means runs (k from 50 to 100, several seeds): the share of runs in which two members land together | Higher |
| S6 | **Word dependence** | Share of members containing the cluster's single most over-represented content word (log-odds against the whole summary set, stopwords and "someone" excluded) | Higher for a **vocabulary** cluster — one word doing the work |
S6 is predicted to indicate the failure, the others to indicate success.
## Protocol
1. **Candidates.** Cluster the summary embeddings with k = 60. A candidate is a cluster with ≥ 5
members, ≥ 4 distinct stories and dominant share ≤ 40% (`measure.py`'s "recurring"; the relative
tightness filter is *not* applied — the selector replaces it).
2. **Two candidate sets from two partitions.** Development: seed 11. Test: seed 12.
3. **Blind labels.** Each set is dumped to a sheet showing, per cluster, only its member summaries in
shuffled order — no story titles, no scores, no sizes beyond the lines themselves. Each cluster is
labelled by reading:
- **2 — one situation:** one nameable social transaction that most members (roughly 70%+)
instantiate. The label records its name.
- **1 — broad:** a shared setting or theme, but the transaction varies across members.
- **0 — not a situation:** vocabulary, register, or nothing nameable.
The test sheet is labelled before any signal is computed for the test set.
4. **Choose the rule on development only.** Compute S1–S6 for the development set. Candidate rules
are single thresholds on one signal, or an AND of two. The chosen rule is the one with the highest
**precision** (share of accepted clusters labelled 2), subject to accepting at least half of the
development set's label-2 clusters; ties go to the simpler rule. Freeze it here, in writing, before
step 5.
*Detail fixed after the labels were written and before any signal was computed:* thresholds are
placed midway between adjacent observed development values, never on one. Ties beyond simplicity
break on more label-2 clusters accepted, then on precision for label ≥ 1. Two reference points
are reported but cannot be chosen: accepting every candidate, and `measure.py`'s tightness filter
(coherence at or above the partition's median sized cluster), called B0.
5. **Test once.** Apply the frozen rule to the test set. Report precision (label 2, and label ≥ 1),
recall of label-2 clusters, and the number accepted. No re-tuning after seeing test results; any
later change is a new rule and needs new data to be tested on.
6. **Near-duplicates.** Among accepted clusters, merge any pair whose centroids have cosine ≥ a
threshold chosen on development by reading the labels' names. Report how many merges happened on
the test set and whether each merged pair's names agree.
## Known limits of this test
- **The labels are one reader's judgement** — the model running the probe. They are written to a
file so the designer can spot-check or overrule them, and the verdict should be read with that in
mind.
- **Development and test come from the same corpus.** Two partitions of the same 838 scenes share
situations, so the test measures stability across clustering, not generalisation to a new corpus.
A genuinely held-out test needs a second corpus, which needs embedding, which needs the inference
host.
- **k is fixed at 60** for the test. The selector's behaviour at other k is not tested.
## Frozen rule (written before any test signal was computed)
Chosen on development (40 candidates, 5 labelled 2; must accept ≥ 3), by `probe/select_rule.py choose`:
> **Accept a candidate when S1 ≥ 0.6119 AND S6 ≤ 0.1603.**
> Cross-story cohesion at least 0.6119, and no single over-represented word in more than 16% of members.
| On development | Accepted | Precision (2) | Precision (≥ 1) | Recall (2) |
| --- | --- | --- | --- | --- |
| Accept every candidate | 40 | 0.12 | 0.62 | 1.00 |
| B0, tightness filter | 17 | 0.24 | 0.76 | 0.80 |
| S1 alone (best single) | 9 | 0.44 | 0.89 | 0.80 |
| **Frozen rule** | **5** | **0.60** | **1.00** | **0.60** |
Accepted on development: D04 (2), D09 (1), D13 (2), D16 (2), D33 (1). Rejected label-2: D15 (S1 0.603),
D30 (S6 0.20). Full leaderboard: `probe/choose-dev.log`; signals: `probe/signals-dev.json`.
Observations recorded now, so they cannot be shaped by the test:
- **This is very likely overfit.** Five positives, fine-grained thresholds, and an AND of two signals
is exactly the setting in which a rule memorises its development set. The test is what decides.
- **S1 does most of the work.** It is the best single signal by a wide margin, and it alone rejects
the vocabulary cluster that motivated this step (D18, S1 0.605).
- **S6 did not behave as predicted.** Its most over-represented word is usually a *story's* word
("juggler", "furs", "eviction"), not generic vocabulary, and D18's is low (0.11). It catches D39
("young", 1.00) — a vocabulary cluster of a different kind. Its contribution to the frozen rule is
a single cut, and may not carry over.
- **S2, S3 and S4 were weak on development**, and S5 middling. None of the predicted directions was
strongly wrong, but only S1 separated well.
**Near-duplicates (step 6): centroid cosine cannot do it, so no automatic merge is applied.** Among
the five clusters the rule accepts on development, no two share a situation, so there was nothing
there to calibrate on. Across all 26 development candidates labelled ≥ 1, the one clear duplicate —
D13 and D30, both "a run-in with the police" — has centroid cosine **0.9026**, while five pairs of
plainly different situations score higher, up to **0.9326** ("money and marriage" against "schemes,
cons and theft"). D05/D23 (artists) sit at 0.9005. Any threshold that merges the duplicate merges
different situations first. Summary embeddings place situations about money, romance and conflict
in one dense region, so closeness between centroids is not identity of situation.
Decision, fixed before the test: the test reports, **by label name**, whether any two accepted
clusters are the same situation, as a measurement of how often duplicates survive selection. A
merge method is an open problem; member overlap across partitions (co-association between two
clusters' members) is the candidate to try next, and is untested.
## Results
**Test run once, 2026-09-14,** on the seed-12 partition (41 candidates, 5 labelled 2), frozen rule
unchanged. Logs: `probe/signals-test.json`, `probe/apply-test.log`.
| On test | Accepted | Precision (2) | Precision (≥ 1) | Recall (2) |
| --- | --- | --- | --- | --- |
| Accept every candidate | 41 | 0.12 | 0.51 | 1.00 |
| B0, tightness filter | 18 | 0.17 | 0.44 | 0.60 |
| **Frozen rule** | **5** | **0.40** | **1.00** | **0.40** |
Accepted: T35 (2) coercion over money; T39 (2) spouse against spouse over infidelity; T05 (1) schemes
and cons over money; T07 (1) courtship; T25 (1) trying to impress or win someone's favour.
Rejected label-2: T16 newcomer seeks acceptance (S1 0.598), T17 proposal or engagement (S1 0.596),
T30 run-in with the police (S1 0.604, S6 0.23).
### Verdict
**The selector is precise and narrow.** Across both sets it accepted 10 clusters, and **all 10 are
situations or broad situational themes; none is a vocabulary, register or grab-bag cluster.**
Accepting everything gives about half that. The tightness filter used in the assumption-1 test is
*worse* than accepting everything on the test set (0.44 against 0.51) — it keeps vocabulary clusters,
which are tight.
**It does not find most of the situations.** Recall of label-2 clusters fell from 0.60 to 0.40, and
precision for label 2 from 0.60 to 0.40 — the expected overfitting. The rejected label-2 clusters sit
just under the S1 cut (0.596–0.604), so the threshold is doing real work at a place where situations
and non-situations overlap. **Five accepted clusters out of 41 is far short of the ~15 situations
per stage a pack needs**, so as it stands selection would make the sufficiency report fail on this
corpus.
**S6 earned its place on the test set, in the way it failed to on development.** Two test clusters are
held together by a single word — T10 "caliph" (S6 0.70) and T18 "young" (1.00) — and both have high
cross-story cohesion (S1 0.628, 0.626). S1 alone would accept them; S6 rejects them. (Observation made
after the test, not a re-tuning: S1 ≥ 0.6119 alone would accept 10 test clusters, 2 labelled 2 and 6 labelled ≥ 1.)
Its original motivation, D18, was caught by S1 instead. Both failure modes exist, and each signal
catches one.
**Duplicates among accepted clusters:** no two carry the same label name. T07 (courtship) and T25
(trying to impress or win someone's favour) overlap in substance; a reader could reasonably call them
one situation. Automatic merging remains unsolved (see step 6).
### What this means for the design
1. **Two-signal selection is a sound first gate, not a complete step.** It should run as a *filter
that is allowed to reject* — its precision is what matters for pack quality — followed by a
second stage that recovers recall.
2. **Recall is the next problem, and clustering is the likely cause, not selection.** The rejected
situations (police run-ins, proposals, social climbing) are recognisable but diluted by
off-situation members at k = 60. Candidates to test, all without inference: re-cluster the members
of rejected clusters at finer k and re-apply the gate; select from the union of several partitions,
using co-association to keep only members that consistently travel together; or grow an accepted
cluster's core by adding nearest cross-story members while S1 stays above the cut.
3. **Merging duplicates needs member overlap, not centroid distance.**
4. **These labels are one reader's.** Every number above rests on them. They are in
`probe/labels-dev.json` and `probe/labels-test.json`, with the sheets they were read from, and are
worth a designer's spot-check — especially the 1-versus-0 boundary, which is the least certain.
5. **The test is within one corpus.** Generalisation to a different corpus is untested and needs
embeddings, which needs the inference host.
---
## Recovering recall (pre-registered 2026-09-14, before any recovery method was run)
The gate is precise and keeps too few. Three ways to find more situations, compared on **fresh data**,
because the test set above is spent. Still no inference.
### Fixed before running
- **The gate does not change.** `S1 ≥ 0.6119 AND S6 ≤ 0.1603`, plus the candidate conditions (≥ 5
members, ≥ 4 stories, dominant ≤ 40%). "Passes the gate" below means all of these.
- **A property of the gate, noted in advance:** S6 counts a word only if at least two members share
it, so a cluster of fewer than 13 members fails whenever any content word repeats (2/12 = 0.167). In
effect the gate requires ~13 members. The methods are built to produce clusters of that size;
anything smaller is expected to fail, and that is the gate's behaviour, not a method's.
- **Fresh partition:** k = 60, **seed 13** — neither development (11) nor test (12).
- **Baseline (B):** the gate applied to seed 13's candidates, as before.
- **M1 — re-cluster the rejected.** Every scene not in a B-accepted cluster is pooled and re-clustered
with k-means, k = round(pool size / 20), seed 13. Result: B's accepted clusters plus pool clusters
that pass the gate.
- **M2 — consensus clustering.** Each scene is represented by its co-association row over the 12 cached
runs (k 50–100, seeds 21 and 22 — no overlap with 11, 12 or 13): how often it landed with every
other scene. Those rows are clustered with k-means, k = 60, seed 13. Result: clusters that pass the gate.
- **M3 — prune, then grow.** For each seed-13 candidate: if it fails the gate, repeatedly remove the
member with the lowest mean similarity to members from *other* stories, until it passes or falls to
13 members (then it is dropped). Then grow every passing cluster: add the non-member scene nearest
its centroid while the result still passes; stop at the first addition that would fail, or at 30
members. Scenes may end up in more than one M3 cluster; the overlap is reported.
### Evaluation
- Every cluster accepted by any method is pooled (identical member sets once), given a shuffled id, and
dumped to a blind sheet showing only member summaries. Which method produced a cluster is written
to a separate map file that is not read until labelling is finished.
- Labels use the same 2 / 1 / 0 scale. Names are reused verbatim from the development and test label
files wherever the situation is the same, so duplicate situations can be counted by name.
- Per method: clusters accepted; precision for label ≥ 1 and for label 2; **distinct situations**
(distinct names among label ≥ 1); scene coverage; overlap between accepted clusters.
- **Success, per method:** at least **twice** B's accepted clusters, **and** precision (≥ 1) of at least
**0.80**, **and** at least twice B's distinct situations. For scale: a two-stage pack wants ~30.
### Known limits
The reader has now read every summary several times; labels are blind to method and signals, not to
content. The methods and the gate are evaluated on one corpus.
### Results (2026-09-14)
Seed 13 gave 42 candidates. 21 distinct accepted clusters were labelled blind (`probe/labels-recover.json`,
3 labelled 2, 13 labelled 1, 5 labelled 0), then scored by `probe/evaluate_recover.py`:
| Method | Accepted | Precision (≥ 1) | Precision (2) | Distinct situations | Scene coverage | Success |
| --- | --- | --- | --- | --- | --- | --- |
| B, gate on candidates | 4 | 1.00 | 0.00 | 4 | 13% | — |
| M1, re-cluster the rejected | 4 | 1.00 | 0.00 | 4 | 13% | **fail** — 0 new |
| M2, consensus clustering | 3 | 1.00 | **1.00** | 3 | 5% | **fail** — too few |
| M3, prune then grow | 16 | 0.69 | 0.00 | 10 | 45% | **fail** — precision |
**No method meets the pre-registered bar.**
- **M1 adds nothing.** Re-clustering the 729 scenes outside B's clusters at k = 36 produced no cluster
that passes the gate. The gate's selectivity does not come from dilution by one bad partition.
- **M2 is the only method that produced sharp situations.** All three of its clusters are labelled 2
— a wealthy suitor weighed against love, a spouse against a spouse over infidelity, coercion by
threat — at 14–17 members, against B's 13–38. Consensus over many partitions isolates the scenes
that genuinely travel together. It produces too few clusters to be the selection step on its own.
- **M3 finds the most, and the gate cannot stop it diluting.** Pruning recovered 12 rejected
candidates and M3 reached 10 distinct situations (B: 4), but 5 of its 16 clusters are not
situations. **Growth was stopped by the 30-member cap in 10 of 16 clusters, not by the gate.**
That exposes the gate's second size effect: S6 is a share, so it gets *easier* to satisfy as a
cluster grows, and S1 moves slowly. The gate calibrated on k = 60 clusters does not constrain a
cluster that is being built up one scene at a time.
- **Duplicates survive.** Across methods, "a spouse against a spouse" appears three times (R006, R015)
and courtship twice within M3 (R008, R010); 68 scenes sit in more than one M3 cluster.
- **The seed-13 baseline itself accepted no label-2 cluster** (0 of 4), where development and test
each had 2–3. The baseline's precision for label 2 is unstable across partitions.
### What this means
The gate, as frozen, is **size-dependent in both directions**: it structurally rejects clusters under
~13 members and barely constrains clusters above ~25. Any method that changes cluster size — pruning,
growing, finer re-clustering — changes what the gate measures. Every recovery method was, in effect,
testing the gate at sizes it was not calibrated for.
Observation only, not pre-registered and not a result: B and M2 together would give 7 clusters,
6 distinct situations, all labelled ≥ 1, and 3 labelled 2.
Directions this points to, none tested:
1. **Make the gate size-invariant before anything else.** Replace S6's share with a count or a
significance test against the corpus, and judge S1 relative to what a random cluster of the same
size scores. Then re-validate — which needs fresh labels on fresh partitions.
2. **Use consensus (M2) to form clusters, not k-means.** It was the only source of sharp situations.
Running more co-association partitions (more seeds, more k) and extracting cores at several
agreement levels may raise its yield without losing precision.
3. **Grow only with a size-invariant stopping rule.** M3's recall is real; its stopping rule is not.
---
## A size-invariant gate, G2 (pre-registered 2026-09-14, before any G2 signal was computed)
### Why each old signal depended on size
- **S6** is a share with a floor of 2 members: 2/size. Under ~13 members any repeated word fails the
cut; above ~25 almost nothing does. A share is not evidence unless the count is improbable.
- **S1**'s *expectation* does not depend on size — it is a mean over pairs — but its *noise* does:
small clusters reach a high S1 by chance, which is part of why k-means' small clusters look tight.
- **Nothing looked at individual members.** Adding one off-situation scene to 30 moves S1 by
roughly a thirtieth of the difference, so growth was never stopped.
### G2 signals
| Signal | Definition | Why it is size-invariant |
| --- | --- | --- |
| **S1** | Unchanged: mean summary cosine over cross-story member pairs | Unbiased at any size |
| **Z1** | (S1 − μₘ) / σₘ, where μₘ and σₘ come from random member sets of the same size m (≥ 4 stories, dominant ≤ 40%): 300 samples for m ≤ 40, 120 above, sizes 5–150, seeded | Requires more evidence where chance does more |
| **MMIN** | The lowest, over members, of a member's mean cosine to members of *other* stories | A per-member quantity; one diluting member fails it at any size |
| **W** | Among words shared by ≥ 2 members, those whose count is significant by a hypergeometric upper-tail test at Bonferroni 0.01 over the corpus vocabulary; W is the largest share among them (0 if none) | Counts only improbable sharing, then measures how dominant it is |
Word extraction, stopwords and the corpus are identical to `signals.py`. Pairwise similarities are
precomputed once. Candidate conditions (≥ 5 members, ≥ 4 stories, dominant ≤ 40%) still apply.
### Choosing thresholds
- **Tuning set:** every labelled cluster so far — development (40), test (41) and the recovery pool
(21): 102 clusters, 63 labelled ≥ 1, sizes 5–39. All have been seen; none is used for validation.
- **Rule form:** S1 ≥ a AND Z1 ≥ b AND MMIN ≥ c AND W < d, where each term may be switched off. Grid:
a and c at the deciles of the tuning set's values or off; b ∈ {off, 1, 2, 3, 4, 5}; d ∈ {off, 0.25,
0.30, 0.35, 0.40, 0.50, 0.60}.
- **Objective:** the most tuning clusters labelled ≥ 1 accepted, subject to precision (≥ 1) of at least
**0.85**; ties on more label-2 accepted, then fewer active terms. The objective changes from the
first gate (which maximised label-2 precision) because the failure now is recall, and label 1 —
a situational theme — is usable material for a pack.
- The tuning report also gives, per size band (≤ 12, 13–25, ≥ 26), how many clusters the old and new
gates accept and at what precision.
### Validation on fresh data (nothing below has been computed)
- **V1, small clusters:** k = 100, seed 15; 30 candidates sampled at random (seed 15).
- **V2, large clusters:** k = 40, seed 14; 30 candidates sampled at random (seed 14).
- **V3, growth:** k = 60, seed 16; every candidate G2 accepts is grown by adding the non-member nearest
its centroid while the result still passes G2, up to 45 members. The stop reason — gate or cap — is
recorded.
- All V1 and V2 samples and all V3 grown clusters are pooled, deduplicated, shuffled and labelled blind,
with the set and method in a separate map file.
- **Success, all three required:**
1. G2's precision (≥ 1) is at least **0.80 on V1 and on V2 separately**, and G2 accepts at least 3
clusters in each.
2. In V1, G2 accepts **more** small clusters than the old gate.
3. In V3, **most growths are stopped by the gate, not the cap**, and grown clusters have precision
(≥ 1) of at least 0.80.
The old gate is scored on V1 and V2 alongside, for comparison.
### Frozen G2 (written after tuning, before any validation cluster was built)
`probe/tune_gate2.py`, log `probe/tune-gate2.log`. (Correction: the tuning set has 62 clusters
labelled ≥ 1, not 63.)
> **Accept a candidate when S1 ≥ 0.6037 AND Z1 ≥ 5 AND W < 0.60.** MMIN is off.
| Tuning set, 102 clusters | Accepted | Label ≥ 1 | Precision (≥ 1) | Label 2 |
| --- | --- | --- | --- | --- |
| Accept all | 102 | 62 | 0.61 | 13 |
| Old gate | 31 | 26 | 0.84 | 8 |
| **G2** | **40** | **34** | **0.85** | **9** |
| Size band | Clusters (≥ 1) | Old gate: accepted / P1+ | G2: accepted / P1+ |
| --- | --- | --- | --- |
| ≤ 12 | 20 (7) | 0 / — | 1 / 1.00 |
| 13–25 | 52 (34) | 13 / 0.92 | 18 / 0.89 |
| ≥ 26 | 30 (21) | 18 / 0.78 | 21 / 0.81 |
**Observations and predictions, recorded before validation:**
- **G2 is not size-invariant at the small end, and the objective chose that.** Z1 ≥ 5 is a fixed bar
on a statistic whose value, for a fixed true cohesion, *grows* with size (σₘ shrinks as m grows).
It replaces S6's structural rejection of small clusters with a milder one: 1 small cluster accepted
instead of 0. **Prediction: criterion 2 passes narrowly or fails, and V1 yields fewer than 3
accepted clusters, failing criterion 1.**
- **MMIN, the signal designed to stop dilution, was not selected**, and W's cut (0.60) is loose enough
to act only on near-total single-word clusters. Growing a cluster raises Z1 and barely moves S1.
**Prediction: most V3 growths run to the cap, failing criterion 3.**
- The gain over the old gate is real but modest, and it is in the mid and large bands: 8 more clusters
labelled ≥ 1 accepted at unchanged precision.
- The thresholds are not changed in light of these predictions. The validation tests the rule the
pre-registered procedure produced.
### Validation results (2026-09-14)
73 fresh clusters labelled blind (`probe/labels-validate.json`: 6 labelled 2, 33 labelled 1, 34
labelled 0), scored by `probe/validate_gate2.py score` (log `probe/validate-score.log`).
| Set | Sampled (≥ 1) | Old gate: accepted / ≥ 1 / P1+ | G2: accepted / ≥ 1 / P1+ |
| --- | --- | --- | --- |
| V1, k = 100 | 30 (13) | 3 / 2 / 0.67 | **10 / 8 / 0.80** (2 labelled 2) |
| V2, k = 40 | 30 (16) | 1 / 1 / 1.00 | 4 / 3 / 0.75 |
| V3, grown | 13 | — | 13 grown: **11 stopped by the cap**, 2 by the gate; 10 ≥ 1, P1+ 0.77 |
| Criterion | Result |
| --- | --- |
| 1. P1+ ≥ 0.80 and ≥ 3 accepted, on V1 and V2 | **fail** — V1 passes (0.80, 10), V2 does not (0.75, 4) |
| 2. G2 accepts more V1 clusters than the old gate | **pass** — 10 against 3 |
| 3. V3 mostly stopped by the gate, P1+ ≥ 0.80 | **fail** — 2 of 13, 0.77 |
| **Overall** | **fail** |
By size, across V1 and V2 together:
| Size | Old gate: accepted (≥ 1) | G2: accepted (≥ 1) |
| --- | --- | --- |
| ≤ 9 | 0 | **0** — including two clusters labelled 2 (5 and 6 members) |
| 10–12 | 0 | 3 (3) |
| 13–25 | 3 (2) | 7 (5) |
| ≥ 26 | 1 (1) | 4 (3) |
| All | 4 (3), P1+ 0.75 | 14 (11), P1+ 0.79 |
### Verdict on G2
**G2 is a real improvement on the old gate, and it is not size-invariant.** It accepts three to four
times as many clusters at slightly better precision, and it opens the 10–25 member range the old gate
all but closed. Both pre-registered predictions of failure held:
- **The small end is still closed below 10 members.** Z1 ≥ 5 is unreachable for a cluster of 5–9
however cohesive it is. The two sharpest small clusters on the sheet — a fare dispute (5 members,
label 2) and a spouse against a spouse (6 members, label 2) — were both rejected.
- **Growth is still unconstrained.** Growing a cluster raises Z1 and barely moves S1, so both gates
accept a 45-member cluster as readily as a 15-member one. The one per-member signal built to stop
this, MMIN, was not selected by the tuning objective.
- **Large clusters are still mostly rejected** (4 of 30 in V2), including clear themes such as the
saloon, dining rooms and illness.
**What was learned about size-invariance itself:** a fixed threshold on a z-score is not size-invariant;
it reverses the old bias rather than removing it. A size-invariant cohesion test needs a threshold
that scales with the null — for example a fixed *percentile* of the same-size random distribution
combined with an effect-size floor on S1 — and a stopping rule for growth has to be a per-member
test that the tuning cannot switch off. Both are untested proposals, and testing them needs another
fresh partition and another round of blind labels, because every cluster labelled so far has now
informed a design decision.
---
## G3: a percentile of the same-size null, and a mandatory member floor (pre-registered 2026-09-14, before any G3 signal was computed)
### Changes from G2
- **Z1 is replaced by PCT**: the fraction of same-size random member sets (≥ 4 stories, dominant ≤ 40%)
whose S1 is below the cluster's. 1,000 samples per size for 5–60 members, 200 for 61–150, seeded.
A percentile threshold asks "is this cluster as unusual as the top x% of random sets *of its own
size*", which a z-score with a fixed bar does not.
- **S1 floor is mandatory.** At large sizes the null is narrow and PCT saturates at 1.0 for any
cluster slightly above average, so an effect-size floor has to carry the weight there.
- **MMIN is mandatory**: every member's mean cosine to members of other stories must reach the floor.
It is the growth stop; tuning chooses its value but cannot switch it off.
- **W** is unchanged from G2 and remains optional.
### Tuning
- **Tuning set: all 175 labelled clusters** — development 40, test 41, recovery 21, G2 validation 73.
- **Rule form:** S1 ≥ a AND PCT ≥ p AND MMIN ≥ c AND W < d. a and c at the tuning set's deciles (no "off");
p ∈ {0.90, 0.95, 0.975, 0.99, 0.999}; d ∈ {off, 0.30, 0.40, 0.50, 0.60}.
- **Objective unchanged from G2:** most clusters labelled ≥ 1 accepted, subject to precision (≥ 1) of at
least 0.85; ties on more label-2 accepted, then fewer active terms.
**Risk named in advance:** MMIN is a minimum over per-member means. Small clusters have noisier
member means but fewer of them; large clusters have steadier means but more chances to contain one
low member. Its net size effect is not predictable, and it could re-introduce a size bias of its own.
### Validation on fresh partitions (seeds never used before)
- **V0, very small:** k = 150, seed 20; 25 candidates sampled (seed 20).
- **V1, small:** k = 100, seed 17; 20 sampled (seed 17).
- **V2, large:** k = 40, seed 18; 20 sampled (seed 18).
- **V3, growth:** k = 60, seed 19; every G3-accepted candidate grown toward its nearest scene while it
passes G3, cap 45, stop reason recorded.
- Pooled, deduplicated, shuffled and labelled blind as before. The old gate and G2 are scored on the
same clusters.
**Success, all four required:**
1. G3 precision (≥ 1) ≥ **0.80** and ≥ 3 accepted, on V1 and on V2 separately.
2. **Small clusters:** among V0 and V1 clusters of ≤ 9 members, G3 accepts **at least 2**, with precision
(≥ 1) of at least **0.67**, and more than G2 accepts.
3. **Growth:** most V3 growths stopped by the gate, and grown clusters have precision (≥ 1) ≥ 0.80.
4. G3 accepts at least as many clusters labelled ≥ 1 across V0–V2 as G2, at precision no lower than G2's.
### Frozen G3 (written after tuning, before any validation cluster was built)
`probe/tune_gate3.py`, log `probe/tune-gate3.log`.
> **Accept a candidate when S1 ≥ 0.6175 AND PCT ≥ 0.90 AND MMIN ≥ 0.5296 AND W < 0.60.**
| Tuning set, 175 clusters (101 ≥ 1, 19 label 2) | Accepted | Label ≥ 1 | Precision (≥ 1) | Label 2 |
| --- | --- | --- | --- | --- |
| Old gate | 47 | 38 | 0.81 | 9 |
| G2 | 67 | 55 | 0.82 | 12 |
| **G3** | **47** | **40** | **0.85** | **13** |
| Size band | Clusters (≥ 1) | Old: accepted/≥ 1 | G2 | G3 |
| --- | --- | --- | --- | --- |
| ≤ 9 | 26 (5) | 0/0 | 0/0 | **4/3** |
| 10–12 | 21 (11) | 0/0 | 4/4 | 4/4 |
| 13–25 | 75 (47) | 17/15 | 26/22 | 17/17 |
| ≥ 26 | 53 (38) | 30/23 | 37/29 | 22/16 |
**Observations and predictions, recorded before validation:**
- **G3 is the first gate to accept any cluster under 10 members** (4 on tuning, 3 of them ≥ 1).
**Prediction: criterion 2 passes.**
- **The percentile threshold settled at its loosest value, 0.90, and the S1 floor rose** from G2's 0.6037
to 0.6175. The floor, not the percentile, is doing most of the work at mid and large sizes, and G3
accepts 15 fewer large clusters than G2 on tuning. **Prediction: criterion 4 fails** — G3 finds fewer
clusters labelled ≥ 1 than G2 (40 against 55 on tuning), though at higher precision. Criterion 1 is
at risk on V2 for the same reason.
- **MMIN at 0.53 is the growth stop.** Whether a nearest-to-centroid scene falls below it is not
predictable from tuning, which contains no grown-under-G3 clusters. No prediction for criterion 3.
- The trade is visible already: G3 exchanges recall at the large end for precision everywhere and for
admission at the small end. Whether that is the right trade for a pack is a design question the
validation cannot answer.
### Validation results (2026-09-14)
75 fresh clusters labelled blind (`probe/labels-validate3.json`: 4 labelled 2, 24 labelled 1, 47
labelled 0), scored by `probe/validate_gate3.py score` (log `probe/validate3-score.log`).
| Set | Sampled (≥ 1) | Old: acc / ≥ 1 / P1+ | G2: acc / ≥ 1 / P1+ | G3: acc / ≥ 1 / P1+ |
| --- | --- | --- | --- | --- |
| V0, k = 150 | 25 (5) | 2 / 1 / 0.50 | 8 / 4 / 0.50 | 11 / 4 / 0.36 |
| V1, k = 100 | 20 (5) | 1 / 0 / 0.00 | 1 / 1 / 1.00 | 4 / 2 / 0.50 |
| V2, k = 40 | 20 (12) | 0 / 0 / — | 4 / 4 / 1.00 | 2 / 2 / 1.00 |
| ≤ 9 members, V0 + V1 | 27 (4) | 1 / 0 / 0.00 | 3 / 1 / 0.33 | 11 / 3 / 0.27 |
| V0–V2 pooled | 65 (22) | 3 / 1 / 0.33 | **13 / 9 / 0.69** | 17 / 8 / 0.47 |
| V3, grown under G3 | 10 | — | — | 7 stopped by the cap, 3 by the gate; P1+ 0.60 |
**All four criteria fail**, and G3 is worse than G2 on fresh data: more clusters accepted, fewer of
them situations. G3's gain on the tuning set (P1+ 0.85, 13 label-2) did not survive a new partition.
### Why G3 failed — measured, not inferred
- **PCT carries no information.** Across all 175 tuning clusters, 173 reach PCT ≥ 0.90 and 141 reach
1.0 — including 72 of the 74 clusters labelled 0. **The random-set null is the wrong null.** k-means
groups each scene with its nearest neighbours, so every cluster it produces is far tighter than a
random set of scenes of the same size, whether or not it is a situation. A cohesion test has to ask
whether a cluster is tighter than *other clusters of its size*, not than random scenes.
- **MMIN barely separates labels and favours small clusters.** Median MMIN is 0.536 for label 0,
0.543 for label 1 and 0.564 for label 2; by size it is 0.547 for ≤ 9 members, 0.537 for 10–25 and
0.552 for ≥ 26. A minimum over few, noisy per-member means is not reliably lower for small clusters,
so the floor admitted 11 small clusters of which 8 are not situations. This is the risk recorded
before tuning.
- **Growth is still not stopped by the gate** (7 of 10 to the cap). A nearest-to-centroid scene
typically clears a member floor calibrated on whole clusters.
### Where this leaves selection
Across three gates and five blind-labelled evaluations, **G2 remains the best**: pooled precision
(≥ 1) of 0.79 on its own validation and 0.69 here, with the fewest non-situations admitted.
None of the three is size-invariant, and none stops growth.
Two conclusions look robust enough to build on:
1. **Cohesion alone does not identify a situation.** Clusters labelled 0, 1 and 2 overlap heavily on
every cohesion statistic tried (S1, Z1, PCT, MMIN). What separates the best-labelled clusters is
visible to a reader — one transaction — and is not a geometric property of summary embeddings at
the resolution tried.
2. **The labelling budget is the binding constraint.** Each gate needs fresh partitions and 70–80 new
blind labels to test honestly, and every labelled cluster so far has informed a design. Continuing
to iterate gate designs on this corpus without a different kind of evidence is unlikely to converge.
Candidate directions, none tested: a null built from **k-means clusters of the same size on
shuffled-story data** rather than random scenes; a selection step that uses the **inference model to
name each cluster's transaction** and checks agreement across members (needs the GPU); or accepting
G2 as the gate and moving the problem downstream to event construction, where a human pass over
~40 candidate clusters is cheap compared with the labelling this probe has already required.
---
## G4: the model reads the cluster (pre-registered 2026-09-14, before any model call)
The finding above is that what makes a cluster a situation is visible to a reader and not to the
geometry. So let a reader decide — the inference model — and measure it against the blind labels.
### Method
- For each cluster, up to **15 member summaries** (a seeded sample when larger), shuffled and numbered,
go to the model in one call with the prompt below. It returns JSON: the single social transaction
most of them clearly show, or NONE, and the numbers of the summaries that clearly show it.
- **FIT** = (number of valid summary numbers listed) / (number shown); 0 when the answer is NONE or
the reply cannot be parsed. Parse failures are counted and reported.
- **Gate: accept when FIT ≥ θ**, plus the usual candidate conditions (≥ 5 members, ≥ 4 stories,
dominant ≤ 40%), which every labelled cluster already meets.
- Deterministic: temperature 0, 2,048-token context, `format: json`. One request at a time, as before.
- **The prompt contains no example of a situation** — the 3B model copied the examples it was given
in the summarisation pilot.
Prompt (system), fixed:
> You read short summaries of scenes from different stories and decide whether they share one kind
> of social situation.
>
> A social situation is a specific transaction between people: who wants what from whom, and what
> is at stake. Describe it in your own words.
>
> Rules:
> - Find the single situation that the largest number of summaries clearly show.
> - It is normal for many summaries not to fit. Include a summary only if it clearly shows that
> situation.
> - A shared place, mood, profession or word is not a situation.
> - If no situation is clearly shown by at least three summaries, the situation is NONE.
>
> Reply with JSON only: {"situation": "<at most 10 words, or NONE>", "fits": [<numbers of the
> summaries that clearly show it>]}
The user message is the numbered list. For `qwen3:14b` it ends with `/no_think`, and any think block
is stripped before parsing.
### Models
- **Primary: `qwen2.5:3b-instruct`** — the design requires a small local model; the verdict rests on it.
- **Secondary: `qwen3:14b`** — reported, to show whether capability is what limits the result.
### Data — no new labels
The model is never fitted to labels; the only fitted quantity is θ.
- **θ is chosen on development (40 clusters)** only, with the objective used for G2 and G3: most
clusters labelled ≥ 1 accepted, subject to precision (≥ 1) ≥ 0.85; ties on more label-2 accepted,
then the higher θ. Grid: θ ∈ {0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0}.
- **Primary evaluation: the G2 and G3 validation sets, 148 clusters**, which tuned neither G2 nor G4,
so G2 can be compared fairly on them. **Secondary:** test and recovery, 62 clusters.
- Identical member sets appearing in more than one file are judged once.
- **Pilot, allowed once:** the first 5 development clusters, to check the JSON parses. If more than 1
of 5 fails to parse, one prompt revision is allowed and recorded here before any other call.
### Success, on the 148-cluster primary set, with the 3B model
1. Precision (≥ 1) at least **0.80**.
2. **More** clusters labelled ≥ 1 accepted than G2 accepts on the same set.
3. Among accepted clusters of ≤ 9 members, precision (≥ 1) at least 0.67, if there are at least 3.
Also reported: G4 AND G2; FIT distribution by label; the situations named for accepted clusters
against the label names; and the 14B results against the same bar.
**Pilot (21:23):** 5 of 5 replies parsed with the 3B model, so the prompt is unchanged. About 1.6 s
per cluster. Noted without drawing a conclusion from five: the situations named are often generic
("Confrontation and hidden motives" for D04, a label-2 spouse-against-spouse cluster, FIT 0.47), and
a label-1 argument cluster scored FIT 0.87.
### Results, `qwen2.5:3b-instruct` (21:24–21:29, 245 distinct clusters in 286 s, 1 parse failure)
Log `probe/evaluate-judge-3b.log`. θ chosen on development: **0.6** (9 accepted, P1+ 0.89).
| Primary set, 148 clusters (67 ≥ 1) | Accepted | ≥ 1 | Label 2 | P1+ | Recall |
| --- | --- | --- | --- | --- | --- |
| Accept all | 148 | 67 | 10 | 0.45 | 1.00 |
| **G2** | **49** | **36** | 5 | **0.73** | **0.54** |
| G4, FIT ≥ 0.6 | 50 | 24 | 4 | 0.48 | 0.36 |
| G4 AND G2 | 18 | 11 | 1 | 0.61 | 0.16 |
On the secondary set (62) G4 reaches P1+ 0.71 against G2's 0.82. **All three criteria fail**
(P1+ 0.48; 24 ≥ 1 against G2's 36; 13 accepted small clusters at P1+ 0.23).
**FIT does not separate the labels:** mean 0.39 for label 0, 0.49 for label 1 — and **0.42 for label
2**, lower than for label 1. Precision at the primary set (0.48) is barely above accepting everything
(0.45).
**Why, visible in the replies:** the 3B model names situations at a level of abstraction where almost
anything fits — "confrontation", "power dynamics in social hierarchies", "tension between parties",
"coercion and manipulation" — and then lists most summaries as fitting it. A 5-member cluster labelled
0 scored FIT 1.00 as "power struggle/conflict". When it is specific it is often right ("marital
conflict" for a spouse-against-spouse cluster, "con or scam", "argue over payment"), but specificity
is not what FIT rewards. The rule against shared place, mood and profession was followed; a rule
against *abstraction* was never written, and the model supplies abstraction freely.
### 14B pilot and the one allowed revision
The `qwen3:14b` pilot failed to parse **5 of 5**: every reply's content was empty. The model spent its
whole 200-token budget in a separate thinking channel (Ollama returns it outside `content`), and the
`/no_think` suffix did not stop it. **Revision, the one allowed by the pilot rule, affecting only the
14B run:** requests set Ollama's `think: false` and allow 300 output tokens. The prompt text,
sampling, θ procedure and success criteria are unchanged. The five empty results were deleted before
re-running the pilot.
Re-run pilot: **5 of 5 parsed**, about 2.4 s per cluster. The situations named are more specific than
the 3B model's ("Wealthy person pressures someone into marriage", "Struggling artist facing societal or
personal challenges"), though an abstract one still scored FIT 0.80 ("conflict between individuals with
differing perspectives").
### Results, `qwen3:14b` (21:31–21:45, 245 clusters in 832 s, 0 parse failures)
Log `probe/evaluate-judge-14b.log`. θ chosen on development: **0.6** (8 accepted, P1+ 0.88).
| Primary set, 148 clusters (67 ≥ 1) | Accepted | ≥ 1 | Label 2 | P1+ | Recall |
| --- | --- | --- | --- | --- | --- |
| Accept all | 148 | 67 | 10 | 0.45 | 1.00 |
| **G2** | **49** | **36** | 5 | **0.73** | **0.54** |
| G4 (14B), FIT ≥ 0.6 | 42 | 22 | 4 | 0.52 | 0.33 |
| G4 (14B) AND G2 | 10 | 10 | 1 | 1.00 | 0.15 |
By size, G4 (14B) on the primary set: ≤ 9 members 17 accepted, P1+ 0.29; 10–25, 14 accepted, P1+ 0.50;
≥ 26, 11 accepted, **P1+ 0.91**. Secondary set (62): G4 P1+ 0.60 against G2's 0.82.
**All three criteria fail with the 14B model as well** (P1+ 0.52; 22 ≥ 1 against G2's 36; 17 small
clusters accepted at P1+ 0.29).
FIT now rises with the label — mean 0.40, 0.47, 0.52 for labels 0, 1, 2 — where the 3B model's did
not, so capability helps. But the separation is small next to the spread, and the 14B model still
names abstractions that fit almost anything ("Someone wants something from someone else", "Power
dynamics between individuals", "Someone holds power over someone else"), most often for small
clusters.
GPU host during both runs: no faults logged; peak 284.9 W (1 s), mean 163 W under load, 65 °C, PCIe
Gen 3 x4 throughout. Logs in `probe/gpu-logs/*2122*`.
### Verdict on G4
**Neither model's judgement is a better gate than G2**, and the pre-registered verdict (on the 3B
model) is a fail. The failure is specific and repeatable: asked for "the single situation most
summaries share", a model answers at whatever level of abstraction makes the most summaries fit, and
FIT rewards exactly that. Small clusters make it worse, because with 5–9 summaries an abstraction that
covers four or five of them is always available.
Two observations, recorded as observations because neither was a pre-registered criterion:
- **G4 (14B) AND G2 accepted 10 clusters on the primary set, and all 10 are labelled ≥ 1.** The two
gates fail in different ways — G2 admits cohesive clusters that are not situations, the model admits
abstract "situations" that are not cohesive — and their intersection is precise. It is also
narrow: recall 0.15.
- **For clusters of 26 or more, the 14B model's precision is 0.91.** The abstraction problem is a
small-cluster problem: 15 shown summaries rarely all fit one vague phrase *and* get listed.
What would test the obvious next step, none of it done: a prompt that asks for the situation **in
terms of roles and a concrete want** and rejects answers without both; scoring the *specificity* of
the named situation (for example, requiring the named situation to be closer to the member summaries
than to the corpus average in embedding space); and G2 AND G4 evaluated as a pre-registered gate on a
fresh partition, since every labelled cluster now informs this observation.