Research toward building a content pack from a story corpus, kept on its own branch and independent of the game. Records the selection experiments against blind labels, and settles selection as gate G2 followed by a human review: review.py writes REVIEW.md and a review.json form, apply_review.py checks the filled form and writes situations.json for the next stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
745 lines
46 KiB
Markdown
745 lines
46 KiB
Markdown
# story-to-pack — selecting situations
|
||
|
||
> **Decision (2026-09-15): selection is gate G2 followed by a structured human review.** `probe/review.py`
|
||
> writes a review package (`REVIEW.md`, `review.json`, `candidates.json`); the reviewer fills in
|
||
> `review.json`; `probe/apply_review.py` checks it and writes `situations.json` for the next stage. See
|
||
> the README. The record below is how that decision was reached: four gate designs and two
|
||
> model-judged gates, each tested against blind labels.
|
||
|
||
The step after clustering: from the clusters of summarised scenes, choose the ones that are genuine
|
||
**recurring situations** — the raw material of a recurring event — and reject the rest. Companion to
|
||
`DESIGN.md`; this is the step its assumption-1 result called for.
|
||
|
||
**Status (2026-09-14): designed and pre-registered below, before any feature was computed or any
|
||
label was read against a feature.** Results are appended at the end, and nothing above the Results
|
||
heading is edited after the test runs, except to fix a typo.
|
||
|
||
---
|
||
|
||
## The problem, from the assumption-1 result
|
||
|
||
Summarise-then-embed broke the story-identity effect: clusters now draw from many stories. But of five
|
||
sampled tight-recurring clusters, one was a **vocabulary cluster** — fortune-tellers, adventurers, a
|
||
found tyre, a graft inquiry — held together by "seeks", "fortune", "mysterious", and mostly from three
|
||
stories. The concentration and tightness filters in `measure.py` let it through. A selector has to
|
||
reject clusters like that one and keep clusters like "a spouse schemes against or confronts a spouse
|
||
over infidelity".
|
||
|
||
## Constraints
|
||
|
||
- **No inference.** Selection uses only what is already on disk: the summary embeddings, the prose
|
||
embeddings, the summary text, and which story each scene came from. It is pure computation, like the
|
||
numeric search, so it can be re-run and tuned for free. (It also keeps the inference host free for
|
||
other work.)
|
||
- **Deterministic.** Seeded clustering; the same inputs give the same selection.
|
||
- **Pure Python**, like the rest of the probe.
|
||
|
||
## What a selected situation is
|
||
|
||
A set of scenes, from **at least 4 different stories**, with no story contributing more than 40%,
|
||
that a reader can name as **one social transaction**: who wants what from whom. It becomes one
|
||
recurring event; its member scenes become that event's description variants and option material.
|
||
|
||
Rejected: story-dominated clusters (already mechanical), **vocabulary clusters** (shared words, no
|
||
shared transaction), **register clusters** (shared tone or setting, e.g. "men walking through the
|
||
city"), and **near-duplicates** of an already-selected situation.
|
||
|
||
## Candidate signals
|
||
|
||
All computed per cluster, from existing vectors and text. Each carries a pre-registered prediction of
|
||
which way it moves for a genuine situation.
|
||
|
||
| # | Signal | Definition | Prediction for a situation |
|
||
| --- | --- | --- | --- |
|
||
| S1 | **Cross-story cohesion** | Mean summary-embedding cosine over member pairs from *different* stories | Higher |
|
||
| S2 | **Cross/within ratio** | S1 divided by the mean cosine over member pairs from the *same* story (1.0 if there are none) | Closer to 1 — a situation is as similar across stories as within one; a cluster held together by a few stories' scenes is not |
|
||
| S3 | **Effective stories** | Inverse Simpson index of the member story counts, divided by cluster size | Higher — members spread evenly rather than a few stories plus stragglers |
|
||
| S4 | **Prose agreement** | Mean *prose*-embedding cosine over cross-story member pairs, minus the corpus-wide mean cross-story prose cosine | Higher — a second, independent view: scenes of the same situation should read a little alike even in the original text |
|
||
| S5 | **Stability** | Mean pairwise co-association of members across 12 k-means runs (k from 50 to 100, several seeds): the share of runs in which two members land together | Higher |
|
||
| S6 | **Word dependence** | Share of members containing the cluster's single most over-represented content word (log-odds against the whole summary set, stopwords and "someone" excluded) | Higher for a **vocabulary** cluster — one word doing the work |
|
||
|
||
S6 is predicted to indicate the failure, the others to indicate success.
|
||
|
||
## Protocol
|
||
|
||
1. **Candidates.** Cluster the summary embeddings with k = 60. A candidate is a cluster with ≥ 5
|
||
members, ≥ 4 distinct stories and dominant share ≤ 40% (`measure.py`'s "recurring"; the relative
|
||
tightness filter is *not* applied — the selector replaces it).
|
||
2. **Two candidate sets from two partitions.** Development: seed 11. Test: seed 12.
|
||
3. **Blind labels.** Each set is dumped to a sheet showing, per cluster, only its member summaries in
|
||
shuffled order — no story titles, no scores, no sizes beyond the lines themselves. Each cluster is
|
||
labelled by reading:
|
||
- **2 — one situation:** one nameable social transaction that most members (roughly 70%+)
|
||
instantiate. The label records its name.
|
||
- **1 — broad:** a shared setting or theme, but the transaction varies across members.
|
||
- **0 — not a situation:** vocabulary, register, or nothing nameable.
|
||
|
||
The test sheet is labelled before any signal is computed for the test set.
|
||
4. **Choose the rule on development only.** Compute S1–S6 for the development set. Candidate rules
|
||
are single thresholds on one signal, or an AND of two. The chosen rule is the one with the highest
|
||
**precision** (share of accepted clusters labelled 2), subject to accepting at least half of the
|
||
development set's label-2 clusters; ties go to the simpler rule. Freeze it here, in writing, before
|
||
step 5.
|
||
|
||
*Detail fixed after the labels were written and before any signal was computed:* thresholds are
|
||
placed midway between adjacent observed development values, never on one. Ties beyond simplicity
|
||
break on more label-2 clusters accepted, then on precision for label ≥ 1. Two reference points
|
||
are reported but cannot be chosen: accepting every candidate, and `measure.py`'s tightness filter
|
||
(coherence at or above the partition's median sized cluster), called B0.
|
||
5. **Test once.** Apply the frozen rule to the test set. Report precision (label 2, and label ≥ 1),
|
||
recall of label-2 clusters, and the number accepted. No re-tuning after seeing test results; any
|
||
later change is a new rule and needs new data to be tested on.
|
||
6. **Near-duplicates.** Among accepted clusters, merge any pair whose centroids have cosine ≥ a
|
||
threshold chosen on development by reading the labels' names. Report how many merges happened on
|
||
the test set and whether each merged pair's names agree.
|
||
|
||
## Known limits of this test
|
||
|
||
- **The labels are one reader's judgement** — the model running the probe. They are written to a
|
||
file so the designer can spot-check or overrule them, and the verdict should be read with that in
|
||
mind.
|
||
- **Development and test come from the same corpus.** Two partitions of the same 838 scenes share
|
||
situations, so the test measures stability across clustering, not generalisation to a new corpus.
|
||
A genuinely held-out test needs a second corpus, which needs embedding, which needs the inference
|
||
host.
|
||
- **k is fixed at 60** for the test. The selector's behaviour at other k is not tested.
|
||
|
||
## Frozen rule (written before any test signal was computed)
|
||
|
||
Chosen on development (40 candidates, 5 labelled 2; must accept ≥ 3), by `probe/select_rule.py choose`:
|
||
|
||
> **Accept a candidate when S1 ≥ 0.6119 AND S6 ≤ 0.1603.**
|
||
> Cross-story cohesion at least 0.6119, and no single over-represented word in more than 16% of members.
|
||
|
||
| On development | Accepted | Precision (2) | Precision (≥ 1) | Recall (2) |
|
||
| --- | --- | --- | --- | --- |
|
||
| Accept every candidate | 40 | 0.12 | 0.62 | 1.00 |
|
||
| B0, tightness filter | 17 | 0.24 | 0.76 | 0.80 |
|
||
| S1 alone (best single) | 9 | 0.44 | 0.89 | 0.80 |
|
||
| **Frozen rule** | **5** | **0.60** | **1.00** | **0.60** |
|
||
|
||
Accepted on development: D04 (2), D09 (1), D13 (2), D16 (2), D33 (1). Rejected label-2: D15 (S1 0.603),
|
||
D30 (S6 0.20). Full leaderboard: `probe/choose-dev.log`; signals: `probe/signals-dev.json`.
|
||
|
||
Observations recorded now, so they cannot be shaped by the test:
|
||
|
||
- **This is very likely overfit.** Five positives, fine-grained thresholds, and an AND of two signals
|
||
is exactly the setting in which a rule memorises its development set. The test is what decides.
|
||
- **S1 does most of the work.** It is the best single signal by a wide margin, and it alone rejects
|
||
the vocabulary cluster that motivated this step (D18, S1 0.605).
|
||
- **S6 did not behave as predicted.** Its most over-represented word is usually a *story's* word
|
||
("juggler", "furs", "eviction"), not generic vocabulary, and D18's is low (0.11). It catches D39
|
||
("young", 1.00) — a vocabulary cluster of a different kind. Its contribution to the frozen rule is
|
||
a single cut, and may not carry over.
|
||
- **S2, S3 and S4 were weak on development**, and S5 middling. None of the predicted directions was
|
||
strongly wrong, but only S1 separated well.
|
||
|
||
**Near-duplicates (step 6): centroid cosine cannot do it, so no automatic merge is applied.** Among
|
||
the five clusters the rule accepts on development, no two share a situation, so there was nothing
|
||
there to calibrate on. Across all 26 development candidates labelled ≥ 1, the one clear duplicate —
|
||
D13 and D30, both "a run-in with the police" — has centroid cosine **0.9026**, while five pairs of
|
||
plainly different situations score higher, up to **0.9326** ("money and marriage" against "schemes,
|
||
cons and theft"). D05/D23 (artists) sit at 0.9005. Any threshold that merges the duplicate merges
|
||
different situations first. Summary embeddings place situations about money, romance and conflict
|
||
in one dense region, so closeness between centroids is not identity of situation.
|
||
|
||
Decision, fixed before the test: the test reports, **by label name**, whether any two accepted
|
||
clusters are the same situation, as a measurement of how often duplicates survive selection. A
|
||
merge method is an open problem; member overlap across partitions (co-association between two
|
||
clusters' members) is the candidate to try next, and is untested.
|
||
|
||
## Results
|
||
|
||
**Test run once, 2026-09-14,** on the seed-12 partition (41 candidates, 5 labelled 2), frozen rule
|
||
unchanged. Logs: `probe/signals-test.json`, `probe/apply-test.log`.
|
||
|
||
| On test | Accepted | Precision (2) | Precision (≥ 1) | Recall (2) |
|
||
| --- | --- | --- | --- | --- |
|
||
| Accept every candidate | 41 | 0.12 | 0.51 | 1.00 |
|
||
| B0, tightness filter | 18 | 0.17 | 0.44 | 0.60 |
|
||
| **Frozen rule** | **5** | **0.40** | **1.00** | **0.40** |
|
||
|
||
Accepted: T35 (2) coercion over money; T39 (2) spouse against spouse over infidelity; T05 (1) schemes
|
||
and cons over money; T07 (1) courtship; T25 (1) trying to impress or win someone's favour.
|
||
Rejected label-2: T16 newcomer seeks acceptance (S1 0.598), T17 proposal or engagement (S1 0.596),
|
||
T30 run-in with the police (S1 0.604, S6 0.23).
|
||
|
||
### Verdict
|
||
|
||
**The selector is precise and narrow.** Across both sets it accepted 10 clusters, and **all 10 are
|
||
situations or broad situational themes; none is a vocabulary, register or grab-bag cluster.**
|
||
Accepting everything gives about half that. The tightness filter used in the assumption-1 test is
|
||
*worse* than accepting everything on the test set (0.44 against 0.51) — it keeps vocabulary clusters,
|
||
which are tight.
|
||
|
||
**It does not find most of the situations.** Recall of label-2 clusters fell from 0.60 to 0.40, and
|
||
precision for label 2 from 0.60 to 0.40 — the expected overfitting. The rejected label-2 clusters sit
|
||
just under the S1 cut (0.596–0.604), so the threshold is doing real work at a place where situations
|
||
and non-situations overlap. **Five accepted clusters out of 41 is far short of the ~15 situations
|
||
per stage a pack needs**, so as it stands selection would make the sufficiency report fail on this
|
||
corpus.
|
||
|
||
**S6 earned its place on the test set, in the way it failed to on development.** Two test clusters are
|
||
held together by a single word — T10 "caliph" (S6 0.70) and T18 "young" (1.00) — and both have high
|
||
cross-story cohesion (S1 0.628, 0.626). S1 alone would accept them; S6 rejects them. (Observation made
|
||
after the test, not a re-tuning: S1 ≥ 0.6119 alone would accept 10 test clusters, 2 labelled 2 and 6 labelled ≥ 1.)
|
||
Its original motivation, D18, was caught by S1 instead. Both failure modes exist, and each signal
|
||
catches one.
|
||
|
||
**Duplicates among accepted clusters:** no two carry the same label name. T07 (courtship) and T25
|
||
(trying to impress or win someone's favour) overlap in substance; a reader could reasonably call them
|
||
one situation. Automatic merging remains unsolved (see step 6).
|
||
|
||
### What this means for the design
|
||
|
||
1. **Two-signal selection is a sound first gate, not a complete step.** It should run as a *filter
|
||
that is allowed to reject* — its precision is what matters for pack quality — followed by a
|
||
second stage that recovers recall.
|
||
2. **Recall is the next problem, and clustering is the likely cause, not selection.** The rejected
|
||
situations (police run-ins, proposals, social climbing) are recognisable but diluted by
|
||
off-situation members at k = 60. Candidates to test, all without inference: re-cluster the members
|
||
of rejected clusters at finer k and re-apply the gate; select from the union of several partitions,
|
||
using co-association to keep only members that consistently travel together; or grow an accepted
|
||
cluster's core by adding nearest cross-story members while S1 stays above the cut.
|
||
3. **Merging duplicates needs member overlap, not centroid distance.**
|
||
4. **These labels are one reader's.** Every number above rests on them. They are in
|
||
`probe/labels-dev.json` and `probe/labels-test.json`, with the sheets they were read from, and are
|
||
worth a designer's spot-check — especially the 1-versus-0 boundary, which is the least certain.
|
||
5. **The test is within one corpus.** Generalisation to a different corpus is untested and needs
|
||
embeddings, which needs the inference host.
|
||
|
||
---
|
||
|
||
## Recovering recall (pre-registered 2026-09-14, before any recovery method was run)
|
||
|
||
The gate is precise and keeps too few. Three ways to find more situations, compared on **fresh data**,
|
||
because the test set above is spent. Still no inference.
|
||
|
||
### Fixed before running
|
||
|
||
- **The gate does not change.** `S1 ≥ 0.6119 AND S6 ≤ 0.1603`, plus the candidate conditions (≥ 5
|
||
members, ≥ 4 stories, dominant ≤ 40%). "Passes the gate" below means all of these.
|
||
- **A property of the gate, noted in advance:** S6 counts a word only if at least two members share
|
||
it, so a cluster of fewer than 13 members fails whenever any content word repeats (2/12 = 0.167). In
|
||
effect the gate requires ~13 members. The methods are built to produce clusters of that size;
|
||
anything smaller is expected to fail, and that is the gate's behaviour, not a method's.
|
||
- **Fresh partition:** k = 60, **seed 13** — neither development (11) nor test (12).
|
||
- **Baseline (B):** the gate applied to seed 13's candidates, as before.
|
||
- **M1 — re-cluster the rejected.** Every scene not in a B-accepted cluster is pooled and re-clustered
|
||
with k-means, k = round(pool size / 20), seed 13. Result: B's accepted clusters plus pool clusters
|
||
that pass the gate.
|
||
- **M2 — consensus clustering.** Each scene is represented by its co-association row over the 12 cached
|
||
runs (k 50–100, seeds 21 and 22 — no overlap with 11, 12 or 13): how often it landed with every
|
||
other scene. Those rows are clustered with k-means, k = 60, seed 13. Result: clusters that pass the gate.
|
||
- **M3 — prune, then grow.** For each seed-13 candidate: if it fails the gate, repeatedly remove the
|
||
member with the lowest mean similarity to members from *other* stories, until it passes or falls to
|
||
13 members (then it is dropped). Then grow every passing cluster: add the non-member scene nearest
|
||
its centroid while the result still passes; stop at the first addition that would fail, or at 30
|
||
members. Scenes may end up in more than one M3 cluster; the overlap is reported.
|
||
|
||
### Evaluation
|
||
|
||
- Every cluster accepted by any method is pooled (identical member sets once), given a shuffled id, and
|
||
dumped to a blind sheet showing only member summaries. Which method produced a cluster is written
|
||
to a separate map file that is not read until labelling is finished.
|
||
- Labels use the same 2 / 1 / 0 scale. Names are reused verbatim from the development and test label
|
||
files wherever the situation is the same, so duplicate situations can be counted by name.
|
||
- Per method: clusters accepted; precision for label ≥ 1 and for label 2; **distinct situations**
|
||
(distinct names among label ≥ 1); scene coverage; overlap between accepted clusters.
|
||
- **Success, per method:** at least **twice** B's accepted clusters, **and** precision (≥ 1) of at least
|
||
**0.80**, **and** at least twice B's distinct situations. For scale: a two-stage pack wants ~30.
|
||
|
||
### Known limits
|
||
|
||
The reader has now read every summary several times; labels are blind to method and signals, not to
|
||
content. The methods and the gate are evaluated on one corpus.
|
||
|
||
### Results (2026-09-14)
|
||
|
||
Seed 13 gave 42 candidates. 21 distinct accepted clusters were labelled blind (`probe/labels-recover.json`,
|
||
3 labelled 2, 13 labelled 1, 5 labelled 0), then scored by `probe/evaluate_recover.py`:
|
||
|
||
| Method | Accepted | Precision (≥ 1) | Precision (2) | Distinct situations | Scene coverage | Success |
|
||
| --- | --- | --- | --- | --- | --- | --- |
|
||
| B, gate on candidates | 4 | 1.00 | 0.00 | 4 | 13% | — |
|
||
| M1, re-cluster the rejected | 4 | 1.00 | 0.00 | 4 | 13% | **fail** — 0 new |
|
||
| M2, consensus clustering | 3 | 1.00 | **1.00** | 3 | 5% | **fail** — too few |
|
||
| M3, prune then grow | 16 | 0.69 | 0.00 | 10 | 45% | **fail** — precision |
|
||
|
||
**No method meets the pre-registered bar.**
|
||
|
||
- **M1 adds nothing.** Re-clustering the 729 scenes outside B's clusters at k = 36 produced no cluster
|
||
that passes the gate. The gate's selectivity does not come from dilution by one bad partition.
|
||
- **M2 is the only method that produced sharp situations.** All three of its clusters are labelled 2
|
||
— a wealthy suitor weighed against love, a spouse against a spouse over infidelity, coercion by
|
||
threat — at 14–17 members, against B's 13–38. Consensus over many partitions isolates the scenes
|
||
that genuinely travel together. It produces too few clusters to be the selection step on its own.
|
||
- **M3 finds the most, and the gate cannot stop it diluting.** Pruning recovered 12 rejected
|
||
candidates and M3 reached 10 distinct situations (B: 4), but 5 of its 16 clusters are not
|
||
situations. **Growth was stopped by the 30-member cap in 10 of 16 clusters, not by the gate.**
|
||
That exposes the gate's second size effect: S6 is a share, so it gets *easier* to satisfy as a
|
||
cluster grows, and S1 moves slowly. The gate calibrated on k = 60 clusters does not constrain a
|
||
cluster that is being built up one scene at a time.
|
||
- **Duplicates survive.** Across methods, "a spouse against a spouse" appears three times (R006, R015)
|
||
and courtship twice within M3 (R008, R010); 68 scenes sit in more than one M3 cluster.
|
||
- **The seed-13 baseline itself accepted no label-2 cluster** (0 of 4), where development and test
|
||
each had 2–3. The baseline's precision for label 2 is unstable across partitions.
|
||
|
||
### What this means
|
||
|
||
The gate, as frozen, is **size-dependent in both directions**: it structurally rejects clusters under
|
||
~13 members and barely constrains clusters above ~25. Any method that changes cluster size — pruning,
|
||
growing, finer re-clustering — changes what the gate measures. Every recovery method was, in effect,
|
||
testing the gate at sizes it was not calibrated for.
|
||
|
||
Observation only, not pre-registered and not a result: B and M2 together would give 7 clusters,
|
||
6 distinct situations, all labelled ≥ 1, and 3 labelled 2.
|
||
|
||
Directions this points to, none tested:
|
||
|
||
1. **Make the gate size-invariant before anything else.** Replace S6's share with a count or a
|
||
significance test against the corpus, and judge S1 relative to what a random cluster of the same
|
||
size scores. Then re-validate — which needs fresh labels on fresh partitions.
|
||
2. **Use consensus (M2) to form clusters, not k-means.** It was the only source of sharp situations.
|
||
Running more co-association partitions (more seeds, more k) and extracting cores at several
|
||
agreement levels may raise its yield without losing precision.
|
||
3. **Grow only with a size-invariant stopping rule.** M3's recall is real; its stopping rule is not.
|
||
|
||
---
|
||
|
||
## A size-invariant gate, G2 (pre-registered 2026-09-14, before any G2 signal was computed)
|
||
|
||
### Why each old signal depended on size
|
||
|
||
- **S6** is a share with a floor of 2 members: 2/size. Under ~13 members any repeated word fails the
|
||
cut; above ~25 almost nothing does. A share is not evidence unless the count is improbable.
|
||
- **S1**'s *expectation* does not depend on size — it is a mean over pairs — but its *noise* does:
|
||
small clusters reach a high S1 by chance, which is part of why k-means' small clusters look tight.
|
||
- **Nothing looked at individual members.** Adding one off-situation scene to 30 moves S1 by
|
||
roughly a thirtieth of the difference, so growth was never stopped.
|
||
|
||
### G2 signals
|
||
|
||
| Signal | Definition | Why it is size-invariant |
|
||
| --- | --- | --- |
|
||
| **S1** | Unchanged: mean summary cosine over cross-story member pairs | Unbiased at any size |
|
||
| **Z1** | (S1 − μₘ) / σₘ, where μₘ and σₘ come from random member sets of the same size m (≥ 4 stories, dominant ≤ 40%): 300 samples for m ≤ 40, 120 above, sizes 5–150, seeded | Requires more evidence where chance does more |
|
||
| **MMIN** | The lowest, over members, of a member's mean cosine to members of *other* stories | A per-member quantity; one diluting member fails it at any size |
|
||
| **W** | Among words shared by ≥ 2 members, those whose count is significant by a hypergeometric upper-tail test at Bonferroni 0.01 over the corpus vocabulary; W is the largest share among them (0 if none) | Counts only improbable sharing, then measures how dominant it is |
|
||
|
||
Word extraction, stopwords and the corpus are identical to `signals.py`. Pairwise similarities are
|
||
precomputed once. Candidate conditions (≥ 5 members, ≥ 4 stories, dominant ≤ 40%) still apply.
|
||
|
||
### Choosing thresholds
|
||
|
||
- **Tuning set:** every labelled cluster so far — development (40), test (41) and the recovery pool
|
||
(21): 102 clusters, 63 labelled ≥ 1, sizes 5–39. All have been seen; none is used for validation.
|
||
- **Rule form:** S1 ≥ a AND Z1 ≥ b AND MMIN ≥ c AND W < d, where each term may be switched off. Grid:
|
||
a and c at the deciles of the tuning set's values or off; b ∈ {off, 1, 2, 3, 4, 5}; d ∈ {off, 0.25,
|
||
0.30, 0.35, 0.40, 0.50, 0.60}.
|
||
- **Objective:** the most tuning clusters labelled ≥ 1 accepted, subject to precision (≥ 1) of at least
|
||
**0.85**; ties on more label-2 accepted, then fewer active terms. The objective changes from the
|
||
first gate (which maximised label-2 precision) because the failure now is recall, and label 1 —
|
||
a situational theme — is usable material for a pack.
|
||
- The tuning report also gives, per size band (≤ 12, 13–25, ≥ 26), how many clusters the old and new
|
||
gates accept and at what precision.
|
||
|
||
### Validation on fresh data (nothing below has been computed)
|
||
|
||
- **V1, small clusters:** k = 100, seed 15; 30 candidates sampled at random (seed 15).
|
||
- **V2, large clusters:** k = 40, seed 14; 30 candidates sampled at random (seed 14).
|
||
- **V3, growth:** k = 60, seed 16; every candidate G2 accepts is grown by adding the non-member nearest
|
||
its centroid while the result still passes G2, up to 45 members. The stop reason — gate or cap — is
|
||
recorded.
|
||
- All V1 and V2 samples and all V3 grown clusters are pooled, deduplicated, shuffled and labelled blind,
|
||
with the set and method in a separate map file.
|
||
- **Success, all three required:**
|
||
1. G2's precision (≥ 1) is at least **0.80 on V1 and on V2 separately**, and G2 accepts at least 3
|
||
clusters in each.
|
||
2. In V1, G2 accepts **more** small clusters than the old gate.
|
||
3. In V3, **most growths are stopped by the gate, not the cap**, and grown clusters have precision
|
||
(≥ 1) of at least 0.80.
|
||
|
||
The old gate is scored on V1 and V2 alongside, for comparison.
|
||
|
||
### Frozen G2 (written after tuning, before any validation cluster was built)
|
||
|
||
`probe/tune_gate2.py`, log `probe/tune-gate2.log`. (Correction: the tuning set has 62 clusters
|
||
labelled ≥ 1, not 63.)
|
||
|
||
> **Accept a candidate when S1 ≥ 0.6037 AND Z1 ≥ 5 AND W < 0.60.** MMIN is off.
|
||
|
||
| Tuning set, 102 clusters | Accepted | Label ≥ 1 | Precision (≥ 1) | Label 2 |
|
||
| --- | --- | --- | --- | --- |
|
||
| Accept all | 102 | 62 | 0.61 | 13 |
|
||
| Old gate | 31 | 26 | 0.84 | 8 |
|
||
| **G2** | **40** | **34** | **0.85** | **9** |
|
||
|
||
| Size band | Clusters (≥ 1) | Old gate: accepted / P1+ | G2: accepted / P1+ |
|
||
| --- | --- | --- | --- |
|
||
| ≤ 12 | 20 (7) | 0 / — | 1 / 1.00 |
|
||
| 13–25 | 52 (34) | 13 / 0.92 | 18 / 0.89 |
|
||
| ≥ 26 | 30 (21) | 18 / 0.78 | 21 / 0.81 |
|
||
|
||
**Observations and predictions, recorded before validation:**
|
||
|
||
- **G2 is not size-invariant at the small end, and the objective chose that.** Z1 ≥ 5 is a fixed bar
|
||
on a statistic whose value, for a fixed true cohesion, *grows* with size (σₘ shrinks as m grows).
|
||
It replaces S6's structural rejection of small clusters with a milder one: 1 small cluster accepted
|
||
instead of 0. **Prediction: criterion 2 passes narrowly or fails, and V1 yields fewer than 3
|
||
accepted clusters, failing criterion 1.**
|
||
- **MMIN, the signal designed to stop dilution, was not selected**, and W's cut (0.60) is loose enough
|
||
to act only on near-total single-word clusters. Growing a cluster raises Z1 and barely moves S1.
|
||
**Prediction: most V3 growths run to the cap, failing criterion 3.**
|
||
- The gain over the old gate is real but modest, and it is in the mid and large bands: 8 more clusters
|
||
labelled ≥ 1 accepted at unchanged precision.
|
||
- The thresholds are not changed in light of these predictions. The validation tests the rule the
|
||
pre-registered procedure produced.
|
||
|
||
### Validation results (2026-09-14)
|
||
|
||
73 fresh clusters labelled blind (`probe/labels-validate.json`: 6 labelled 2, 33 labelled 1, 34
|
||
labelled 0), scored by `probe/validate_gate2.py score` (log `probe/validate-score.log`).
|
||
|
||
| Set | Sampled (≥ 1) | Old gate: accepted / ≥ 1 / P1+ | G2: accepted / ≥ 1 / P1+ |
|
||
| --- | --- | --- | --- |
|
||
| V1, k = 100 | 30 (13) | 3 / 2 / 0.67 | **10 / 8 / 0.80** (2 labelled 2) |
|
||
| V2, k = 40 | 30 (16) | 1 / 1 / 1.00 | 4 / 3 / 0.75 |
|
||
| V3, grown | 13 | — | 13 grown: **11 stopped by the cap**, 2 by the gate; 10 ≥ 1, P1+ 0.77 |
|
||
|
||
| Criterion | Result |
|
||
| --- | --- |
|
||
| 1. P1+ ≥ 0.80 and ≥ 3 accepted, on V1 and V2 | **fail** — V1 passes (0.80, 10), V2 does not (0.75, 4) |
|
||
| 2. G2 accepts more V1 clusters than the old gate | **pass** — 10 against 3 |
|
||
| 3. V3 mostly stopped by the gate, P1+ ≥ 0.80 | **fail** — 2 of 13, 0.77 |
|
||
| **Overall** | **fail** |
|
||
|
||
By size, across V1 and V2 together:
|
||
|
||
| Size | Old gate: accepted (≥ 1) | G2: accepted (≥ 1) |
|
||
| --- | --- | --- |
|
||
| ≤ 9 | 0 | **0** — including two clusters labelled 2 (5 and 6 members) |
|
||
| 10–12 | 0 | 3 (3) |
|
||
| 13–25 | 3 (2) | 7 (5) |
|
||
| ≥ 26 | 1 (1) | 4 (3) |
|
||
| All | 4 (3), P1+ 0.75 | 14 (11), P1+ 0.79 |
|
||
|
||
### Verdict on G2
|
||
|
||
**G2 is a real improvement on the old gate, and it is not size-invariant.** It accepts three to four
|
||
times as many clusters at slightly better precision, and it opens the 10–25 member range the old gate
|
||
all but closed. Both pre-registered predictions of failure held:
|
||
|
||
- **The small end is still closed below 10 members.** Z1 ≥ 5 is unreachable for a cluster of 5–9
|
||
however cohesive it is. The two sharpest small clusters on the sheet — a fare dispute (5 members,
|
||
label 2) and a spouse against a spouse (6 members, label 2) — were both rejected.
|
||
- **Growth is still unconstrained.** Growing a cluster raises Z1 and barely moves S1, so both gates
|
||
accept a 45-member cluster as readily as a 15-member one. The one per-member signal built to stop
|
||
this, MMIN, was not selected by the tuning objective.
|
||
- **Large clusters are still mostly rejected** (4 of 30 in V2), including clear themes such as the
|
||
saloon, dining rooms and illness.
|
||
|
||
**What was learned about size-invariance itself:** a fixed threshold on a z-score is not size-invariant;
|
||
it reverses the old bias rather than removing it. A size-invariant cohesion test needs a threshold
|
||
that scales with the null — for example a fixed *percentile* of the same-size random distribution
|
||
combined with an effect-size floor on S1 — and a stopping rule for growth has to be a per-member
|
||
test that the tuning cannot switch off. Both are untested proposals, and testing them needs another
|
||
fresh partition and another round of blind labels, because every cluster labelled so far has now
|
||
informed a design decision.
|
||
|
||
---
|
||
|
||
## G3: a percentile of the same-size null, and a mandatory member floor (pre-registered 2026-09-14, before any G3 signal was computed)
|
||
|
||
### Changes from G2
|
||
|
||
- **Z1 is replaced by PCT**: the fraction of same-size random member sets (≥ 4 stories, dominant ≤ 40%)
|
||
whose S1 is below the cluster's. 1,000 samples per size for 5–60 members, 200 for 61–150, seeded.
|
||
A percentile threshold asks "is this cluster as unusual as the top x% of random sets *of its own
|
||
size*", which a z-score with a fixed bar does not.
|
||
- **S1 floor is mandatory.** At large sizes the null is narrow and PCT saturates at 1.0 for any
|
||
cluster slightly above average, so an effect-size floor has to carry the weight there.
|
||
- **MMIN is mandatory**: every member's mean cosine to members of other stories must reach the floor.
|
||
It is the growth stop; tuning chooses its value but cannot switch it off.
|
||
- **W** is unchanged from G2 and remains optional.
|
||
|
||
### Tuning
|
||
|
||
- **Tuning set: all 175 labelled clusters** — development 40, test 41, recovery 21, G2 validation 73.
|
||
- **Rule form:** S1 ≥ a AND PCT ≥ p AND MMIN ≥ c AND W < d. a and c at the tuning set's deciles (no "off");
|
||
p ∈ {0.90, 0.95, 0.975, 0.99, 0.999}; d ∈ {off, 0.30, 0.40, 0.50, 0.60}.
|
||
- **Objective unchanged from G2:** most clusters labelled ≥ 1 accepted, subject to precision (≥ 1) of at
|
||
least 0.85; ties on more label-2 accepted, then fewer active terms.
|
||
|
||
**Risk named in advance:** MMIN is a minimum over per-member means. Small clusters have noisier
|
||
member means but fewer of them; large clusters have steadier means but more chances to contain one
|
||
low member. Its net size effect is not predictable, and it could re-introduce a size bias of its own.
|
||
|
||
### Validation on fresh partitions (seeds never used before)
|
||
|
||
- **V0, very small:** k = 150, seed 20; 25 candidates sampled (seed 20).
|
||
- **V1, small:** k = 100, seed 17; 20 sampled (seed 17).
|
||
- **V2, large:** k = 40, seed 18; 20 sampled (seed 18).
|
||
- **V3, growth:** k = 60, seed 19; every G3-accepted candidate grown toward its nearest scene while it
|
||
passes G3, cap 45, stop reason recorded.
|
||
- Pooled, deduplicated, shuffled and labelled blind as before. The old gate and G2 are scored on the
|
||
same clusters.
|
||
|
||
**Success, all four required:**
|
||
|
||
1. G3 precision (≥ 1) ≥ **0.80** and ≥ 3 accepted, on V1 and on V2 separately.
|
||
2. **Small clusters:** among V0 and V1 clusters of ≤ 9 members, G3 accepts **at least 2**, with precision
|
||
(≥ 1) of at least **0.67**, and more than G2 accepts.
|
||
3. **Growth:** most V3 growths stopped by the gate, and grown clusters have precision (≥ 1) ≥ 0.80.
|
||
4. G3 accepts at least as many clusters labelled ≥ 1 across V0–V2 as G2, at precision no lower than G2's.
|
||
|
||
### Frozen G3 (written after tuning, before any validation cluster was built)
|
||
|
||
`probe/tune_gate3.py`, log `probe/tune-gate3.log`.
|
||
|
||
> **Accept a candidate when S1 ≥ 0.6175 AND PCT ≥ 0.90 AND MMIN ≥ 0.5296 AND W < 0.60.**
|
||
|
||
| Tuning set, 175 clusters (101 ≥ 1, 19 label 2) | Accepted | Label ≥ 1 | Precision (≥ 1) | Label 2 |
|
||
| --- | --- | --- | --- | --- |
|
||
| Old gate | 47 | 38 | 0.81 | 9 |
|
||
| G2 | 67 | 55 | 0.82 | 12 |
|
||
| **G3** | **47** | **40** | **0.85** | **13** |
|
||
|
||
| Size band | Clusters (≥ 1) | Old: accepted/≥ 1 | G2 | G3 |
|
||
| --- | --- | --- | --- | --- |
|
||
| ≤ 9 | 26 (5) | 0/0 | 0/0 | **4/3** |
|
||
| 10–12 | 21 (11) | 0/0 | 4/4 | 4/4 |
|
||
| 13–25 | 75 (47) | 17/15 | 26/22 | 17/17 |
|
||
| ≥ 26 | 53 (38) | 30/23 | 37/29 | 22/16 |
|
||
|
||
**Observations and predictions, recorded before validation:**
|
||
|
||
- **G3 is the first gate to accept any cluster under 10 members** (4 on tuning, 3 of them ≥ 1).
|
||
**Prediction: criterion 2 passes.**
|
||
- **The percentile threshold settled at its loosest value, 0.90, and the S1 floor rose** from G2's 0.6037
|
||
to 0.6175. The floor, not the percentile, is doing most of the work at mid and large sizes, and G3
|
||
accepts 15 fewer large clusters than G2 on tuning. **Prediction: criterion 4 fails** — G3 finds fewer
|
||
clusters labelled ≥ 1 than G2 (40 against 55 on tuning), though at higher precision. Criterion 1 is
|
||
at risk on V2 for the same reason.
|
||
- **MMIN at 0.53 is the growth stop.** Whether a nearest-to-centroid scene falls below it is not
|
||
predictable from tuning, which contains no grown-under-G3 clusters. No prediction for criterion 3.
|
||
- The trade is visible already: G3 exchanges recall at the large end for precision everywhere and for
|
||
admission at the small end. Whether that is the right trade for a pack is a design question the
|
||
validation cannot answer.
|
||
|
||
### Validation results (2026-09-14)
|
||
|
||
75 fresh clusters labelled blind (`probe/labels-validate3.json`: 4 labelled 2, 24 labelled 1, 47
|
||
labelled 0), scored by `probe/validate_gate3.py score` (log `probe/validate3-score.log`).
|
||
|
||
| Set | Sampled (≥ 1) | Old: acc / ≥ 1 / P1+ | G2: acc / ≥ 1 / P1+ | G3: acc / ≥ 1 / P1+ |
|
||
| --- | --- | --- | --- | --- |
|
||
| V0, k = 150 | 25 (5) | 2 / 1 / 0.50 | 8 / 4 / 0.50 | 11 / 4 / 0.36 |
|
||
| V1, k = 100 | 20 (5) | 1 / 0 / 0.00 | 1 / 1 / 1.00 | 4 / 2 / 0.50 |
|
||
| V2, k = 40 | 20 (12) | 0 / 0 / — | 4 / 4 / 1.00 | 2 / 2 / 1.00 |
|
||
| ≤ 9 members, V0 + V1 | 27 (4) | 1 / 0 / 0.00 | 3 / 1 / 0.33 | 11 / 3 / 0.27 |
|
||
| V0–V2 pooled | 65 (22) | 3 / 1 / 0.33 | **13 / 9 / 0.69** | 17 / 8 / 0.47 |
|
||
| V3, grown under G3 | 10 | — | — | 7 stopped by the cap, 3 by the gate; P1+ 0.60 |
|
||
|
||
**All four criteria fail**, and G3 is worse than G2 on fresh data: more clusters accepted, fewer of
|
||
them situations. G3's gain on the tuning set (P1+ 0.85, 13 label-2) did not survive a new partition.
|
||
|
||
### Why G3 failed — measured, not inferred
|
||
|
||
- **PCT carries no information.** Across all 175 tuning clusters, 173 reach PCT ≥ 0.90 and 141 reach
|
||
1.0 — including 72 of the 74 clusters labelled 0. **The random-set null is the wrong null.** k-means
|
||
groups each scene with its nearest neighbours, so every cluster it produces is far tighter than a
|
||
random set of scenes of the same size, whether or not it is a situation. A cohesion test has to ask
|
||
whether a cluster is tighter than *other clusters of its size*, not than random scenes.
|
||
- **MMIN barely separates labels and favours small clusters.** Median MMIN is 0.536 for label 0,
|
||
0.543 for label 1 and 0.564 for label 2; by size it is 0.547 for ≤ 9 members, 0.537 for 10–25 and
|
||
0.552 for ≥ 26. A minimum over few, noisy per-member means is not reliably lower for small clusters,
|
||
so the floor admitted 11 small clusters of which 8 are not situations. This is the risk recorded
|
||
before tuning.
|
||
- **Growth is still not stopped by the gate** (7 of 10 to the cap). A nearest-to-centroid scene
|
||
typically clears a member floor calibrated on whole clusters.
|
||
|
||
### Where this leaves selection
|
||
|
||
Across three gates and five blind-labelled evaluations, **G2 remains the best**: pooled precision
|
||
(≥ 1) of 0.79 on its own validation and 0.69 here, with the fewest non-situations admitted.
|
||
None of the three is size-invariant, and none stops growth.
|
||
|
||
Two conclusions look robust enough to build on:
|
||
|
||
1. **Cohesion alone does not identify a situation.** Clusters labelled 0, 1 and 2 overlap heavily on
|
||
every cohesion statistic tried (S1, Z1, PCT, MMIN). What separates the best-labelled clusters is
|
||
visible to a reader — one transaction — and is not a geometric property of summary embeddings at
|
||
the resolution tried.
|
||
2. **The labelling budget is the binding constraint.** Each gate needs fresh partitions and 70–80 new
|
||
blind labels to test honestly, and every labelled cluster so far has informed a design. Continuing
|
||
to iterate gate designs on this corpus without a different kind of evidence is unlikely to converge.
|
||
|
||
Candidate directions, none tested: a null built from **k-means clusters of the same size on
|
||
shuffled-story data** rather than random scenes; a selection step that uses the **inference model to
|
||
name each cluster's transaction** and checks agreement across members (needs the GPU); or accepting
|
||
G2 as the gate and moving the problem downstream to event construction, where a human pass over
|
||
~40 candidate clusters is cheap compared with the labelling this probe has already required.
|
||
|
||
---
|
||
|
||
## G4: the model reads the cluster (pre-registered 2026-09-14, before any model call)
|
||
|
||
The finding above is that what makes a cluster a situation is visible to a reader and not to the
|
||
geometry. So let a reader decide — the inference model — and measure it against the blind labels.
|
||
|
||
### Method
|
||
|
||
- For each cluster, up to **15 member summaries** (a seeded sample when larger), shuffled and numbered,
|
||
go to the model in one call with the prompt below. It returns JSON: the single social transaction
|
||
most of them clearly show, or NONE, and the numbers of the summaries that clearly show it.
|
||
- **FIT** = (number of valid summary numbers listed) / (number shown); 0 when the answer is NONE or
|
||
the reply cannot be parsed. Parse failures are counted and reported.
|
||
- **Gate: accept when FIT ≥ θ**, plus the usual candidate conditions (≥ 5 members, ≥ 4 stories,
|
||
dominant ≤ 40%), which every labelled cluster already meets.
|
||
- Deterministic: temperature 0, 2,048-token context, `format: json`. One request at a time, as before.
|
||
- **The prompt contains no example of a situation** — the 3B model copied the examples it was given
|
||
in the summarisation pilot.
|
||
|
||
Prompt (system), fixed:
|
||
|
||
> You read short summaries of scenes from different stories and decide whether they share one kind
|
||
> of social situation.
|
||
>
|
||
> A social situation is a specific transaction between people: who wants what from whom, and what
|
||
> is at stake. Describe it in your own words.
|
||
>
|
||
> Rules:
|
||
> - Find the single situation that the largest number of summaries clearly show.
|
||
> - It is normal for many summaries not to fit. Include a summary only if it clearly shows that
|
||
> situation.
|
||
> - A shared place, mood, profession or word is not a situation.
|
||
> - If no situation is clearly shown by at least three summaries, the situation is NONE.
|
||
>
|
||
> Reply with JSON only: {"situation": "<at most 10 words, or NONE>", "fits": [<numbers of the
|
||
> summaries that clearly show it>]}
|
||
|
||
The user message is the numbered list. For `qwen3:14b` it ends with `/no_think`, and any think block
|
||
is stripped before parsing.
|
||
|
||
### Models
|
||
|
||
- **Primary: `qwen2.5:3b-instruct`** — the design requires a small local model; the verdict rests on it.
|
||
- **Secondary: `qwen3:14b`** — reported, to show whether capability is what limits the result.
|
||
|
||
### Data — no new labels
|
||
|
||
The model is never fitted to labels; the only fitted quantity is θ.
|
||
|
||
- **θ is chosen on development (40 clusters)** only, with the objective used for G2 and G3: most
|
||
clusters labelled ≥ 1 accepted, subject to precision (≥ 1) ≥ 0.85; ties on more label-2 accepted,
|
||
then the higher θ. Grid: θ ∈ {0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0}.
|
||
- **Primary evaluation: the G2 and G3 validation sets, 148 clusters**, which tuned neither G2 nor G4,
|
||
so G2 can be compared fairly on them. **Secondary:** test and recovery, 62 clusters.
|
||
- Identical member sets appearing in more than one file are judged once.
|
||
- **Pilot, allowed once:** the first 5 development clusters, to check the JSON parses. If more than 1
|
||
of 5 fails to parse, one prompt revision is allowed and recorded here before any other call.
|
||
|
||
### Success, on the 148-cluster primary set, with the 3B model
|
||
|
||
1. Precision (≥ 1) at least **0.80**.
|
||
2. **More** clusters labelled ≥ 1 accepted than G2 accepts on the same set.
|
||
3. Among accepted clusters of ≤ 9 members, precision (≥ 1) at least 0.67, if there are at least 3.
|
||
|
||
Also reported: G4 AND G2; FIT distribution by label; the situations named for accepted clusters
|
||
against the label names; and the 14B results against the same bar.
|
||
|
||
**Pilot (21:23):** 5 of 5 replies parsed with the 3B model, so the prompt is unchanged. About 1.6 s
|
||
per cluster. Noted without drawing a conclusion from five: the situations named are often generic
|
||
("Confrontation and hidden motives" for D04, a label-2 spouse-against-spouse cluster, FIT 0.47), and
|
||
a label-1 argument cluster scored FIT 0.87.
|
||
|
||
### Results, `qwen2.5:3b-instruct` (21:24–21:29, 245 distinct clusters in 286 s, 1 parse failure)
|
||
|
||
Log `probe/evaluate-judge-3b.log`. θ chosen on development: **0.6** (9 accepted, P1+ 0.89).
|
||
|
||
| Primary set, 148 clusters (67 ≥ 1) | Accepted | ≥ 1 | Label 2 | P1+ | Recall |
|
||
| --- | --- | --- | --- | --- | --- |
|
||
| Accept all | 148 | 67 | 10 | 0.45 | 1.00 |
|
||
| **G2** | **49** | **36** | 5 | **0.73** | **0.54** |
|
||
| G4, FIT ≥ 0.6 | 50 | 24 | 4 | 0.48 | 0.36 |
|
||
| G4 AND G2 | 18 | 11 | 1 | 0.61 | 0.16 |
|
||
|
||
On the secondary set (62) G4 reaches P1+ 0.71 against G2's 0.82. **All three criteria fail**
|
||
(P1+ 0.48; 24 ≥ 1 against G2's 36; 13 accepted small clusters at P1+ 0.23).
|
||
|
||
**FIT does not separate the labels:** mean 0.39 for label 0, 0.49 for label 1 — and **0.42 for label
|
||
2**, lower than for label 1. Precision at the primary set (0.48) is barely above accepting everything
|
||
(0.45).
|
||
|
||
**Why, visible in the replies:** the 3B model names situations at a level of abstraction where almost
|
||
anything fits — "confrontation", "power dynamics in social hierarchies", "tension between parties",
|
||
"coercion and manipulation" — and then lists most summaries as fitting it. A 5-member cluster labelled
|
||
0 scored FIT 1.00 as "power struggle/conflict". When it is specific it is often right ("marital
|
||
conflict" for a spouse-against-spouse cluster, "con or scam", "argue over payment"), but specificity
|
||
is not what FIT rewards. The rule against shared place, mood and profession was followed; a rule
|
||
against *abstraction* was never written, and the model supplies abstraction freely.
|
||
|
||
### 14B pilot and the one allowed revision
|
||
|
||
The `qwen3:14b` pilot failed to parse **5 of 5**: every reply's content was empty. The model spent its
|
||
whole 200-token budget in a separate thinking channel (Ollama returns it outside `content`), and the
|
||
`/no_think` suffix did not stop it. **Revision, the one allowed by the pilot rule, affecting only the
|
||
14B run:** requests set Ollama's `think: false` and allow 300 output tokens. The prompt text,
|
||
sampling, θ procedure and success criteria are unchanged. The five empty results were deleted before
|
||
re-running the pilot.
|
||
|
||
Re-run pilot: **5 of 5 parsed**, about 2.4 s per cluster. The situations named are more specific than
|
||
the 3B model's ("Wealthy person pressures someone into marriage", "Struggling artist facing societal or
|
||
personal challenges"), though an abstract one still scored FIT 0.80 ("conflict between individuals with
|
||
differing perspectives").
|
||
|
||
### Results, `qwen3:14b` (21:31–21:45, 245 clusters in 832 s, 0 parse failures)
|
||
|
||
Log `probe/evaluate-judge-14b.log`. θ chosen on development: **0.6** (8 accepted, P1+ 0.88).
|
||
|
||
| Primary set, 148 clusters (67 ≥ 1) | Accepted | ≥ 1 | Label 2 | P1+ | Recall |
|
||
| --- | --- | --- | --- | --- | --- |
|
||
| Accept all | 148 | 67 | 10 | 0.45 | 1.00 |
|
||
| **G2** | **49** | **36** | 5 | **0.73** | **0.54** |
|
||
| G4 (14B), FIT ≥ 0.6 | 42 | 22 | 4 | 0.52 | 0.33 |
|
||
| G4 (14B) AND G2 | 10 | 10 | 1 | 1.00 | 0.15 |
|
||
|
||
By size, G4 (14B) on the primary set: ≤ 9 members 17 accepted, P1+ 0.29; 10–25, 14 accepted, P1+ 0.50;
|
||
≥ 26, 11 accepted, **P1+ 0.91**. Secondary set (62): G4 P1+ 0.60 against G2's 0.82.
|
||
|
||
**All three criteria fail with the 14B model as well** (P1+ 0.52; 22 ≥ 1 against G2's 36; 17 small
|
||
clusters accepted at P1+ 0.29).
|
||
|
||
FIT now rises with the label — mean 0.40, 0.47, 0.52 for labels 0, 1, 2 — where the 3B model's did
|
||
not, so capability helps. But the separation is small next to the spread, and the 14B model still
|
||
names abstractions that fit almost anything ("Someone wants something from someone else", "Power
|
||
dynamics between individuals", "Someone holds power over someone else"), most often for small
|
||
clusters.
|
||
|
||
GPU host during both runs: no faults logged; peak 284.9 W (1 s), mean 163 W under load, 65 °C, PCIe
|
||
Gen 3 x4 throughout. Logs in `probe/gpu-logs/*2122*`.
|
||
|
||
### Verdict on G4
|
||
|
||
**Neither model's judgement is a better gate than G2**, and the pre-registered verdict (on the 3B
|
||
model) is a fail. The failure is specific and repeatable: asked for "the single situation most
|
||
summaries share", a model answers at whatever level of abstraction makes the most summaries fit, and
|
||
FIT rewards exactly that. Small clusters make it worse, because with 5–9 summaries an abstraction that
|
||
covers four or five of them is always available.
|
||
|
||
Two observations, recorded as observations because neither was a pre-registered criterion:
|
||
|
||
- **G4 (14B) AND G2 accepted 10 clusters on the primary set, and all 10 are labelled ≥ 1.** The two
|
||
gates fail in different ways — G2 admits cohesive clusters that are not situations, the model admits
|
||
abstract "situations" that are not cohesive — and their intersection is precise. It is also
|
||
narrow: recall 0.15.
|
||
- **For clusters of 26 or more, the 14B model's precision is 0.91.** The abstraction problem is a
|
||
small-cluster problem: 15 shown summaries rarely all fit one vague phrase *and* get listed.
|
||
|
||
What would test the obvious next step, none of it done: a prompt that asks for the situation **in
|
||
terms of roles and a concrete want** and rejects answers without both; scoring the *specificity* of
|
||
the named situation (for example, requiring the named situation to be closer to the member summaries
|
||
than to the corpus average in embedding space); and G2 AND G4 evaluated as a pre-registered gate on a
|
||
fresh partition, since every labelled cluster now informs this observation.
|