Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
This commit is contained in:
co-authored by
Claude Opus 5
parent
fa3769d0fe
commit
e5617b86ba
@@ -12,28 +12,42 @@ versioned alongside the pack format it targets, without touching `main`.
|
||||
|
||||
## Status
|
||||
|
||||
Selection is settled as **gate G2 followed by a structured human review** (see below). Every automated
|
||||
Selection was settled as **gate G2 followed by a structured human review** (see below). Every automated
|
||||
alternative tried — three more geometric gates and two model-judged gates — was worse on fresh,
|
||||
blind-labelled data. Nothing downstream of selection (paths and directions, numbers, pack output) is
|
||||
built yet.
|
||||
blind-labelled data. The first real review then found G2's groups were **topics** (artists, police,
|
||||
hotels), not situations. **Now under test:** re-describe every scene as its main person's
|
||||
**predicament** with qwen3:14b, regroup, and read the groups (SELECTION.md, "Predicaments"). Nothing
|
||||
downstream of selection (paths and directions, numbers, pack output) is built yet.
|
||||
|
||||
## Layout
|
||||
|
||||
`probe/` holds the v1 data every G1–G4 result was measured on. `probe/v2/` is the corrected corpus
|
||||
(97 stories, 839 scenes: v1's splitter merged three quoted-title stories into their neighbours), built by
|
||||
running the same scripts from inside `v2/` (`python3 ../split.py`, and so on).
|
||||
|
||||
```
|
||||
probe/
|
||||
corpus/ O. Henry's Four Million cycle, Project Gutenberg (public domain)
|
||||
split.py, chunk.py volumes -> stories.json -> chunks.json (838 scenes)
|
||||
split.py, chunk.py volumes -> stories.json -> chunks.json (838 scenes in v1, 839 in v2)
|
||||
carry_summaries.py v1 summaries onto v2 scenes; only new scene text goes to the model
|
||||
redescribe.py scene -> its main person's predicament, via the model -> predicaments.json
|
||||
topic_words.py, topic_share.py how much groups are held together by a job, relationship or place word
|
||||
pilot_page.py side-by-side page for the 3B / 14B predicament pilot -> pilot-compare.html
|
||||
groups_page.py reading page for regrouped predicaments -> v2/groups-*.html
|
||||
ollama.py the only module that talks to an inference host (host from STP_OLLAMA)
|
||||
summarise.py scene -> one sentence, via the model -> summaries.json
|
||||
strip_names.py names removed deterministically -> summaries_clean.json
|
||||
embed.py texts -> vectors -> embeddings.json, summary_embeddings.json
|
||||
measure.py, signals.py, gate2.py, gate3.py, judge.py … the selection experiments
|
||||
labels-*.json, sheet-*.txt, candidates-*.json, *-map.json blind labels and what they labelled
|
||||
review.py G2 on a partition -> a review package for a human
|
||||
review_format.py the review form: schema, checks, conversion
|
||||
apply_review.py filled review.json -> situations.json
|
||||
test_review.py tests for the review form
|
||||
review/<run>/ review packages: REVIEW.md, review.json, candidates.json, situations.json
|
||||
review.py G2 on a partition -> REVIEW.html and candidates.json
|
||||
review_page.py renders the page; picks out role words for each group
|
||||
review_page.html the page template
|
||||
review_page_logic.cjs the rules the page applies (also run by the Node tests)
|
||||
review_format.py the saved review's format, checks and conversion
|
||||
apply_review.py saved review -> situations.json
|
||||
test_review.py, test_review_page.mjs tests
|
||||
review/<run>/ REVIEW.html, candidates.json, and situations.json once applied
|
||||
```
|
||||
|
||||
Not committed (see `.gitignore`): the inference host's power and kernel logs, and caches that the
|
||||
@@ -52,29 +66,38 @@ export STP_OLLAMA=http://<inference-host>:11434
|
||||
|
||||
## Reviewing situations
|
||||
|
||||
The review happens in a web page. Nobody edits JSON.
|
||||
|
||||
```sh
|
||||
cd probe
|
||||
python3 review.py # k = 60, seed 24 -> review/k60-s24/
|
||||
python3 review.py # k = 60, seed 24 -> review/k60-s24/REVIEW.html
|
||||
```
|
||||
|
||||
That writes three files into the run directory:
|
||||
1. **Open `review/k60-s24/REVIEW.html` in a browser** (double-click it; nothing to install). The first
|
||||
screen explains the job with examples.
|
||||
2. **One group per screen.** For each: keep it, drop it, or mark it the same as a group already kept;
|
||||
untick scenes that don't fit; and fill in "A ___ wants ___ from ___", with the people named in its
|
||||
scenes offered as click-to-fill words. The page shows as you go whether a group still has enough
|
||||
scenes, and colours each group in the strip across the top. Keys: K keep, D drop, ← →.
|
||||
3. **Progress saves in the browser** after every change, so you can close it and come back.
|
||||
4. **"Finish & save"** lists anything unfinished and saves `review-k60-s24.json`, usually to Downloads.
|
||||
5. **Then:**
|
||||
|
||||
- **`REVIEW.md`** — what the program found, and exactly what the reviewer is asked to decide for
|
||||
each candidate, with every scene listed. Read this.
|
||||
- **`review.json`** — the form. The only file the reviewer edits.
|
||||
- **`candidates.json`** — what the program found, for the checker. Do not edit.
|
||||
```sh
|
||||
python3 apply_review.py review/k60-s24
|
||||
```
|
||||
|
||||
When the form is filled in:
|
||||
It finds the saved review in Downloads by itself, checks everything again, and names any problem the
|
||||
way the page does ("Group 4: kept, but its name is not finished"). Nothing is written until the
|
||||
review passes; then it writes **`situations.json`**: each situation's sentence and roles, the scenes
|
||||
that show it with their stories, and which groups it came from.
|
||||
|
||||
`candidates.json` in the run directory is the program's record for the checker; leave it alone.
|
||||
|
||||
```sh
|
||||
python3 apply_review.py review/k60-s24
|
||||
python3 -m unittest test_review # the saved review's checks, the page renderer, apply_review.py
|
||||
node --test test_review_page.mjs # the rules the page applies as you review
|
||||
```
|
||||
|
||||
It reports every problem at once and writes nothing until the form is complete and consistent; then
|
||||
it writes **`situations.json`**, the machine-readable input for the next stage: each situation's
|
||||
wording, roles and stakes, the scenes that show it (with story and summary), and which candidates it
|
||||
came from.
|
||||
|
||||
```sh
|
||||
python3 -m unittest test_review # the review form's checks
|
||||
```
|
||||
The page's drawing code is only checked for syntax and for starting up; how it looks and behaves in a
|
||||
real browser has to be tried by a person.
|
||||
|
||||
Reference in New Issue
Block a user