Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
This commit is contained in:
co-authored by
Claude Opus 5
parent
fa3769d0fe
commit
e5617b86ba
@@ -1,20 +1,21 @@
|
||||
"""Build a human-review package: the clusters gate G2 accepts, plus the nearest misses.
|
||||
"""Build a review page: the groups gate G2 accepts, plus the nearest misses.
|
||||
|
||||
python3 review.py # k = 60, seed 24 -> review/k60-s24/
|
||||
python3 review.py --k 60 --seed 31 --borderline 8
|
||||
|
||||
Writes into the run directory:
|
||||
REVIEW.md what was found, and exactly what the reviewer decides for each candidate
|
||||
review.json the form to fill in: the only file the reviewer edits
|
||||
candidates.json what was found, for apply_review.py to check the form against (do not edit)
|
||||
REVIEW.html open it in a browser: the whole review happens there
|
||||
candidates.json what was found, for apply_review.py to check the saved review against
|
||||
|
||||
Then: python3 apply_review.py review/k60-s24 -> situations.json
|
||||
When the review is saved from the page:
|
||||
python3 apply_review.py review/k60-s24 -> situations.json
|
||||
|
||||
No inference: this uses the summaries and embeddings already on disk.
|
||||
"""
|
||||
import argparse, datetime, json, pathlib, sys
|
||||
import argparse, datetime, hashlib, json, pathlib
|
||||
import gate2 as g
|
||||
import review_format as rf
|
||||
import review_page
|
||||
|
||||
|
||||
def gate_failures(sig, gate):
|
||||
@@ -31,7 +32,7 @@ def gate_failures(sig, gate):
|
||||
|
||||
|
||||
def shortfall(sig, gate):
|
||||
"""How far a rejected candidate is from passing; smaller is closer."""
|
||||
"""How far a rejected group is from passing; smaller is closer."""
|
||||
s = 0.0
|
||||
if gate.get('S1') is not None: s += max(0.0, gate['S1'] - sig['S1']) / 0.01
|
||||
if gate.get('Z1') is not None: s += max(0.0, gate['Z1'] - sig['Z1'])
|
||||
@@ -39,160 +40,60 @@ def shortfall(sig, gate):
|
||||
return s
|
||||
|
||||
|
||||
def cell(text):
|
||||
return text.replace('|', '\\|').replace('\n', ' ')
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
ap.add_argument('--k', type=int, default=60, help='number of k-means clusters (default 60)')
|
||||
ap.add_argument('--seed', type=int, default=24, help='k-means seed (default 24)')
|
||||
ap.add_argument('--borderline', type=int, default=8, help='near misses to offer for optional review (default 8)')
|
||||
ap.add_argument('--borderline', type=int, default=8, help='near misses offered as optional extras (default 8)')
|
||||
ap.add_argument('--out', help='run directory (default review/k<K>-s<SEED>)')
|
||||
ap.add_argument('--force', action='store_true', help='overwrite an existing review.json')
|
||||
args = ap.parse_args()
|
||||
|
||||
run_id = f'k{args.k}-s{args.seed}'
|
||||
out = pathlib.Path(args.out or f'review/{run_id}')
|
||||
if (out / 'review.json').exists() and not args.force:
|
||||
sys.exit(f'{out / "review.json"} already exists and may hold a reviewer\'s work. '
|
||||
f'Use --force to overwrite it, or --out for a different directory.')
|
||||
|
||||
gate = json.loads(pathlib.Path('gate2-frozen.json').read_text())
|
||||
candidates_found = [c for c in g.kmeans(args.k, args.seed) if g.eligible(c)]
|
||||
scored = [(c, g.signals(c)) for c in candidates_found]
|
||||
found = [c for c in g.kmeans(args.k, args.seed) if g.eligible(c)]
|
||||
scored = [(c, g.signals(c)) for c in found]
|
||||
accepted = sorted([cs for cs in scored if g.g2_passes(cs[1], gate)], key=lambda cs: -cs[1]['S1'])
|
||||
rejected = sorted([cs for cs in scored if not g.g2_passes(cs[1], gate)], key=lambda cs: shortfall(cs[1], gate))
|
||||
borderline = rejected[:args.borderline]
|
||||
|
||||
listed = [(f'C{i + 1:02d}', 'accepted', c, s) for i, (c, s) in enumerate(accepted)]
|
||||
listed += [(f'B{i + 1:02d}', 'borderline', c, s) for i, (c, s) in enumerate(borderline)]
|
||||
centroids = {cid: g.centroid(c) for cid, _, c, _ in listed}
|
||||
|
||||
candidates = {}
|
||||
for cid, kind, members, sig in listed:
|
||||
others = [(g.dot(centroids[cid], centroids[o]), o) for o in centroids if o != cid]
|
||||
cos, similar = max(others) if others else (0.0, None)
|
||||
member_rows = [{'member': rf.member_id(cid, n + 1), 'chunk': chunk, 'story': g.chunks[chunk]['story'],
|
||||
'title': g.chunks[chunk]['title'].title(), 'summary': g.text[chunk]}
|
||||
for n, chunk in enumerate(sorted(members))]
|
||||
candidates[cid] = {
|
||||
'kind': kind,
|
||||
'stats': {'scenes': sig['size'], 'stories': sig['stories'], 'largest_story_share': round(sig['dominant'], 3),
|
||||
'S1': round(sig['S1'], 4), 'Z1': round(sig['Z1'], 2), 'W': round(sig['W'], 3), 'W_word': sig['W_word']},
|
||||
'gate_failures': gate_failures(sig, gate),
|
||||
'most_similar': {'id': similar, 'cosine': round(cos, 3)},
|
||||
'members': [{'member': rf.member_id(cid, n + 1), 'chunk': chunk, 'story': g.chunks[chunk]['story'],
|
||||
'title': g.chunks[chunk]['title'], 'summary': g.text[chunk]}
|
||||
for n, chunk in enumerate(sorted(members))],
|
||||
'role_hints': review_page.role_hints(member_rows),
|
||||
'members': member_rows,
|
||||
}
|
||||
# Ties a saved review to exactly these groups, so progress for a different set of groups
|
||||
# that happens to share the run name is never loaded into this page.
|
||||
fingerprint = hashlib.sha1(json.dumps({cid: [m['chunk'] for m in c['members']] for cid, c in candidates.items()},
|
||||
sort_keys=True).encode()).hexdigest()[:12]
|
||||
document = {
|
||||
'schema_version': rf.SCHEMA_VERSION,
|
||||
'run': {'id': run_id, 'k': args.k, 'seed': args.seed,
|
||||
'run': {'id': run_id, 'k': args.k, 'seed': args.seed, 'fingerprint': fingerprint,
|
||||
'created': datetime.datetime.now().astimezone().isoformat(timespec='seconds'),
|
||||
'gate': {'name': 'G2', **gate}, 'scenes': g.n, 'stories': len(set(g.story)),
|
||||
'candidates_found': len(candidates_found), 'accepted': len(accepted), 'borderline': len(borderline)},
|
||||
'candidates_found': len(found), 'accepted': len(accepted), 'borderline': len(borderline)},
|
||||
'candidates': candidates,
|
||||
}
|
||||
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
(out / 'candidates.json').write_text(json.dumps(document, indent=1, ensure_ascii=False), encoding='utf-8')
|
||||
(out / 'review.json').write_text(json.dumps(rf.blank_form(document), indent=1, ensure_ascii=False), encoding='utf-8')
|
||||
(out / 'REVIEW.md').write_text(render(document, out), encoding='utf-8')
|
||||
(out / 'REVIEW.html').write_text(review_page.render_page(document), encoding='utf-8')
|
||||
|
||||
print(f'{run_id}: {len(candidates_found)} candidates, G2 accepts {len(accepted)}, '
|
||||
f'{len(borderline)} borderline offered')
|
||||
print(f'wrote {out}/REVIEW.md, {out}/review.json, {out}/candidates.json')
|
||||
print(f'next: fill in {out}/review.json, then python3 apply_review.py {out}')
|
||||
|
||||
|
||||
def render(doc, out):
|
||||
run, cands = doc['run'], doc['candidates']
|
||||
accepted = [cid for cid, c in cands.items() if c['kind'] == 'accepted']
|
||||
borderline = [cid for cid, c in cands.items() if c['kind'] == 'borderline']
|
||||
gate = run['gate']
|
||||
expected_rejects = round(len(accepted) * 0.27)
|
||||
L = []
|
||||
L.append(f'# Situation review — run {run["id"]}\n')
|
||||
L.append(f'Generated {run["created"]} by `probe/review.py` (k-means k = {run["k"]}, seed {run["seed"]}, gate G2).\n')
|
||||
L.append('## What the program found\n')
|
||||
L.append(f'- {run["scenes"]} scenes from {run["stories"]} stories, grouped into {run["k"]} clusters.')
|
||||
L.append(f'- {run["candidates_found"]} clusters are large and varied enough to be candidates '
|
||||
f'(at least {rf.MIN_MEMBERS} scenes from at least {rf.MIN_STORIES} stories, no story above {rf.MAX_DOMINANT:.0%}).')
|
||||
L.append(f'- **{len(accepted)} candidates passed the automatic gate (G2): {", ".join(accepted) or "none"}.** '
|
||||
f'On blind-labelled test data about 3 in 4 clusters that G2 accepts are real situations, '
|
||||
f'so expect to reject roughly {expected_rejects} of these.')
|
||||
L.append(f'- {len(borderline)} more candidates narrowly failed the gate: {", ".join(borderline) or "none"}. '
|
||||
f'They are at the end. **Reviewing them is optional.**\n')
|
||||
L.append('## What you are asked to do\n')
|
||||
L.append(f'Budget about a minute per candidate. For each candidate below, read its scenes, then fill in '
|
||||
f'its entry in **`review.json`** in this directory. That is the only file to edit; '
|
||||
f'`candidates.json` is the program\'s record and must stay as it is.\n')
|
||||
L.append('For each candidate, answer:\n')
|
||||
L.append('1. **Do most of these scenes show one recurring situation** — one person wanting something '
|
||||
'specific from another? → `decision`')
|
||||
L.append('2. **If so, what is it?** One sentence: who wants what from whom. → `situation`, `actor`, `counterpart`, `stakes`')
|
||||
L.append('3. **Which scenes do not show it?** → `exclude`')
|
||||
L.append('4. **Is it the same situation as another candidate?** → `decision: "merge"` and `merge_into`\n')
|
||||
L.append('| Field | What to put | When |')
|
||||
L.append('| --- | --- | --- |')
|
||||
L.append('| `decision` | `"accept"` — most scenes show one situation. `"reject"` — they do not. `"merge"` — the same situation as another candidate you accepted. Near misses (B…) start as `"skip"` and may be changed to `"accept"` or `"merge"`. | always |')
|
||||
L.append('| `situation` | One sentence, at most 200 characters: who wants what from whom. | accept |')
|
||||
L.append('| `actor` | The role that wants something — a role such as "a lodger", never a character\'s name. | accept |')
|
||||
L.append('| `counterpart` | The role they want it from. | accept |')
|
||||
L.append('| `stakes` | What is won or lost, in a few words. | asked for; not required |')
|
||||
L.append('| `exclude` | Scene ids that do **not** show the situation, e.g. `["C04.02", "C04.11"]`. | when some do not fit |')
|
||||
L.append('| `merge_into` | The id of the accepted candidate this one duplicates. | merge |')
|
||||
L.append('| `notes` | Anything worth keeping. | optional |\n')
|
||||
L.append('Guidance:\n')
|
||||
L.append('- **A shared place, mood, job or word is not a situation.** "Scenes in restaurants" is a reject; '
|
||||
'"a diner wants credit from a proprietor who wants payment" is an accept.')
|
||||
L.append('- **Stay concrete.** "Conflict", "power dynamics" and "someone wants something" describe everything; '
|
||||
'if that is the best you can say, reject.')
|
||||
L.append(f'- **An accepted situation needs at least {rf.MIN_MEMBERS} scenes from at least {rf.MIN_STORIES} stories '
|
||||
f'after exclusions, with no story above {rf.MAX_DOMINANT:.0%}.** The checker tells you if exclusions break that.')
|
||||
L.append('- **"Most similar" is a hint from the embeddings and is often wrong.** Merge only when the situations are the same.')
|
||||
L.append('- Scenes are machine summaries with names replaced by "someone"; the story title is the original.\n')
|
||||
L.append('An invented example of a completed entry (not from this corpus):\n')
|
||||
L.append('```json')
|
||||
L.append(json.dumps({'id': 'C00', 'kind': 'accepted', 'decision': 'accept',
|
||||
'situation': 'A junior clerk wants a raise from an employer who wants more work for the same pay.',
|
||||
'actor': 'a junior clerk', 'counterpart': 'an employer', 'stakes': 'his wage and his job',
|
||||
'exclude': ['C00.04'], 'merge_into': None, 'notes': ''}, indent=1))
|
||||
L.append('```\n')
|
||||
L.append('## When you are done\n')
|
||||
L.append('```sh')
|
||||
L.append(f'python3 apply_review.py {out}')
|
||||
L.append('```\n')
|
||||
L.append('It checks the whole form and lists **every** problem at once — a missing decision, a scene id that '
|
||||
'is not in that candidate, a merge into something you rejected, exclusions that leave too few '
|
||||
'stories. Nothing is written until the form passes. Then it writes **`situations.json`**: each '
|
||||
'situation\'s wording, roles and stakes, the scenes that show it with their stories, and the '
|
||||
'candidates it came from. That file is what the next stage reads.\n')
|
||||
|
||||
def section(cid):
|
||||
c = cands[cid]
|
||||
st = c['stats']
|
||||
L.append(f'### {cid} — {st["scenes"]} scenes from {st["stories"]} stories\n')
|
||||
if c['kind'] == 'accepted':
|
||||
L.append(f'Passed the gate: cohesion S1 {st["S1"]:.3f} (needs {gate["S1"]:.3f}), Z1 {st["Z1"]:.1f} '
|
||||
f'(needs {gate["Z1"]}); largest story share {st["largest_story_share"]:.0%}. ')
|
||||
else:
|
||||
L.append('Missed the gate: ' + '; '.join(c['gate_failures']) + '. ')
|
||||
sim = c['most_similar']
|
||||
if sim['id']:
|
||||
L.append(f'Most similar candidate: {sim["id"]} (cosine {sim["cosine"]:.2f} — a hint only).\n')
|
||||
L.append('| Scene | Story | Summary |')
|
||||
L.append('| --- | --- | --- |')
|
||||
for m in c['members']:
|
||||
L.append(f'| {m["member"]} | {cell(m["title"].title())} | {cell(m["summary"])} |')
|
||||
L.append('')
|
||||
|
||||
L.append('## Candidates that passed the gate — a decision is required\n')
|
||||
for cid in accepted:
|
||||
section(cid)
|
||||
L.append('## Near misses — optional\n')
|
||||
L.append('These failed the gate narrowly. Leave them as `"skip"` unless one clearly shows a situation.\n')
|
||||
for cid in borderline:
|
||||
section(cid)
|
||||
return '\n'.join(L) + '\n'
|
||||
print(f'{run_id}: {len(found)} candidate groups, G2 accepts {len(accepted)}, {len(borderline)} optional extras')
|
||||
print(f'open this in a browser: {(out / "REVIEW.html").resolve()}')
|
||||
print(f'when the review is saved: python3 apply_review.py {out}')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
|
||||
Reference in New Issue
Block a user