Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
This commit is contained in:
co-authored by
Claude Opus 5
parent
fa3769d0fe
commit
e5617b86ba
@@ -0,0 +1,58 @@
|
||||
"""Render REVIEW.html, the page a reviewer works in, and pick out role words for each group.
|
||||
|
||||
No clustering code, so the tests load it instantly. The page is one self-contained file:
|
||||
the template (review_page.html) with the rules (review_page_logic.cjs) and the run's data
|
||||
inlined, so it opens by double-clicking with no server and nothing installed.
|
||||
"""
|
||||
import collections, html, json, pathlib, re
|
||||
import review_format as rf
|
||||
|
||||
HERE = pathlib.Path(__file__).resolve().parent
|
||||
|
||||
# People the summaries mention by role. Offered as click-to-fill words on the page, because
|
||||
# naming the two people in a situation was the hardest part of the first review. A fixed list
|
||||
# rather than anything clever: it only has to suggest, and the reviewer can type anything.
|
||||
ROLE_WORDS = set("""
|
||||
wife husband daughter son father mother uncle aunt nephew niece brother sister cousin friend
|
||||
suitor fiance fiancee bride groom lover admirer rival sweetheart widow widower
|
||||
boss employer employee clerk typist stenographer secretary partner manager foreman
|
||||
cop policeman officer captain detective marshal judge lawyer
|
||||
landlady landlord tenant lodger boarder roomer housekeeper
|
||||
waiter waitress cook chef bartender proprietor owner shopkeeper storekeeper merchant salesman
|
||||
saleslady shopgirl customer patron guest host servant maid butler chambermaid
|
||||
driver cabby cabman coachman chauffeur
|
||||
artist painter poet writer author editor reporter journalist actor actress singer dancer model
|
||||
doctor nurse patient preacher teacher student professor
|
||||
stranger neighbour neighbor visitor newcomer millionaire banker broker gambler thief crook
|
||||
burglar tramp beggar vagrant panhandler worker girl woman man boy child
|
||||
""".split())
|
||||
|
||||
|
||||
def role_hints(members, limit=8):
|
||||
"""The role words that turn up most in a group's summaries."""
|
||||
counts = collections.Counter()
|
||||
for m in members:
|
||||
for word in re.findall(r"[a-z]+", m['summary'].lower()):
|
||||
if word in ROLE_WORDS:
|
||||
counts[word] += 1
|
||||
return [word for word, _ in counts.most_common(limit)]
|
||||
|
||||
|
||||
def page_data(document):
|
||||
return {'run': document['run'], 'rules': rf.RULES, 'candidates': document['candidates']}
|
||||
|
||||
|
||||
def render_page(document):
|
||||
template = (HERE / 'review_page.html').read_text(encoding='utf-8')
|
||||
logic = (HERE / 'review_page_logic.cjs').read_text(encoding='utf-8')
|
||||
for marker in ('/*__LOGIC__*/', '/*__DATA__*/null', '__RUN_ID__'):
|
||||
if marker not in template:
|
||||
raise ValueError(f'review_page.html is missing the {marker} marker')
|
||||
# Data inside a <script> must not be able to close it: a summary containing "</script>"
|
||||
# would end the script early. Escaping "</" is enough and is still valid JSON.
|
||||
data = json.dumps(page_data(document), ensure_ascii=False).replace('</', '<\\/')
|
||||
logic = logic.replace('</script', '<\\/script')
|
||||
return (template
|
||||
.replace('/*__LOGIC__*/', logic)
|
||||
.replace('/*__DATA__*/null', data)
|
||||
.replace('__RUN_ID__', html.escape(document['run']['id'])))
|
||||
Reference in New Issue
Block a user