The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
22 lines
1.2 KiB
Python
22 lines
1.2 KiB
Python
"""Words that tie a scene to a topic: who is in it and where it happens.
|
|
|
|
The first human review found the summary groups were topics (artists, police, courtship,
|
|
hotels) rather than situations. These words are how that is measured, and how leakage into
|
|
re-described predicaments is caught. A fixed list: it only has to catch the common cases.
|
|
"""
|
|
from review_page import ROLE_WORDS
|
|
|
|
PLACE_WORDS = set("""
|
|
hotel restaurant cafe café saloon bar store shop office park street square church theatre theater
|
|
stage farm ranch camp mine court courtroom jail cell station train boat ship ferry house room flat
|
|
apartment tenement boardinghouse mansion factory laundry bakery kitchen table dinner lunch breakfast
|
|
party wedding ball club casino studio gallery magazine newspaper bank
|
|
""".split())
|
|
|
|
# Words for a person with no role at all. "A man" is not a topic; the first baseline had a group
|
|
# scored 90% topic-bound for saying "man", so these are left out of the measure (still offered as
|
|
# role hints on the review page). Fixed before any predicament was generated.
|
|
GENERIC_PEOPLE = {'man', 'woman', 'girl', 'boy', 'child', 'worker', 'stranger', 'friend'}
|
|
|
|
TOPIC_WORDS = (ROLE_WORDS - GENERIC_PEOPLE) | PLACE_WORDS
|