The pipeline's embed-and-cluster step is dead, and this commit holds both the evidence for that and the step proposed to replace it. Predicaments. Scenes are re-described as "what the person is up against", with no names, jobs or places, then embedded and clustered (redescribe.py, topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading both side by side. Two defects the pilot exposed are fixed: split.py missed titles in quotes and a contents subtitle after a dash, so three stories had been merged into their neighbours, and strip_names.py read New York place names as people. The corrected corpus is probe/v2 (97 stories, 839 scenes); carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from 20% to 13% at k=60, short of the pre-registered 10%. Hand references. Three corpora were read scene by scene and written up by hand, under the same prompt rules the local models get, as a baseline to judge them against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a readable page. No inference was used for any of them. Catalogue. probe/catalogue maps every hand group in the three references onto 36 situation entries, with an answer key per corpus and one recurrence rule applied to all three. classify.py assigns a scene one entry or none, leave-one-corpus- out; score.py checks it against the key, with a self-test on random labels. Why clustering is out: hand-written predicaments, embedded and clustered exactly as the model's were, agree with the hand grouping at ARI 0.05 — no better than the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even shortlist: the hand label is the nearest entry 13% of the time and in the top 8 half the time. The classification runs are not here. The dev and test runs are pre-registered in probe/catalogue/README.md with the bar set beforehand, and are blocked on the inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene partial output in out/ is not a result. Review page. The situation review is now a browser page rather than JSON edited by hand (review_page.py, review_page_logic.cjs with Node tests, format schema v2). It has never been rendered in a real browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
71 lines
3.7 KiB
Python
71 lines
3.7 KiB
Python
"""Re-describe each scene as the predicament its main person is in, with nothing tying it to the story.
|
|
|
|
The first human review found the summary groups were topics (artists, police, courtship, hotels)
|
|
because summaries keep jobs and places, and those dominate the embeddings. The unit a pack needs is
|
|
a predicament: what someone is up against and what they want out of it (SELECTION.md, "Predicaments").
|
|
This rewrites the scene's own prose into that, leaving out whatever makes a topic.
|
|
|
|
STP_OLLAMA=http://host:11434 STP_CHAT_MODEL=qwen3:14b python3 redescribe.py --limit 40 --out pilot-predicaments-14b.json
|
|
STP_OLLAMA=http://host:11434 python3 redescribe.py # every scene -> predicaments.json
|
|
|
|
The output has the same layout as summaries.json ({"summaries": [{"chunk", "story", "summary"}]}),
|
|
so strip_names.py and embed.py read it unchanged. One request at a time; progress is saved after
|
|
every scene, so an interrupted run resumes.
|
|
"""
|
|
import argparse, json, os, pathlib, re, time
|
|
import ollama
|
|
from topic_words import TOPIC_WORDS
|
|
|
|
ap = argparse.ArgumentParser()
|
|
ap.add_argument('--limit', type=int, default=None)
|
|
ap.add_argument('--out', default='predicaments.json')
|
|
args = ap.parse_args()
|
|
MODEL = os.environ.get('STP_CHAT_MODEL', 'qwen2.5:3b-instruct')
|
|
|
|
# Fixed before any call (SELECTION.md). No example of the output vocabulary: the 3B model copied
|
|
# the examples it was given when summarising.
|
|
SYSTEM = """You describe the predicament of the main person in a passage of fiction, in ONE sentence.
|
|
|
|
A predicament is what that person is up against and what they want out of it: the kind of trouble or opportunity that could happen to anyone, in any line of work, in any town.
|
|
|
|
Rules:
|
|
- Exactly one sentence, at most 20 words.
|
|
- Leave out everything that ties it to this story: no names, no jobs or professions, no places, no businesses, and no objects that only belong to this story.
|
|
- Say what the person is up against, and what they want.
|
|
- Do not describe the writing, the mood, or who is telling the story.
|
|
- Output only the sentence."""
|
|
|
|
# A full stop after these is not the end of a sentence; nor is one after a single capital (an initial).
|
|
END = re.compile(r'(?<!\bMr)(?<!\bMrs)(?<!\bMs)(?<!\bDr)(?<!\bSt)(?<!\bNo)(?<![\s.][A-Z])[.!?](?:\s|$)')
|
|
|
|
|
|
def first_sentence(text):
|
|
text = re.sub(r'<think>.*?</think>', '', text, flags=re.S)
|
|
text = re.sub(r'\s+', ' ', text).strip().strip('"')
|
|
m = END.search(text)
|
|
return text[:m.end()].strip() if m else text
|
|
|
|
|
|
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
|
|
if args.limit:
|
|
chunks = chunks[:args.limit]
|
|
out = pathlib.Path(args.out)
|
|
state = json.loads(out.read_text()) if out.exists() else {'model': MODEL, 'prompt': SYSTEM, 'summaries': []}
|
|
if state['model'] != MODEL or state['prompt'] != SYSTEM:
|
|
raise SystemExit(f'{out} was made with a different model or prompt; choose another --out')
|
|
done = state['summaries']
|
|
print(f'{len(chunks)} scenes, resuming at {len(done)}, model {MODEL}', flush=True)
|
|
|
|
t0 = time.time()
|
|
for i in range(len(done), len(chunks)):
|
|
think = False if MODEL.startswith('qwen3') else None
|
|
text, _, _ = ollama.chat(MODEL, SYSTEM, chunks[i]['text'], num_ctx=2048, num_predict=80, think=think)
|
|
done.append({'chunk': i, 'story': chunks[i]['story'], 'summary': first_sentence(text), 'model_output': text})
|
|
out.write_text(json.dumps(state, indent=1, ensure_ascii=False), encoding='utf-8')
|
|
if (i + 1) % 10 == 0 or i + 1 == len(chunks):
|
|
print(f' {i + 1}/{len(chunks)} {time.time() - t0:.0f}s', flush=True)
|
|
|
|
leaks = [x for x in done if set(re.findall(r"[a-zé]+", x['summary'].lower())) & TOPIC_WORDS]
|
|
print(f'done: {len(done)} predicaments, {time.time() - t0:.0f}s this session; '
|
|
f'{len(leaks)} still mention a job, relationship or place word')
|