Files
TheLadder/tools/story-to-pack/probe/redescribe.py
T
JesseMarkowitzandClaude Opus 5 e5617b86ba Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.

Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.

Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.

Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.

Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.

The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.

Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
2026-09-20 17:19:42 -04:00

71 lines
3.7 KiB
Python

"""Re-describe each scene as the predicament its main person is in, with nothing tying it to the story.
The first human review found the summary groups were topics (artists, police, courtship, hotels)
because summaries keep jobs and places, and those dominate the embeddings. The unit a pack needs is
a predicament: what someone is up against and what they want out of it (SELECTION.md, "Predicaments").
This rewrites the scene's own prose into that, leaving out whatever makes a topic.
STP_OLLAMA=http://host:11434 STP_CHAT_MODEL=qwen3:14b python3 redescribe.py --limit 40 --out pilot-predicaments-14b.json
STP_OLLAMA=http://host:11434 python3 redescribe.py # every scene -> predicaments.json
The output has the same layout as summaries.json ({"summaries": [{"chunk", "story", "summary"}]}),
so strip_names.py and embed.py read it unchanged. One request at a time; progress is saved after
every scene, so an interrupted run resumes.
"""
import argparse, json, os, pathlib, re, time
import ollama
from topic_words import TOPIC_WORDS
ap = argparse.ArgumentParser()
ap.add_argument('--limit', type=int, default=None)
ap.add_argument('--out', default='predicaments.json')
args = ap.parse_args()
MODEL = os.environ.get('STP_CHAT_MODEL', 'qwen2.5:3b-instruct')
# Fixed before any call (SELECTION.md). No example of the output vocabulary: the 3B model copied
# the examples it was given when summarising.
SYSTEM = """You describe the predicament of the main person in a passage of fiction, in ONE sentence.
A predicament is what that person is up against and what they want out of it: the kind of trouble or opportunity that could happen to anyone, in any line of work, in any town.
Rules:
- Exactly one sentence, at most 20 words.
- Leave out everything that ties it to this story: no names, no jobs or professions, no places, no businesses, and no objects that only belong to this story.
- Say what the person is up against, and what they want.
- Do not describe the writing, the mood, or who is telling the story.
- Output only the sentence."""
# A full stop after these is not the end of a sentence; nor is one after a single capital (an initial).
END = re.compile(r'(?<!\bMr)(?<!\bMrs)(?<!\bMs)(?<!\bDr)(?<!\bSt)(?<!\bNo)(?<![\s.][A-Z])[.!?](?:\s|$)')
def first_sentence(text):
text = re.sub(r'<think>.*?</think>', '', text, flags=re.S)
text = re.sub(r'\s+', ' ', text).strip().strip('"')
m = END.search(text)
return text[:m.end()].strip() if m else text
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
if args.limit:
chunks = chunks[:args.limit]
out = pathlib.Path(args.out)
state = json.loads(out.read_text()) if out.exists() else {'model': MODEL, 'prompt': SYSTEM, 'summaries': []}
if state['model'] != MODEL or state['prompt'] != SYSTEM:
raise SystemExit(f'{out} was made with a different model or prompt; choose another --out')
done = state['summaries']
print(f'{len(chunks)} scenes, resuming at {len(done)}, model {MODEL}', flush=True)
t0 = time.time()
for i in range(len(done), len(chunks)):
think = False if MODEL.startswith('qwen3') else None
text, _, _ = ollama.chat(MODEL, SYSTEM, chunks[i]['text'], num_ctx=2048, num_predict=80, think=think)
done.append({'chunk': i, 'story': chunks[i]['story'], 'summary': first_sentence(text), 'model_output': text})
out.write_text(json.dumps(state, indent=1, ensure_ascii=False), encoding='utf-8')
if (i + 1) % 10 == 0 or i + 1 == len(chunks):
print(f' {i + 1}/{len(chunks)} {time.time() - t0:.0f}s', flush=True)
leaks = [x for x in done if set(re.findall(r"[a-zé]+", x['summary'].lower())) & TOPIC_WORDS]
print(f'done: {len(done)} predicaments, {time.time() - t0:.0f}s this session; '
f'{len(leaks)} still mention a job, relationship or place word')