Files
TheLadder/tools/story-to-pack/probe/pilot_page.py
T
JesseMarkowitzandClaude Opus 5 e5617b86ba Replace clustering with a catalogue, and hand-write the references to judge it against
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.

Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.

Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.

Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.

Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.

The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.

Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
2026-09-20 17:19:42 -04:00

116 lines
6.2 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Side-by-side page for reading the predicament pilot: the scene, its summary, and what each model made of it.
python3 strip_names.py pilot-predicaments-3b.json pilot-3b-clean.json
python3 strip_names.py pilot-predicaments-14b.json pilot-14b-clean.json
python3 pilot_page.py pilot-3b-clean.json pilot-14b-clean.json > pilot-compare.html
Takes strip_names.py's output, so the page shows what the full run would actually embed. (strip_names
runs as a script on import, so it is not imported here.) Job, relationship and place words are
highlighted. A pick per scene is kept in the browser only; the tally at the top is what to report
back. No inference.
"""
import collections, html, json, pathlib, re, sys
from topic_words import TOPIC_WORDS
def clean(text):
return text # inputs are already name-stripped
a_path, b_path = sys.argv[1], sys.argv[2]
cols = [json.loads(pathlib.Path(p).read_text(encoding='utf-8')) for p in (a_path, b_path)]
labels = [c['model'] for c in cols]
chunks = json.loads(pathlib.Path('chunks.json').read_text(encoding='utf-8'))
summary = {x['chunk']: x['summary'] for x in
json.loads(pathlib.Path('summaries_clean.json').read_text(encoding='utf-8'))['summaries']}
def marked(text):
out = []
for part in re.split(r"([A-Za-zé]+)", text):
esc = html.escape(part)
out.append(f'<mark>{esc}</mark>' if part.lower() in TOPIC_WORDS else esc)
return ''.join(out)
def leaks(text):
return len(set(re.findall(r"[a-zé]+", text.lower())) & TOPIC_WORDS)
def opening(text):
return ' '.join(text.split()[:3])
rows, stats = [], []
for col in cols:
texts = [clean(x['summary']) for x in col['summaries']]
stats.append({
'leaky': sum(leaks(t) > 0 for t in texts),
'city': sum('city' in t.lower() for t in texts),
'words': sum(len(t.split()) for t in texts) / len(texts),
'openings': collections.Counter(opening(t) for t in texts).most_common(3),
})
n = min(len(c['summaries']) for c in cols)
for i in range(n):
x = cols[0]['summaries'][i]
ch = chunks[x['chunk']]
cells = ''.join(
f'<div class="pred"><span class="who">{html.escape(labels[j])}</span>'
f'<p>{marked(clean(cols[j]["summaries"][i]["summary"]))}</p>'
f'<label><input type="radio" name="p{i}" value="{j}"> better</label></div>'
for j in range(2))
rows.append(f'''<section data-i="{i}">
<header><span class="n">{i + 1}</span> <b>{html.escape(ch["title"])}</b></header>
<details><summary>Scene summary: {marked(summary.get(x["chunk"], ""))}</summary>
<p class="scene">{html.escape(ch["text"][:1200])}{"…" if len(ch["text"]) > 1200 else ""}</p></details>
<div class="pair">{cells}</div>
<label class="both"><input type="radio" name="p{i}" value="n"> neither is a predicament</label>
</section>''')
stat_rows = ''.join(
f'<tr><td>{html.escape(labels[j])}</td><td>{s["leaky"]} of {n}</td><td>{s["city"]}</td>'
f'<td>{s["words"]:.0f}</td><td>{"; ".join(f"“{html.escape(o)}…” ×{c}" for o, c in s["openings"])}</td></tr>'
for j, s in enumerate(stats))
print(f'''<!doctype html><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
<title>Predicament pilot</title>
<style>
:root{{--bg:#f7f6f2;--card:#fff;--ink:#222;--dim:#666;--line:#ddd;--mark:#ffe08a;--pick:#e3f1e6}}
@media (prefers-color-scheme:dark){{:root{{--bg:#1b1b1d;--card:#252528;--ink:#e6e6e6;--dim:#999;--line:#3a3a3e;--mark:#6b5415;--pick:#23402b}}}}
body{{background:var(--bg);color:var(--ink);font:15px/1.5 system-ui,sans-serif;margin:0;padding:24px 16px}}
main{{max-width:980px;margin:auto}} h1{{margin:0 0 4px}} .dim{{color:var(--dim)}}
table{{border-collapse:collapse;margin:12px 0;font-size:14px}} td,th{{border-bottom:1px solid var(--line);padding:4px 10px;text-align:left}}
.wrap{{overflow-x:auto}}
#tally{{position:sticky;top:0;background:var(--bg);padding:8px 0;border-bottom:1px solid var(--line);z-index:1}}
section{{background:var(--card);border:1px solid var(--line);border-radius:8px;padding:12px 14px;margin:14px 0}}
header .n{{color:var(--dim);margin-right:4px}} summary{{cursor:pointer;color:var(--dim);margin:6px 0}}
.scene{{white-space:pre-wrap;font-size:13px;color:var(--dim)}}
.pair{{display:grid;grid-template-columns:1fr 1fr;gap:10px}} @media (max-width:640px){{.pair{{grid-template-columns:1fr}}}}
.pred{{border:1px solid var(--line);border-radius:6px;padding:8px 10px}} .pred p{{margin:4px 0 8px}}
.pred:has(input:checked){{background:var(--pick)}} .who{{font-size:12px;color:var(--dim)}}
mark{{background:var(--mark);color:inherit;border-radius:2px}} .both{{display:block;margin-top:8px;font-size:13px;color:var(--dim)}}
</style>
<main>
<h1>Predicament pilot</h1>
<p class="dim">The first {n} scenes, each re-described by two models as the main person's predicament. Names are
already removed as the full run would. <mark>Highlighted</mark> words are jobs, relationships and places, the
words the old groups formed around. Pick the better sentence in each scene, or "neither". The column order
is fixed: {html.escape(labels[0])} on the left.</p>
<p class="dim">What to look for: does the sentence say what the person is up against and what they want? Could
it happen to anyone, anywhere? Would two scenes with the same predicament end up with similar sentences?</p>
<div class="wrap"><table><tr><th>Model</th><th>Has a highlighted word</th><th>Says "city"</th><th>Words, mean</th><th>Most common openings</th></tr>{stat_rows}</table></div>
<div id="tally"></div>
{"".join(rows)}
</main>
<script>
const KEY='stp-pilot:'+{json.dumps(labels)}.join('|'), labels={json.dumps(labels)};
let saved={{}}; try{{saved=JSON.parse(localStorage.getItem(KEY)||'{{}}')}}catch(e){{}}
for(const [k,v] of Object.entries(saved)){{const el=document.querySelector(`input[name="${{k}}"][value="${{v}}"]`); if(el) el.checked=true}}
function tally(){{
const c={{0:0,1:0,n:0}}; let done=0;
document.querySelectorAll('input[type=radio]:checked').forEach(r=>{{c[r.value]++;done++}});
document.getElementById('tally').textContent=`Picked ${{done}} of {n}: ${{labels[0]}} ${{c[0]}}, ${{labels[1]}} ${{c[1]}}, neither ${{c.n}}`;
}}
document.addEventListener('change',e=>{{if(e.target.type!=='radio')return; saved[e.target.name]=e.target.value;
try{{localStorage.setItem(KEY,JSON.stringify(saved))}}catch(err){{}} tally()}});
tally();
</script>''')