Fourteen corpora split and classified: aesop, bierce, chekhov, holmes,
keefe, lawson, lorimer, maupassant, nobody, plaintales, poe, torchy,
wallingford and winesburg, each with its splitter and the hand-written
groups and pages; catalogue v2 and v3; and the shared splitters
gutenberg_chunks.py, se_split.py and se_build.py.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rwKTmug58sEsJ72AuWsEi
The pipeline's embed-and-cluster step is dead, and this commit holds both the
evidence for that and the step proposed to replace it.
Predicaments. Scenes are re-described as "what the person is up against", with
no names, jobs or places, then embedded and clustered (redescribe.py,
topic_words.py, topic_share.py). The pilot chose qwen3:14b over 3b by reading
both side by side. Two defects the pilot exposed are fixed: split.py missed
titles in quotes and a contents subtitle after a dash, so three stories had been
merged into their neighbours, and strip_names.py read New York place names as
people. The corrected corpus is probe/v2 (97 stories, 839 scenes);
carry_summaries.py reuses the 829 unchanged v1 summaries. Topic share fell from
20% to 13% at k=60, short of the pre-registered 10%.
Hand references. Three corpora were read scene by scene and written up by hand,
under the same prompt rules the local models get, as a baseline to judge them
against: O. Henry (probe/v2/claude, 839 scenes, 20 situations), Wharton's
Descent of Man (probe/wharton, 262 scenes, 16 groups) and Jacobs's The Lady of
the Barge (probe/jacobs, 157 scenes, 19 groups). Each has its own README and a
readable page. No inference was used for any of them.
Catalogue. probe/catalogue maps every hand group in the three references onto 36
situation entries, with an answer key per corpus and one recurrence rule applied
to all three. classify.py assigns a scene one entry or none, leave-one-corpus-
out; score.py checks it against the key, with a self-test on random labels.
Why clustering is out: hand-written predicaments, embedded and clustered exactly
as the model's were, agree with the hand grouping at ARI 0.05 — no better than
the 14B text's 0.07. Better rewriting cannot rescue it. Embeddings cannot even
shortlist: the hand label is the nearest entry 13% of the time and in the top 8
half the time.
The classification runs are not here. The dev and test runs are pre-registered
in probe/catalogue/README.md with the bar set beforehand, and are blocked on the
inference host, whose GPU has fallen off the PCIe bus three times. The 30-scene
partial output in out/ is not a result.
Review page. The situation review is now a browser page rather than JSON edited
by hand (review_page.py, review_page_logic.cjs with Node tests, format schema
v2). It has never been rendered in a real browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014BygvsUXV9eU6oHkTCkKZ1
Research toward building a content pack from a story corpus, kept on its own
branch and independent of the game. Records the selection experiments against
blind labels, and settles selection as gate G2 followed by a human review:
review.py writes REVIEW.md and a review.json form, apply_review.py checks the
filled form and writes situations.json for the next stage.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C6UDQ9o6L6Ey173U7XVou6
Cold Fork is a freight yard: 29 events across two stages, a four-day cycle, and
its own relationship and skill vocabulary. src/engine/ was not touched. Adding it
needed three changes outside the engine: one save slot per pack, app.js holding
the current pack, and turn.js no longer naming the mailroom (DECISIONS #34).
build.js now handles imports wrapped across lines. A conformance suite runs over
every registered pack. DECISIONS #35 records three calls on open questions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RxdJFbLq1rBKsUeticGV1g
The stamping scheme was wrong in a way worth recording. tools/stamp.js rewrote
src/build-info.js before every serve and build, so after each commit the
committed stamp named the *previous* commit — as it does right now, reading
76072ed+ while HEAD is 5c1784d — and the next serve rewrote it and dirtied the
tree again. A generated value does not belong in a tracked file if anything
routinely regenerates it.
src/build-info.js is now permanent and reads `commit: 'dev'`. tools/build.js
substitutes the real commit into the bundled output only, and fails loudly if
the substitution finds nothing to replace. So dist/theladder.html names the
commit that produced it, running from source honestly reads "0.3.0 · dev", and
the working tree never churns. tools/stamp.js is gone and `npm run serve` is a
plain static server again.
Also here, for picking this up later: a "Where things stand" section in the
README with the five open questions in the order they are likely to matter —
whether money stops mattering late, the negotiation gate that four early options
sit behind, the feedback widget nobody uses, how thin Dispatch is next to the
mailroom, and third-tier versus second-pack.
121 tests. docs/DECISIONS.md §25 corrected to describe what the code now does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VMSFHyVPitUoosW5wyEADj
A 182-turn session spent 30% of its turns below zero, went negative on turn 18
and never recovered, and chose the same hard-times option — "take every shift
going" — thirty-seven times.
The cause was a tuning change made without re-deriving what it meant. Rent went
from 385 to 455 while chasing a different problem, which took the mailroom
baseline from the designed -$3/day to -$13/day. Worse, at the debt that run
accumulated the weekly service charge came to $180 against an escape option
worth $120. The hole was inescapable by arithmetic, whatever the player did.
Nothing caught it. Every archetype either optimised money or got promoted out of
the problem before it bit, and the dominance test passed because rent_bounced
was only 17% of turns.
Fixes: mailroom wages 980 -> 1120, restoring the -$3/day baseline; the escape
option 120 -> 220, so it is worth more than a week's rent; rent_bounced from
weight 5000 with no cooldown to 900 on a cooldown of 3, because being broke
should colour a run rather than replace it; debt_collector from weight 25 to 45,
since at 25 it appeared six times in 262 turns and debt became a ratchet.
Three guards so this class of bug cannot recur quietly.
The baseline is now asserted directly: a test sums the pack's own upkeep over
twenty fortnights and requires the mailroom to net between -90 and +10, and
Dispatch to be better but not so much better that money stops mattering. That
would have failed the moment rent changed. Outcome tests over simulated play are
a slow and noisy way to detect a number that is simply wrong.
A `lifer` archetype refuses any option that would change stage — generically, by
looking for an effect on `stage` — and otherwise plays for people. It reproduces
the session that found this; no other archetype can.
And no archetype may spend more than a quarter of a run below zero. Hard times
is a state a player passes through; living in it is the failure mode.
Also fixed: the analyser reported "every ~-5 turns". One export can hold several
playthroughs and turn numbers restart with each, so spans are now accumulated
per run and pooled rather than measured across the seam between two.
121 tests. Reasoning in docs/DECISIONS.md §29-30.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VMSFHyVPitUoosW5wyEADj
Five things from the second playtest.
A status line at the top of every screen: the game, the pack, your current
title, and the version plus the commit it came from. `tools/stamp.js` writes the
commit into src/build-info.js before serving and before building, so a
distributed dist/theladder.html states exactly which commit produced it. An
uncommitted working tree gets a `+` suffix — "4cab9fc+" is honest in a way a
bare SHA would not be.
The whole game is now playable on Enter. Every screen marks one control and the
app focuses it after each render: the first *available* option while choosing,
the next-turn button once the result is in, "keep going" on the retrospective.
Tab and Shift+Tab reach the other options. Locked options needed no special
handling — the browser's tab order already skips disabled controls, so they stay
visible without being in the way.
"What is notable? under What changed?" — two fixes. `notable` is the
retrospective's own list of moments; a display rule can now say
`inLedger: false` and it is left out, because "Notable — changed" tells nobody
anything. Flags are the opposite and stay, but as sentences the pack supplies
rather than as booleans: "The envelope is in your locker." instead of "the
envelope: now true".
Exported logs live in logs/, gitignored, and `npm run analyze` reads the newest
one there when given no argument. A log records what a real person did turn by
turn, including whatever they typed into the feedback box — committing one once
would put it in the history for good.
Version set to 0.3.0 on the reasoning that the engine was 0.1 and the interface
0.2; say if you want a different scheme and I will renumber.
119 tests. Reasoning in docs/DECISIONS.md §25-28.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VMSFHyVPitUoosW5wyEADj
The mailroom is no longer the whole game. Around turn 50, a player who has
built standing and either Marlene's goodwill or the nerve to ask gets offered
the Dispatch post — and can take it, take it and put Trevor up for the mailroom,
or turn it down, which is a real strategy rather than a mistake and comes back
after a cooldown.
Dispatch is eleven events of its own. The money problem is largely solved up
there and replaced by other people's days depending on yours, including the
mailroom's — which you used to be. Trevor branches on whether you recommended
him: he either runs the mailroom and has opinions about how it used to be run,
or he stays put and starts addressing you by your full title, kindly, in front
of people.
None of this is a promotion mechanic in the engine. It is
`{ path: 'stage', op: 'set', value: 'dispatch' }` on an option, since `stage`
was already how events are filtered. The reserved-write rule now allows it
narrowly — `set` only, and only to a stage that has events, both checked by the
validator, so a typo cannot strand the player in an empty pool. Wages differ by
rank through two upkeep entries gated on `stage`. The hard-times events dropped
their stage entirely: rent is due wherever you work.
The second half of this is the thing the last playtest asked for. A 183-turn
session, and no way to see how much of it was new ground.
`npm run analyze <exported-log.json>` now reports turns against distinct
situations, how often each recurred, which options were most used, which locked
gates players kept meeting, and anything they typed. The retrospective shows the
player-facing version.
It found a regression in this very change on its first run. Dispatch started
with eight events against the mailroom's nineteen, so its two floor events were
46% of every promoted run — one every three turns. Promotion was moving the
player into *thinner* content. Three more events and a weight rebalance took the
top two to 24.5%, and the heaviest recurrence from every ~3 turns to every ~5.
A test now holds that line.
Also corrected: the coverage test was asking whether an archetype ever *took* an
option, which is a property of the bot, not the pack. It had flagged
`rent_bounced.ask_trevor` as dead content when it had been offered, unlocked,
419 times across a sweep and simply never chosen. Options are now checked on
availability; events are still checked on firing.
114 tests. Reasoning in docs/DECISIONS.md §22-24.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VMSFHyVPitUoosW5wyEADj
The MVP is playable: a turn screen, a career retrospective, an optional
feedback widget, local-storage autosave with file export/import, and a build
that flattens everything into one self-contained HTML file.
A turn has two phases — the scenario and its options, then what the choice did
and, separately, what the day cost regardless. Numbers that move without the
player seeing why are most of what makes them meaningless. Locked options are
shown greyed with the requirement spelled out rather than hidden, so a player
can see the door they cannot open yet.
The interface is named entirely by the pack. `pack.display` maps a state path to
a label, a format and an order, which is why the header reads "Standing" for a
path called `reputation` and why debt disappears while it is zero.
The first playtest then found three things, all of them fair.
*"Wasn't clear what daily drain was for."* The ledger was labelling changes by
the path that moved — "Money −$18" — when the upkeep entry that caused them had
a perfectly good name. It now says "Coffee, transit, lunch". The labels were in
the data the whole time and never reached the screen.
*"Should probably pay rent weekly / get paid every other week… would be good to
have day of week and week # shown."* Packs can now declare a calendar, and the
engine derives the day name, cycle and index from the turn number before the
event is drawn — so content can require a Friday and the header can say
"Thursday · Week 2 · Day 12". Upkeep entries take `every` and `offset`, making
wages fall on alternate Fridays and rent every Monday. The turn screen carries
a diary line — "Wages tomorrow · Rent in 4 days" — computed generically from
whatever a pack schedules.
*"Always in debt — could never get ahead."* The old economy bled $25 a day
regardless of play. It is now roughly break-even at baseline, with a
`spare_shift` event that appears *because* you are broke: the way out of a hole
should be visible from inside it. Across archetypes a careful player ends around
$1,400, an unplanned one treads water, and a careless one sinks into real debt.
That retune broke the coverage test, usefully. Random play stopped reaching hard
times at all, which made the entire debt branch look dead — it was not, since
random play is not a plausible player. Coverage is now the union across five
archetypes, including one that is broke *and* well-liked, because the "ask a
friend for money" options sit in a state no single-axis strategy reaches.
Not covered by any of this: styling and dark mode, which only a human with a
browser can check. This machine has none that can render a page.
109 tests. Reasoning in docs/DECISIONS.md §14-21.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VMSFHyVPitUoosW5wyEADj