v0.8.0.9 — the bot plans its switching turn, stops wasting its draws, and the engine walks each route once

The developer bot, re-measured decision by decision against the bot before it,
goes from about -0.3 revenue a game to about 4.8:

- plans the whole switching turn before its first Move (sim/switch-planner.ts),
  +2.89 over 1600 paired seeds; closes TODO #53
- takes a face-up card only if it could play it, +1.52 over 1600 seeds
- stops running Second Sections by accident in the New Train phase, +0.32
- lays track by what the district can do afterwards, +0.12 over 6400 seeds,
  run-arounds in 22 of 60 districts against 9

The engine is 2.8x faster with play proven identical: a route cache scoped to
one unchanged position, applyIntent split into prepareIntent + commitEvents,
and less allocation in exploreMoves. npm test now leaves out the bot
simulations, which run as npm run test:sim.

No rule changed; games in progress resume. Rejected candidates and the
Second Section card question are in CHANGELOG.md and TODO.md (#104-#106).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017nnuCv8UodHucFfx3LWEoX
This commit is contained in:
Jesse.Markowitz
2026-09-15 15:30:42 -04:00
co-authored by Claude Opus 5
parent 072029b1f7
commit 76c6e103b3
13 changed files with 1550 additions and 112 deletions
+271 -10
View File
@@ -81,7 +81,7 @@ Not items. Things that are true of every change, and that have gone wrong when s
5. **Replays and saved games** — #14 #47 #48 #49 #50 #51 #52
6. **Rules** — #12 #80 #82 #83 #85
7. **Play balance** — #61 #62 #63 #64 #67 #68 #69 #70 #71 #72 #73 #66 #65 #74
8. **The bot** — #41 #57 #59 #53 #54 #58 #55 #56 #60
8. **The bot** — #104 #105 #106 #41 #57 #59 #54 #58 #55 #56 #60
9. **Code health and housekeeping** — #46 #45 #84 #87
10. **Documentation and assets** — #15a #86 #88
@@ -100,20 +100,119 @@ Everything else in this file waits behind a release; this waits behind an aftern
**Test runs WERE made across 0.7.4 through 0.7.9** (Jesse, 2026-09-07) and produced no change
requests — the two bugs that did come out of them are Gitea#21 and #22, fixed in v0.7.9.1. So this
section is not "nobody has touched it since 0.7.4"; it is the narrower and still-true claim that the
specific paths below have not been exercised at a table. **More testing is planned at the end of the
0.7.9 series, before 0.8.0 starts** — that is the moment to close these, not a separate errand.
specific paths below have not been exercised at a table.
**The gate moved.** It was "before 0.8.0 starts"; 0.8.0 shipped anyway, through v0.8.0.8, so the
session now runs against that build and covers what it added as well. See **Preparing the session**
below — written 2026-09-10 because the measurement it rests on is the whole point: **three of the
four things this section is named for do not happen by themselves.**
### Preparing the session
**MEASURED, 2026-09-10, across ten full competitive games driven to completion.** What a table will
meet without trying, and what it will not:
| interruption | fires in | so |
| --- | --- | --- |
| Superintendent clearance (§8.1) | **9/10 games** | you will meet it; just play |
| a train held at the Limits | 7/10 | ditto |
| Extras started and queued | 10/10 | ditto |
| collisions | 7/10 | ditto |
| Red Flags set / spent | 7/10, 6/10 | ditto |
| **the Yard Office offer** | **0/10** | must be set up |
| **the Red Flag hold and its prompt** | **0/10** | must be set up |
| **extended play (`dayExtended`)** | **0/10** | must be set up |
Those last three are exactly what #39 and #35 are NAMED for. They are not broken — they are
conditional, and the conditions are these, read out of `advance.ts` rather than guessed:
- **Yard Office** (`advance.ts` ~1290) needs the destination district to contain a card carrying the
`yardOffice` **enhancement**, AND an arriving train with **no coach** in its consist, AND a usable
route. The bot never builds one, so **somebody has to build a Yard Office and then let a freight
train arrive.**
- **Red Flag hold** (`advance.ts` ~1232) needs the destination player to be **holding the Red Flags
maneuver card**, AND an arrival that would genuinely collide — §8.3's own two ways: no free A/D
track, or cars fouling the Running Track. So: **hold that card and let your A/D tracks fill.**
- **Extended play** needs the timetable to RUN OUT, which a five-Day game does not do. Deal it with
**`days: 1`** — that is exactly what the 2026-08-29 API verification did, and why it got there.
**What the session needs**
- **Two people, two browsers, two devices.** #35's remaining gap is specifically what a SECOND
player sees while waiting on a first, and whether "waiting on Carol" still reads once Carol has
closed her laptop. That cannot be tested alone, and it is the half that has never been done.
- **Two games, not one.** A short `days: 1` game to reach the extension vote, and an ordinary game
for everything else — with somebody deliberately building a Yard Office and holding Red Flags.
- **#42a is separate and takes five minutes**, solitaire, one person: click every field on the setup
screen and confirm the dealt game matches what was chosen.
**The caution this section exists because of.** #35's own Reference entry records that the
2026-08-29 verification passed over the HTTP API — **which renders no dialog** — and that is exactly
why the v0.7.9 bug survived: the vote sat underneath a modal results dialog whose only control was
Close. What was proven was that the SERVER supports extended play, not that a player can reach it.
Read that into every "verified on `phoenix.local`" line in this file, and into everything v0.8.0
added, all of which is verified by test and simulation and none of it by eye.
**What v0.8.0 added to this list**, none of it played by a person for a whole game and none with a
second human: the watchable board and its ordered steps, the speed control, the pile highlighting and
the Home Office deck tile, "Your Move" being put away while catching up, the Day-end collision line,
and the Salvage Yard naming its top card.
### The checklist
Grouped by what has to be set up, with the item each observation closes. Nothing here needs a
developer present; what it needs is somebody writing down what they saw.
**Game A — `days: 1`, two humans, two browsers.** Reaches the extension vote in one Day.
- [ ] The vote appears **in front of both players**, not underneath the results dialog (#35 — this is
the exact shape of the bug v0.7.9 fixed).
- [ ] While one player has not voted, the other's turn chart says **who** it is waiting on (#35).
- [ ] **Close the second laptop mid-vote.** Does the first player learn why nothing is happening, and
does "waiting on Carol" still read once Carol is gone? (#35 — never tested.)
- [ ] Reopen it. The history panel comes back **populated**, not empty, and the board is current
(the v0.7.9.5 reconnect fix, never seen by a person).
- [ ] Vote yes. The extra Day begins and the official result is **unchanged** from when the
timetable ran out (#35).
**Game B — ordinary length, two humans, bots to fill.** Everything else.
- [ ] Somebody **builds a Yard Office** and lets a freight train (no coach) arrive at it. The offer
interrupts the Mainline Phase and asks a question mid-thought — is it legible, and does it say
which train? (#39)
- [ ] Somebody **holds the Red Flags card** while their A/D tracks are full, so an arrival would
collide. The hold is offered out of phase (#39).
- [ ] A **loaded Extra** is made up and run (#39 — the third of its three).
- [ ] Watch a bot take a whole turn: does the district follow it, does the lit pile catch the eye,
does the caption say who and what? (v0.8.0)
- [ ] Find the speed that suits you and say what it is — it becomes the committed default.
- [ ] Let the board fall behind, then press **Skip**. Nothing is lost; the history has it all.
- [ ] End a Day with a collision on it: the summary reads "N on Day D, N in all" and cannot
contradict itself (v0.8.0.2).
**Solitaire, five minutes, alone.**
- [ ] Click through **every field** on the setup screen and confirm the dealt game matches what was
chosen (#42a).
**Whatever else happens.** The two bugs that came out of the 0.7.4-0.7.9 runs were both things
nobody set out to test. Write down anything that reads wrong, even where the rule underneath is
right — most of this release's defects were legible-but-wrong rather than broken.
- [ ] **#39** — **None of v0.7.4 has been played by a human.** The Yard Office offer, the Red Flag
hold and its out-of-phase prompt, and the loaded-Extra make-up rules are tested end to end,
packed, and running on `phoenix.local` — and nobody has met any of them at a board. **Two are
interruptions that stop the Mainline Phase and put a question in front of somebody
mid-thought**, which is exactly the kind of thing only play reveals. See **Reference · #39**.
mid-thought**, which is exactly the kind of thing only play reveals. **Neither of those two
happens by itself — 0/10 games. See Preparing the session above for what to set up.** See
**Reference · #39**.
- [ ] **#35** — **Extended play has never been played at a real table.** It was verified over the HTTP
API, which renders no dialog — and when a human first reached it in a browser it was unusable
(fixed in v0.7.9). The multiplayer vote has still never been driven through two browsers: what a
second player sees while waiting, and whether "waiting on Carol" reads once Carol has closed her
laptop, are unanswered. See **Reference · #35**.
laptop, are unanswered. **Needs a `days: 1` game — the timetable does not run out in five
Days, so extended play fired in 0/10 measured games.** See **Reference · #35**.
- [ ] **#42a** — **Nobody has clicked through the solitaire setup screen's own fields** and confirmed
the dealt game matches what was chosen. It took three attempts to become reachable at all —
@@ -368,21 +467,37 @@ v0.7.9's collision-floor change (#61).
The developer bot exists to measure the game, not to be a good opponent — so a bot weakness matters
when it stops a measurement being trustworthy. **Read #57 before tuning any weights.**
**Since 2026-09-14 the bot plans its whole switching turn** (`sim/switch-planner.ts`, +2.89 revenue a
game), **takes a face-up card only if it could play it** (+1.52), and **since 2026-09-15 lays track by
what the district can do afterwards** (`bestValuedLay`, +0.12 over 6400 seeds, run-arounds 9/60 → 22/60). Jesse's goal for it is better decisions in simulated runs AND at a real table, with no
non-player advantage — it reads the board, never the deck.
- [ ] **#104** — Weigh a switching turn against drawing and the Freight Agent. Letting the planned gain
gate switching on its own measured nothing (0.1, 0.25) or worse (0.5): `usefulSwitching` already
says yes exactly when a plan gains. What would matter is a VALUE for the other two options to
compare against, which the bot does not have. See **Reference · #104**.
- [ ] **#106** — The Extra trap: a full hand of Extras the A/D cap is holding back cannot be discarded,
so the next draw forces one into a full Office. All 23 train plays past the cap in 40 games were
this. Avoiding the draw measured nothing (−0.03) because it stalled development. See
**Reference · #106**.
- [ ] **#105** — Plan across more than one turn. Jesse is in favour, one turn first to see the impact —
which is now measured. Deferred for a conversation, not declined. See **Reference · #105**.
- [ ] **#41** — The bot never plays Red Flags — zero in 200 games since Gitea#19, and that is deck
luck rather than unwillingness. It takes the danger prompt unconditionally; what it never does
is plant a flag ON PURPOSE to buy a Stage for switching, which needs it to know it wants time.
See **Reference · #41**.
- [ ] **#57** — The bot's priorities are not the problem — measured across ten heuristic variations.
**Read this before tuning weights**; it is the argument that the ceiling is elsewhere. See
**Reference · #57**.
**Read this before tuning weights**; it is the argument that the ceiling is elsewhere. **Part of
"elsewhere" was choosing one Move at a time**: planning the whole switching turn was worth
+2.89 (t = 15.8) in the 2026-09-14 bot-tuning round. See **Reference · #57**.
- [ ] **#59** — The run-around is out of reach of any bot, and the deck is why — measured five ways.
See **Reference · #59**.
- [ ] **#53** — The bot does not know to bring an expedited train back to the station. See **Reference
· #53**.
- [ ] **#54** — The bot cannot spot a car at a stub industry, and the cut-ordering rules made that
visible. See **Reference · #54**.
@@ -1424,6 +1539,12 @@ measuring deck luck rather than reachability, and its comment now says so.
#### #57 — THE BOT'S PRIORITIES ARE NOT THE PROBLEM — measured.
**2026-09-14 — confirmed, and one ceiling found.** Reordering priorities still moves nothing; what
moved the bot was SEARCH. `sim/switch-planner.ts` plans the whole switching turn against a score of
where it ends, and measured +2.89 ± 0.18 (t = 15.79) over 1600 paired seeds — freight loads and
unloads 0.22 → 1.53, Cargo phases with a car spotted 5% → 18%. The "8% of Cargo phases" below was a
fact about how the bot switched, not only about the deck. See `CHANGELOG.md`, 0.8.0.9.
**THE BOT'S PRIORITIES ARE NOT THE PROBLEM — measured.** Ten heuristic variations, each paired
over 400+ seeds. Every reordering of what the bot prefers came out inside the noise; the only
@@ -1449,6 +1570,21 @@ prioritised better. What is left is the economy itself, which is a deck question
#### #59 — THE RUN-AROUND IS OUT OF REACH OF ANY BOT, AND THE DECK IS WHY — measu…
**2026-09-14 — part of "the deck is why" was the bot's own draw.** It was taking ~32 face-up trains and
industries a game it could not play and discarding them again, so the hand rarely held track. Taking
only playable cards (now the default) grew districts 14 → 24 cards and run-arounds 3/60 → 9/60. And
~13 of the ~16 track pieces a game were being laid by the draw turn's "play what is in hand" fallback
at the first legal square, not by `bestTrackLay`; holding them (−0.83) and placing them by
`bestTrackLay`'s score (+0.07, noise) both failed, so the next limit is that SCORING — it cannot tell a
piece that opens an industry site or advances a run-around from one that fills a square. The deck
measurements below were taken before any of this and should be re-read with it in mind.
**2026-09-15 — the scoring, fixed.** `bestValuedLay` scores the layout a lay leaves (reachable industry
sites, a closed run-around, ways off the main) instead of the piece: closed run-arounds in 22 of 60
districts against 9, +0.118 ± 0.029 revenue (t = 4.14, 6400 seeds). The run-around is now reachable
without changing the deck; what is left is a bot that can USE one, which needs more than one turn of
planning (#105).
**THE RUN-AROUND IS OUT OF REACH OF ANY BOT, AND THE DECK IS WHY — measured, five ways.**
"Teach the bot to plan across turns" was tried properly and does not work. Every attempt is
@@ -1490,6 +1626,8 @@ the end of the game, drawing the fault **26 times**. Not an engine bug — the m
exactly as designed — but a clear next bot heuristic: prefer ending a switching turn with any
expedited crew back on the Office square, at least once it has finished the work it went out for.
**CLOSED 2026-09-14** — see **Done · 53**.
#### #54 — THE BOT CANNOT SPOT A CAR AT A STUB INDUSTRY, and the cut-ordering rul…
**THE BOT CANNOT SPOT A CAR AT A STUB INDUSTRY, and the cut-ordering rules made that visible.**
@@ -1511,6 +1649,12 @@ pass should not read the drop as a deck problem.
#### #58 — The bot cannot get a crew next to an industry, so Flying Switch never…
**2026-09-14 — the premise is now a deck fact.** Flying Switch is dealt **0 copies** (not in sheet 5;
Jesse, 2026-08-26), so no bot can fire it: across 30 standard games none was ever drawn. The switching
planner searches the card by default, so it will be used the day it is dealt again. The reachability
sweep's exemption in `sim.test.ts` stays until then.
**The bot cannot get a crew next to an industry, so Flying Switch never fires.** Industries are
now stub-only and the bot places 2.23 a game (was 3.84), in districts averaging under two rows
@@ -1558,6 +1702,114 @@ of zero" over 400 games — but at 400 games the standard error is ±0.33, so a
have looked like nothing. They are nearly free to re-run now and at least one may have been
discarded wrongly.
#### #104 — WEIGH SWITCHING AGAINST THE OTHER TWO OPTIONS, not against a threshold.
Measured 2026-09-14, paired over 400 seeds against the planner: switching only when the planned gain
clears a threshold scored −0.33 at 0.5 (t = −5.06), −0.01 at 0.25, +0.06 at 0.1 (t = 1.79). At 0.5 it
refuses turns that only collect cars, which feed later deliveries; below that it agrees with
`usefulSwitching`. The choice that is still made by a fixed ladder is WHICH of §6's three options a
Stage goes to, and the planner can now put a number on one of them. The other two need numbers of
their own — what a draw is worth given the hand and the Departments, what stocking a box is worth given
the cars spotted — before the three can be compared. Also: planning at every Local Operations decision
costs a full search each time, so any version of this has to stay cheap.
**2026-09-15 — measured the ceiling first: there is almost none.** At 907 real Local Operations choices
(75 standard games, seeds outside the usual measurement range), every legal option was tried and the rest
of the game played out by today's bot, 4 times each with the HIDDEN parts reshuffled — the Home Office
deck order and future rolls — and the same reshuffles for every option, so the comparison is paired.
Grouped by the rule that made the choice, the value of each alternative against it:
| the ladder chose | times | switch instead | draw instead | Freight Agent instead |
| --- | --- | --- | --- | --- |
| draw — nothing urgent, develop | 494 | +0.08 ± 0.12 | — | +0.02 ± 0.04 |
| Freight Agent — feed the pipeline | 147 | +0.07 ± 0.06 | −0.01 ± 0.04 | — |
| switch — a train with work at the Office | 108 | — | −0.17 ± 0.10 | −0.25 ± 0.08 |
| draw — an Office upgrade in hand | 100 | +0.21 ± 0.16 (6) | — | −0.08 ± 0.06 |
| switch — walk the crew home | 47 | — | +0.22 ± 0.14 | −0.15 ± 0.08 |
| draw — a train card in hand | 11 | — | — | −0.45 ± 0.24 |
No rule has an alternative that is significantly better; where the table leans, the ladder is usually
the one that is right. So given how the bot plays each option once chosen, the choice itself is close to
optimal, and value functions for draw and Freight Agent have little to find. The one lean worth a look if
this is reopened is walking a stranded crew home (+0.22, t ≈ 1.6). The rollout tool is analysis only —
the bot never sees a rollout.
#### #106 — THE EXTRA TRAP — why the cap on committed trains still lets an Office overfill.
**2026-09-15 — re-measured under today's defaults, and most of it is not the Extra trap.** 20 "no free
A/D track" collisions in 60 standard games, −1.67 revenue a game. No train was held (§8.2 or clearance) in
the Stage before any of them, and only 8 of the 20 trains destroyed were Extras. **Six destroyed a train
of the same number as a Second Section run within the previous two Stages — and all 26 Second Sections
the bot ran in those games were an accident:** the New Train phase's "no car on offer" fallback takes
`options[0]`, and `legalActions` lists `newTrain.secondSection` ahead of the Extra starts, so whenever an
Extra was waiting to start the bot doubled the train due out instead. The Office never had an A/D track
to spare for one.
**A RULES QUESTION FOR JESSE, found on the way — not a bot matter.** Q9 (`implications.md`) defines the
Second Section as a CARD "played on a train that is due out", and `content.ts` defines `SECOND_SECTION`
with 1 copy — but `buildDeck` never deals it, and `check`'s `newTrain.secondSection` asks for no card in
hand. So any player may run a Second Section for free on every train due out. Either the card should be
dealt and required, or the free action is the intended rule and Q9's wording is stale.
Also measured and removed: starting an over-cap Extra where its run never reaches the Office (+0.10,
t = 1.82) — such a start was on offer at 1 of 25 over-cap starts.
**The accident is fixed** (2026-09-15, default): the New Train fallback takes a car, a pass or the
Extra's start, never `options[0]` — +0.32 ± 0.09 (t = 3.64), collisions 0.24 → 0.19. **Still open under
this item:** the forced Extra itself (a full, undiscardable hand of Extras), and the Second Section card
question above.
`choose` removes train-card plays from the options when `trainWouldOverfillTheOffice`, but yields if
that would leave nothing legal. It does leave nothing legal in one ordinary position: the hand is over
the limit (§6.2 requires reducing it) and every card in it is an Extra, which `keepReason` forbids
discarding. Measured 2026-09-14 after the face-up take rule: 23 of 196 train plays in 40 games were past
the cap, every one "play what is in hand" with four Extras held, and the worst seeds each lost 3-4
collisions to it. Declining the draw option in that position measured −0.03 ± 0.11 — collisions fell
0.26 → 0.20 but cards played fell 29.0 → 25.6. Better answers to try: play the Extra at the least
dangerous moment rather than the first, count WHEN each committed train is due at the Office instead of
how many there are, or keep the hand from filling with Extras in the first place.
#### #105 — PLAN ACROSS TURNS — the evidence so far, for the conversation.
For: one-turn planning already reaches most switching work (Cargo phases with a car spotted 5% → 18%),
and what it cannot do is exactly what spans a Stage — leave a car on a spur for the next crew, or start
a run-around and finish it later. The planner already scores staged wanted cars (+0.2), which is a
first, crude step in that direction.
Against, for now: everything between two switching turns is not the player's — a Mainline Phase, trains
arriving, a Load/Unload phase — so a second turn cannot be searched the way the first is without either
simulating those phases (arrivals are on the public timetable, but cars on arriving trains are not
known) or scoring the position between turns more cleverly. And search cost is already what sets the
test suite's running time. Cheapest next step if taken up: a better score for "what the next turn can
still reach", not a deeper search.
**2026-09-15 — the run-around measurement that bears on this.** A candidate that lays track by what the
district can do afterwards (`valueLays`) more than doubled closed run-arounds, 9/60 → 22/60, yet moved
revenue only +0.14 ± 0.06 (t = 2.58, 1600 seeds). Traced over 40 games: the one-turn planner DOES use
the loop — 63 of 131 plans in a district with one end a move on it, 35 run it both ways — but its planned
gain per turn is the same with a run-around as without (0.28 against 0.27). A run-around is for putting a
train's cars in a different order, which pays off in the turns after; a planner that looks one turn
ahead has no way to value it.
**2026-09-15 — two-turn planning, built and measured: it does not pay.** `planTwoTurns` kept the six best
ends of a switching turn, removed the trains that would highball in the Mainline Phase in between (on the
Office square and made up — their cars leave with them), reset the Moves, planned the next turn from each,
and chose by the position after departures plus 0.8 of what the next turn adds. Nothing hidden is read.
| version | revenue, 400 paired seeds | what went wrong |
| --- | --- | --- |
| first | −0.28 ± 0.11 (t = −2.58) | an expedited train left away drew 31 faults in one game: the fault is charged in the gap, which the discounted next turn "recovered"; trains left away 5.3 a game against 3.1 |
| with the gap fault charged in full and 0.5 a Stage per train left away | −0.14 ± 0.06 (t = −2.36) | trains still left away 4.8 a game; the "crew must get back to the Office" choice 4.2 a game against 2.7 |
Why a second turn has so little to find, measured over 30 standard games:
- only **49%** of switching turns have the same train in the district at the next Local Operations choice;
- **6.1 wanted cars a game** do leave aboard departing trains — but a further switching turn from those
exact positions could have spotted only **0.7** of them: most were never deliverable;
- the one-turn planner already gains no more with a run-around than without (0.28 against 0.27).
And what a second turn COSTS is a Stage: trains left away have to be walked home, and those choices come
out of drawing and the Freight Agent (cards played 28.8 → 28.3). A multi-turn bot would have to weigh
switching against the other two options — which is #104, not a deeper search.
### Code health and housekeeping
#### #46 — tsc --noUnusedLocals finds 29 unused declarations across 14 files, and…
@@ -1801,6 +2053,15 @@ where it belongs, and it is still open.
Closed items, kept because several are the only record of a ruling or a lesson. Newest first within
each group.
### Closed in the 2026-09-14 bot-tuning round (unreleased)
53. ~~**The bot did not know to bring an expedited train back to the station.**~~ — done 2026-09-14,
not by a heuristic of its own but as a consequence of planning the switching turn: the planner's
score charges an expedited train left away from the Office a full Revenue point, which is what Q3
charges. Over 400 paired seeds, 19 games drew expedite faults under the rule ladder and **none**
under the planner, worth +1.07 a game (t = 3.83) — the two worst cases had drawn 45 and 42 faults
in a single game. See `CHANGELOG.md`, 0.8.0.9.
### Shipped through v0.7.9.8, from the queue
Closed items, newest first. Kept because several of them are the only record of a ruling or a lesson;