created harness to better test bots. improved bot play

This commit is contained in:
Jesse
2026-08-09 19:24:54 -04:00
parent 35575147cd
commit 4a800a053a
13 changed files with 6263 additions and 5503 deletions
+109 -17
View File
@@ -28,24 +28,84 @@ Ordered within each section by how much it is currently costing us.
(they have no east-west road and dead-end the main) cost it districts 28.0 → 19.7 cards and
revenue ~2.0 → 0.8. It has no plan for where a curve should go once the easy square is gone.
Same root cause as the two items below; fix them together, after the rebalance.
- [ ] **THE BOT'S PRIORITIES ARE NOT THE PROBLEM — measured.** Ten heuristic variations, each paired
over 400+ seeds. Every reordering of what the bot prefers came out inside the noise; the only
thing that moved revenue was refusing to schedule a train the Office cannot hold
(**+1.09 ± 0.16, t = 6.79** at 1600 seeds, revenue 1.19 → 2.28). Notable failures, all
instructive:
- **Refusing to bury the engine costs more than it saves** (−0.35, t = −2.98). It works —
burial falls from 8.4 decisions a game to 0.03 — and freight halves with it, because
coupling is mandatory (§A.4): the moves that bury the engine ARE the moves that pick cars
up. Burial is the price of collecting, not a mistake.
- **Reserving Moves to get home costs 0.55** (t = −2.32), though 62 of 120 trains left on the
board at game end were stranded in the district. The switching work is worth more than the
departures.
- **Granting clearance when the train ahead has one Stage left is −0.98** (t = −5.24). Trains
move in numeric order, so a follower can enter the region the leader still occupies before
the leader moves. "About to leave" is not "gone".
- Preferring coaches at make-up, stocking the platform first, playing Interlocking earlier,
hunting the Depot in the Departments: all within noise, and three of them were exact
no-ops — Interlocking sits in hand alongside a train card **0.04 decisions a game**.
The funnel says why: only **8% of Cargo phases** have a stocked green box, and the bot already
takes 42% of the turns where stocking is productive. The opportunities are not there to be
prioritised better. What is left is the economy itself, which is a deck question.
- [ ] **The bot cannot get a crew next to an industry, so Flying Switch never fires.** Industries are
now stub-only and the bot places 2.23 a game (was 3.84), in districts averaging under two rows
deep. `flyingSwitch` is exempted by name in the reachability sweep in `sim.test.ts`; deleting
that line is the test that this is fixed. Same root cause as the item below.
- [ ] **The bot does not play for a run-around any more, and revenue halved.** With track in the deck
a run-around needs a turnout, a matching curve, straights, a second curve and a second turnout,
all of the right hand, arriving in a three-card hand in a usable order. The bot holds no plan
across turns and discards a piece it cannot use immediately: run-arounds fell 70/100 → 29/100
and revenue 5.80 → 2.87. Two test floors in `sim.test.ts` are pinned below the measurement as
break-detectors rather than targets, and say so. Do the density re-measurement above first —
bot weakness and deck density are currently confounded.
- [ ] **A "load" is stored as a car, so rolling stock cannot be counted.** `outboundBox`,
`inboundBox` and `menAtWork` all hold `RollingStock`, and `freightAgent.stockOutbound` takes a
LOADED CAR out of the Division Yard to fill a green box. So a census of every holder comes to
92 against the 80 dealt at setup — not necessarily duplication, because some of those objects
are cargo in transit rather than cars, but there is no way to tell them apart. Until a load is
its own type, "is any stock being created or destroyed?" is an unanswerable question, and the
supply numbers below cannot be tuned with confidence.
- [ ] **THE RUN-AROUND IS OUT OF REACH OF ANY BOT, AND THE DECK IS WHY — measured, five ways.**
"Teach the bot to plan across turns" was tried properly and does not work. Every attempt is
neutral or negative, and they fail for one reason that the numbers make plain.
| attempt | result |
|---|---|
| hold ALL track for the siding | **−0.70** (t = −3.27) |
| hold only CURVES, the closing piece | **−0.26** (t = −3.22), district 17.9 → 16.7 cards |
| finish a run before cutting another way down | 0.00 — 398/400 games identical |
| treat a second turnout as the closing piece | 0.00 — **400/400 identical** |
| spend a curve only on a square that CLOSES | −0.11, and only 15 games in 400 differ at all |
**The pieces never meet.** Over 12,000 Local Operations turns: a turnout and a curve are in
hand together on **0.3%** of them, and a turnout with a MATCHING-hand curve on **0.2%** — about
once every eight games. A run-around needs five specific pieces of the right hands in a usable
order; the bot does not get to the two-piece prerequisite.
And it is not hand pressure. The hand is FULL — mean 2.66 cards, at the three-card limit on
78% of turns. The bot plays 11.4 track cards a game and discards 1.5, so it spends the pieces
as they arrive because a piece that builds anything outscores holding one that might build
more later. Holding is the only counter, and holding measures worse every way it is tried.
This is a consequence of moving track into the deck, not a bot weakness: 91 run-arounds per 100
games when track was a private 26-piece supply the player chose from, 29/100 once it was drawn,
4/60 now. **If the run-around is meant to be the central switching puzzle — and the rules
present it that way — the supply has to change, not the player.** Options: give track its own
hand or yard the way the prototype did, raise the hand limit for track specifically, or print a
siding as a single card. Nothing else reaches it.
- [ ] **CLEARING AN INBOUND BOX MINTS A CAR — measured at 1.29 a game against a supply of 80.**
Answered, and it is duplication after all. The two directions are not symmetrical:
- **Outbound is paid for.** `stockToOutbound` SPLICES a loaded car out of the Division Yard to
become the load, and `loadCompleted` swaps the emptied car into the Classification Yard as
the loaded one takes its place on the industry track. Objects in, objects out.
- **Inbound is not.** `unloadBegan` (apply.ts) turns one loaded car into an empty car on the
track **plus** a load on MEN|AT|WORK, and `passengersDetrained` does the same to a coach —
one loaded coach becomes an empty coach in the train plus an object in the red box. Then
`inboundCleared` pushes that object into the **Classification Yard as a car**. The cargo
becomes rolling stock, while the car it came out of is already back in service.
Measured over 200 games: cars at the end minus 80, plus collision losses, equals the
`inboundCleared` count in 71/200 games exactly and 258 against 279 in total — the rest is
loads still in flight at the final whistle. So the supply inflates by about 1.3 cars a game.
That is the number `ROLLING_STOCK_SUPPLY` is supposed to control, so **no supply figure below
can be tuned until this is settled**. The fix is a decision, not a patch: either a load stops
being a `RollingStock` and becomes its own type, or `inboundCleared` discards rather than
banking. Found by a conservation audit, not by a failing test.
- [ ] **`state = fold(events)` is not literally true, and the README says it is.** Replaying the
event log onto a fresh state throws: the phase driver mutates state directly and emits a
descriptive event afterwards — `newTrainPhase` does `s.trays.set(...)` and then pushes
`trainMadeUp`. Replay works because it re-applies INTENTS (`fromSave`), not because folding
events reconstructs the position. Nothing is broken today, but the claim underwrites
reconnection and restart recovery, which are unbuilt — so it should be either made true or
restated before anything is built on it.
- [ ] **Engines are not a SUPPLY yet, only a position.** `engineAt` now records where the engine
sits in the tray and the consist shows it, but an engine is still conjured with the tray
rather than drawn from the Division Yard and returned to it. The rules put engines in the
@@ -142,9 +202,36 @@ target is settled and freight carries its intended share.
- [ ] **Train density.** Left alone by decision, but noted: 22 train cards in 140 are drawn less often
than 22 in 115 were, and trains scheduled fell 2.9 → 2.1 as a side effect of the other density
changes.
- [ ] **The victory target itself** (20 over 5 Days). 5 wins in 100, up from 1, and the bot now does
exploit sidings — so that unclaimed gain has been claimed and the target is still missed by a
wide margin (mean 6.0 against 20). This is the next real balance question.
- [ ] **The victory target (20 over 5 Days) is out of reach by a factor of about four, and the
Office ladder is why.** Measured over 800 games with the tuned bot, which no longer throws
revenue away on collisions (0.0 a game, down from 0.4):
| trains scheduled | games | revenue | | Office reached | games | trains | revenue |
|---|---|---|---|---|---|---|---|
| 0 | 110 | 0.67 | | Whistle Post | 297 | 0.81 | 0.62 |
| 1 | 379 | 1.69 | | Depot | 272 | 1.49 | 2.92 |
| 2 | 234 | 3.72 | | Station | 176 | 1.89 | 4.36 |
| 3 | 68 | 5.68 | | Terminal | 55 | 1.93 | 5.04 |
| 4 | 9 | 5.78 | | | | | |
Revenue is almost exactly linear in trains scheduled — about **1.9 a train** — and trains are
capped by A/D capacity, which is the Office tier, which is a card you have to draw. So the
whole economy hangs off one valve: **37% of games never leave the Whistle Post and earn 0.62;
53% of all games earn nothing at all.**
Extrapolating the line, 20 Revenue needs roughly **11 trains and therefore 11 A/D tracks**. A
Terminal has four. The target is not merely missed, it is structurally unreachable under this
deck at this Office ladder — no amount of bot skill closes it, and the best game seen in 800
was 26 against a median of 0.
The three ways out are all yours to choose between, and they are different games:
1. **Lower the target** to what a 5-Day game can produce (6–8 looks like the honest number).
2. **Open the valve** — more Office cards, or a cheaper first upgrade, or more A/D capacity at
the Whistle Post, so the ladder is climbed rather than drawn.
3. **Raise revenue per arrival.** It is 0.46 today; each arrival can in principle pay 2 for
passengers alone. That is the freight/passenger conversion problem, not the traffic problem.
Nothing here is a bot weakness any more, which is what this measurement was waiting on.
---
@@ -228,6 +315,11 @@ target is settled and freight carries its intended share.
close up.
- [ ] **Wide boards scroll.** A 40-card district and a 13-section Division both need horizontal
scrolling. Legible, not compact.
- [x] ~~**Every published replay was dead.**~~ All three replayed **2 intents of roughly 400** and
presented as short games, exactly as the item below predicted. Re-recorded from bot games with
`node src/sim/save-replay.ts`, which verifies each save round-trips before writing it, and
`harness.test.ts` now fails if a published replay stops short. The version-stamp item below is
still worth doing — this catches the breakage, it does not explain it to a player.
- [ ] **Save/restore is not version-aware.** A save from an older ruleset stops replaying rather than
failing loudly, which is the safe direction but says little about what changed. **This has now
bitten once**: both published replays were dead — one got 42 intents into 360, the other 4 of