created harness to better test bots. improved bot play

This commit is contained in:
Jesse
2026-08-09 19:24:54 -04:00
parent 35575147cd
commit 4a800a053a
13 changed files with 6263 additions and 5503 deletions
+156
View File
@@ -21,6 +21,162 @@ page as `v0.1.0 · <sha> · <date>`, so what is deployed can always be identifie
## Unreleased
## 0.2.0 — 2026-08-09
### The bot plays twice as well, and a rule it was never following
**Adopted, after measuring: cap the trains at the Office's A/D capacity, and operate before
drawing.** Together **+1.52 ± 0.17 (t = 9.16)** over 1600 paired seeds — revenue **1.19 → 2.72**, wins
11 → 21 in 1600, collisions 0.38 → 0.02. Both are now default play; the flags that carried them are
gone, and what remains in `BotTweaks` are two ABLATIONS that turn them off, because the question a
measured heuristic needs later is "is this still true?" rather than "does this help?".
The six candidates that measured neutral are deleted rather than left switched off. Adding all of
them on top of the two winners was worth **−0.08**: +1.44 against +1.52 without.
**Jesse's idea — build first, then switch — was right about the axis and wrong about the half that
mattered.** Reserving Day 1 for development measures +0.52; reserving none and simply preferring the
work to the draw measures +0.48 with a far better spread (317 seeds better against 77, where the
build phase gives 329 against 144). With the cap alongside, no build phase beats one Day beats two.
So the keeper is "operate rather than draw", and the building phase — the intuitive part — is inert:
the draw-priority ordering was already developing the district well enough.
It works through the constraint the funnel had been pointing at all along: green-stocked Cargo phases
+60%, freight revenue +39%.
**And then the same tooling proved that constraint was the wrong one.** Stocking green boxes
speculatively — filling them before a car is spotted — raises stocked phases **five-fold**, from 5.6
to 28.8 of 60, and moves freight revenue by **nothing at all** (1.16 → 1.15, and −0.10 revenue
overall, t = −3.62). The green box was never the gate. The gate is **the spotted car**, and the 8%
figure was low because stocking is *conditional on* a car being there — it was measuring the car
supply all along. Switching quality, not Freight Agent turns, is where the freight economy is won.
### The engine was not following two rules it prints
Chasing the Rolling Stock census found the cause, and it is not a modelling ambiguity after all.
§9.3 says an unload requires "*an empty car of that type in the Division Yard*" and §9.2 says
de-training requires "*a white empty coach in the Division Yard*". **Neither requirement was checked,
and both reducers conjured the replacement car instead of taking it** — so every unload and every
de-training minted a car, 1.29 a game against a supply of 80.
Both now take the car from the Division Yard, as the rules say. The census is conserved: censused
after every batch across 80 games, **nothing changes it but a collision**, where before it drifted to
119 against 80. A test does that check now, because this class of bug produces no symptom until a
supply number is tuned against it.
It costs the bot about 0.09 revenue — unloading is meant to consume supply — and that is the point:
`ROLLING_STOCK_SUPPLY` can be tuned against reality now.
### Where the game actually stands
With a bot that no longer wastes revenue, the balance question resolves. Revenue is linear in trains
scheduled at about **1.9 a train**, and trains are capped by A/D capacity, which is the Office tier,
which is a card you have to draw. 37% of games never leave the Whistle Post; **53% earn nothing at
all**; the median game scores 0 and the best of 800 scored 26.
Twenty Revenue would need roughly **eleven trains and eleven A/D tracks**. A Terminal has four. The
target is not missed, it is unreachable — and that is now a deck question with numbers behind it
rather than a suspicion. `TODO.md` carries the three ways out.
### Five ways to teach the bot to plan a siding, and why none of them can work
Asked directly whether the bot could think across turns and stop building dead-end stubs. It can be
taught to; it does not help, and finding out why was worth more than the attempts.
| attempt | result |
|---|---|
| hold ALL track for the siding | **−0.70** (t = −3.27) |
| hold only CURVES — the only piece that can climb back to the main | **−0.26** (t = −3.22) |
| finish an open run before cutting another way down | 0.00 — 398/400 identical |
| treat a second turnout as the piece that closes the loop | 0.00 — **400/400 identical** |
| spend a curve only on a square that actually closes a run | −0.11, 15 games in 400 differ |
**The pieces never meet.** Over 12,000 Local Operations turns, a turnout and a curve are in hand
together on **0.3%** of them, and a turnout with a MATCHING-hand curve on **0.2%** — about once every
eight games. A run-around needs five specific pieces of the right hands arriving in a usable order,
and the bot does not reach the two-piece prerequisite, let alone the fifth.
**It is not hand pressure**, which is what the first two attempts assumed. The hand is full: mean
2.66 cards, at the three-card limit on 78% of turns. The bot plays 11.4 track cards a game against
1.5 discarded, spending each piece as it arrives because a piece that builds something now outscores
holding one that might build more later — and the measurements say it is right to.
So the dead-end stubs are not a planning failure. They are what a five-piece structure looks like
when the pieces are drawn one at a time from a 243-card deck: 91 run-arounds per 100 games when
track was a private supply the player chose from, 29/100 once track was drawn, 4/60 today. `TODO.md`
carries the options, and all of them are deck changes rather than bot changes.
### Ten ways to make the bot better, and one of them works
Every heuristic below was measured paired over 400+ seeds, and confirmed at 1600 before being
believed. Nothing is adopted yet — every flag is off, so the bot plays exactly as it did.
**The only thing that moves revenue is refusing to schedule a train the Office cannot hold.**
`trainCapSlack=0` — never more committed trains than A/D tracks — is **+1.09 ± 0.16 (t = 6.79)**,
taking the bot from 1.19 to 2.28 mean. Every other reordering of the bot's preferences measured
inside the noise. Adding all six of the small ones on top of the cap is worth a further 0.10.
**Three failures worth more than the success.**
*Refusing to bury the engine costs 0.35 a game* (t = −2.98). The tweak works perfectly — burial
falls from 8.4 decisions a game to 0.03 — and freight halves along with it. Coupling is mandatory
(§A.4), so **the moves that bury the engine are the moves that pick cars up**. Burial is the price
of collecting, not a mistake to be coached out. The refinement that only refuses to ARRIVE at the
Office buried is +0.16 and inside the noise.
*Reserving Moves to get home costs 0.55* (t = −2.32), even though 62 of the 120 trains still on the
board at game end were stranded in the district, unable to depart from anywhere but the Office. The
switching work is worth more than the departures it forfeits.
*Granting clearance when the train ahead has one Stage left costs 0.98* (t = −5.24) — a clean
rejection of a rule that looked safe. Q13 collides on catching up, so a follower let onto a card
whose occupant is leaving "cannot" catch it — except trains move in numeric order, so the follower
can enter the region the leader still occupies before the leader has moved. **"About to leave" is
not "gone".**
**And three exact no-ops, each for a different reason.** Playing Interlocking ahead of a train card
changed nothing because the two sit in hand together **0.04 decisions a game**. Stocking the Office
platform first changed nothing because the bot already chooses it 79% of the time. Spending an idle
turn on the Freight Agent changed nothing because it asked the same predicate a branch three steps
earlier already acts on — a tautology, and 400/400 identical games is what one looks like.
The funnel explains the pattern: **8% of Cargo phases have a stocked green box**, and the bot already
takes 42% of the turns where stocking is productive. There is nothing to prioritise better. What
remains is the economy itself, which is a deck question rather than a bot one.
### A conservation audit, and the bug it found
**Clearing an inbound box mints a car — 1.29 a game against a supply of 80.** `TODO.md` had this as
an open question ("no way to tell a load from a car"); it is duplication, and the two directions are
not symmetrical:
- **Outbound is paid for.** `stockToOutbound` splices a loaded car OUT of the Division Yard to become
the load, and `loadCompleted` banks the emptied car as the loaded one takes its place.
- **Inbound is not.** `unloadBegan` turns one loaded car into an empty car *plus* a load on
MEN|AT|WORK, `passengersDetrained` does the same to a coach — and `inboundCleared` then pushes that
load into the **Classification Yard as a car**, while the car it came out of is already back in
service.
Over 200 games, cars-at-the-end minus 80 plus collision losses equals the `inboundCleared` count
exactly in 71/200 games and 258 against 279 overall; the remainder is loads still in flight at the
final whistle. Since `ROLLING_STOCK_SUPPLY` is the number that is supposed to set supply pressure,
no supply figure can be tuned until this is decided. Not fixed here — it is a modelling decision.
Also found: **`state = fold(events)` is not literally true.** Replaying the event log onto a fresh
state throws, because the phase driver mutates state directly and emits a descriptive event
afterwards. Replay works by re-applying intents, not by folding events. Nothing is broken today, but
the README claims the property and reconnection would rest on it.
### The replays were all dead, and now they cannot be
All three published replays managed **2 intents of roughly 400** — the site was serving three
recordings of nothing, exactly as `TODO.md` predicted would happen silently. `save-replay.ts` records
bot games as saves, and **verifies every one round-trips before writing it**: same revenue, same Day,
same intent count. `harness.test.ts` fails if any published replay stops short of its own history.
Six replays published, 18 to 23 Revenue, against a target of 20 and a bot median of 0 — including
one game with four collisions that still cleared the target, and two with none.
## 0.1.1 — 2026-08-08
### A way to tell whether a change to the bot helped