created harness to better test bots. improved bot play
This commit is contained in:
+156
@@ -21,6 +21,162 @@ page as `v0.1.0 · <sha> · <date>`, so what is deployed can always be identifie
|
||||
|
||||
## Unreleased
|
||||
|
||||
## 0.2.0 — 2026-08-09
|
||||
|
||||
### The bot plays twice as well, and a rule it was never following
|
||||
|
||||
**Adopted, after measuring: cap the trains at the Office's A/D capacity, and operate before
|
||||
drawing.** Together **+1.52 ± 0.17 (t = 9.16)** over 1600 paired seeds — revenue **1.19 → 2.72**, wins
|
||||
11 → 21 in 1600, collisions 0.38 → 0.02. Both are now default play; the flags that carried them are
|
||||
gone, and what remains in `BotTweaks` are two ABLATIONS that turn them off, because the question a
|
||||
measured heuristic needs later is "is this still true?" rather than "does this help?".
|
||||
|
||||
The six candidates that measured neutral are deleted rather than left switched off. Adding all of
|
||||
them on top of the two winners was worth **−0.08**: +1.44 against +1.52 without.
|
||||
|
||||
**Jesse's idea — build first, then switch — was right about the axis and wrong about the half that
|
||||
mattered.** Reserving Day 1 for development measures +0.52; reserving none and simply preferring the
|
||||
work to the draw measures +0.48 with a far better spread (317 seeds better against 77, where the
|
||||
build phase gives 329 against 144). With the cap alongside, no build phase beats one Day beats two.
|
||||
So the keeper is "operate rather than draw", and the building phase — the intuitive part — is inert:
|
||||
the draw-priority ordering was already developing the district well enough.
|
||||
|
||||
It works through the constraint the funnel had been pointing at all along: green-stocked Cargo phases
|
||||
+60%, freight revenue +39%.
|
||||
|
||||
**And then the same tooling proved that constraint was the wrong one.** Stocking green boxes
|
||||
speculatively — filling them before a car is spotted — raises stocked phases **five-fold**, from 5.6
|
||||
to 28.8 of 60, and moves freight revenue by **nothing at all** (1.16 → 1.15, and −0.10 revenue
|
||||
overall, t = −3.62). The green box was never the gate. The gate is **the spotted car**, and the 8%
|
||||
figure was low because stocking is *conditional on* a car being there — it was measuring the car
|
||||
supply all along. Switching quality, not Freight Agent turns, is where the freight economy is won.
|
||||
|
||||
### The engine was not following two rules it prints
|
||||
|
||||
Chasing the Rolling Stock census found the cause, and it is not a modelling ambiguity after all.
|
||||
§9.3 says an unload requires "*an empty car of that type in the Division Yard*" and §9.2 says
|
||||
de-training requires "*a white empty coach in the Division Yard*". **Neither requirement was checked,
|
||||
and both reducers conjured the replacement car instead of taking it** — so every unload and every
|
||||
de-training minted a car, 1.29 a game against a supply of 80.
|
||||
|
||||
Both now take the car from the Division Yard, as the rules say. The census is conserved: censused
|
||||
after every batch across 80 games, **nothing changes it but a collision**, where before it drifted to
|
||||
119 against 80. A test does that check now, because this class of bug produces no symptom until a
|
||||
supply number is tuned against it.
|
||||
|
||||
It costs the bot about 0.09 revenue — unloading is meant to consume supply — and that is the point:
|
||||
`ROLLING_STOCK_SUPPLY` can be tuned against reality now.
|
||||
|
||||
### Where the game actually stands
|
||||
|
||||
With a bot that no longer wastes revenue, the balance question resolves. Revenue is linear in trains
|
||||
scheduled at about **1.9 a train**, and trains are capped by A/D capacity, which is the Office tier,
|
||||
which is a card you have to draw. 37% of games never leave the Whistle Post; **53% earn nothing at
|
||||
all**; the median game scores 0 and the best of 800 scored 26.
|
||||
|
||||
Twenty Revenue would need roughly **eleven trains and eleven A/D tracks**. A Terminal has four. The
|
||||
target is not missed, it is unreachable — and that is now a deck question with numbers behind it
|
||||
rather than a suspicion. `TODO.md` carries the three ways out.
|
||||
|
||||
### Five ways to teach the bot to plan a siding, and why none of them can work
|
||||
|
||||
Asked directly whether the bot could think across turns and stop building dead-end stubs. It can be
|
||||
taught to; it does not help, and finding out why was worth more than the attempts.
|
||||
|
||||
| attempt | result |
|
||||
|---|---|
|
||||
| hold ALL track for the siding | **−0.70** (t = −3.27) |
|
||||
| hold only CURVES — the only piece that can climb back to the main | **−0.26** (t = −3.22) |
|
||||
| finish an open run before cutting another way down | 0.00 — 398/400 identical |
|
||||
| treat a second turnout as the piece that closes the loop | 0.00 — **400/400 identical** |
|
||||
| spend a curve only on a square that actually closes a run | −0.11, 15 games in 400 differ |
|
||||
|
||||
**The pieces never meet.** Over 12,000 Local Operations turns, a turnout and a curve are in hand
|
||||
together on **0.3%** of them, and a turnout with a MATCHING-hand curve on **0.2%** — about once every
|
||||
eight games. A run-around needs five specific pieces of the right hands arriving in a usable order,
|
||||
and the bot does not reach the two-piece prerequisite, let alone the fifth.
|
||||
|
||||
**It is not hand pressure**, which is what the first two attempts assumed. The hand is full: mean
|
||||
2.66 cards, at the three-card limit on 78% of turns. The bot plays 11.4 track cards a game against
|
||||
1.5 discarded, spending each piece as it arrives because a piece that builds something now outscores
|
||||
holding one that might build more later — and the measurements say it is right to.
|
||||
|
||||
So the dead-end stubs are not a planning failure. They are what a five-piece structure looks like
|
||||
when the pieces are drawn one at a time from a 243-card deck: 91 run-arounds per 100 games when
|
||||
track was a private supply the player chose from, 29/100 once track was drawn, 4/60 today. `TODO.md`
|
||||
carries the options, and all of them are deck changes rather than bot changes.
|
||||
|
||||
### Ten ways to make the bot better, and one of them works
|
||||
|
||||
Every heuristic below was measured paired over 400+ seeds, and confirmed at 1600 before being
|
||||
believed. Nothing is adopted yet — every flag is off, so the bot plays exactly as it did.
|
||||
|
||||
**The only thing that moves revenue is refusing to schedule a train the Office cannot hold.**
|
||||
`trainCapSlack=0` — never more committed trains than A/D tracks — is **+1.09 ± 0.16 (t = 6.79)**,
|
||||
taking the bot from 1.19 to 2.28 mean. Every other reordering of the bot's preferences measured
|
||||
inside the noise. Adding all six of the small ones on top of the cap is worth a further 0.10.
|
||||
|
||||
**Three failures worth more than the success.**
|
||||
|
||||
*Refusing to bury the engine costs 0.35 a game* (t = −2.98). The tweak works perfectly — burial
|
||||
falls from 8.4 decisions a game to 0.03 — and freight halves along with it. Coupling is mandatory
|
||||
(§A.4), so **the moves that bury the engine are the moves that pick cars up**. Burial is the price
|
||||
of collecting, not a mistake to be coached out. The refinement that only refuses to ARRIVE at the
|
||||
Office buried is +0.16 and inside the noise.
|
||||
|
||||
*Reserving Moves to get home costs 0.55* (t = −2.32), even though 62 of the 120 trains still on the
|
||||
board at game end were stranded in the district, unable to depart from anywhere but the Office. The
|
||||
switching work is worth more than the departures it forfeits.
|
||||
|
||||
*Granting clearance when the train ahead has one Stage left costs 0.98* (t = −5.24) — a clean
|
||||
rejection of a rule that looked safe. Q13 collides on catching up, so a follower let onto a card
|
||||
whose occupant is leaving "cannot" catch it — except trains move in numeric order, so the follower
|
||||
can enter the region the leader still occupies before the leader has moved. **"About to leave" is
|
||||
not "gone".**
|
||||
|
||||
**And three exact no-ops, each for a different reason.** Playing Interlocking ahead of a train card
|
||||
changed nothing because the two sit in hand together **0.04 decisions a game**. Stocking the Office
|
||||
platform first changed nothing because the bot already chooses it 79% of the time. Spending an idle
|
||||
turn on the Freight Agent changed nothing because it asked the same predicate a branch three steps
|
||||
earlier already acts on — a tautology, and 400/400 identical games is what one looks like.
|
||||
|
||||
The funnel explains the pattern: **8% of Cargo phases have a stocked green box**, and the bot already
|
||||
takes 42% of the turns where stocking is productive. There is nothing to prioritise better. What
|
||||
remains is the economy itself, which is a deck question rather than a bot one.
|
||||
|
||||
### A conservation audit, and the bug it found
|
||||
|
||||
**Clearing an inbound box mints a car — 1.29 a game against a supply of 80.** `TODO.md` had this as
|
||||
an open question ("no way to tell a load from a car"); it is duplication, and the two directions are
|
||||
not symmetrical:
|
||||
|
||||
- **Outbound is paid for.** `stockToOutbound` splices a loaded car OUT of the Division Yard to become
|
||||
the load, and `loadCompleted` banks the emptied car as the loaded one takes its place.
|
||||
- **Inbound is not.** `unloadBegan` turns one loaded car into an empty car *plus* a load on
|
||||
MEN|AT|WORK, `passengersDetrained` does the same to a coach — and `inboundCleared` then pushes that
|
||||
load into the **Classification Yard as a car**, while the car it came out of is already back in
|
||||
service.
|
||||
|
||||
Over 200 games, cars-at-the-end minus 80 plus collision losses equals the `inboundCleared` count
|
||||
exactly in 71/200 games and 258 against 279 overall; the remainder is loads still in flight at the
|
||||
final whistle. Since `ROLLING_STOCK_SUPPLY` is the number that is supposed to set supply pressure,
|
||||
no supply figure can be tuned until this is decided. Not fixed here — it is a modelling decision.
|
||||
|
||||
Also found: **`state = fold(events)` is not literally true.** Replaying the event log onto a fresh
|
||||
state throws, because the phase driver mutates state directly and emits a descriptive event
|
||||
afterwards. Replay works by re-applying intents, not by folding events. Nothing is broken today, but
|
||||
the README claims the property and reconnection would rest on it.
|
||||
|
||||
### The replays were all dead, and now they cannot be
|
||||
|
||||
All three published replays managed **2 intents of roughly 400** — the site was serving three
|
||||
recordings of nothing, exactly as `TODO.md` predicted would happen silently. `save-replay.ts` records
|
||||
bot games as saves, and **verifies every one round-trips before writing it**: same revenue, same Day,
|
||||
same intent count. `harness.test.ts` fails if any published replay stops short of its own history.
|
||||
|
||||
Six replays published, 18 to 23 Revenue, against a target of 20 and a bot median of 0 — including
|
||||
one game with four collisions that still cleared the target, and two with none.
|
||||
|
||||
## 0.1.1 — 2026-08-08
|
||||
|
||||
### A way to tell whether a change to the bot helped
|
||||
|
||||
Reference in New Issue
Block a user