added harness to support bots being able to run / test strategies.

This commit is contained in:
Jesse
2026-08-08 15:07:42 -04:00
parent 5d825b97d2
commit 35575147cd
9 changed files with 866 additions and 23 deletions
+74
View File
@@ -21,6 +21,80 @@ page as `v0.1.0 · <sha> · <date>`, so what is deployed can always be identifie
## Unreleased
## 0.1.1 — 2026-08-08
### A way to tell whether a change to the bot helped
The bot averages **1.4 Revenue against a target of 20**, and the obvious next step is to teach it to
play better. That step could not be taken honestly, because there was no way to tell whether a
heuristic had helped: revenue has σ ≈ 9 across games, so two runs of the *identical* bot differ by
about a point through nothing but the deal. Every single-change claim in this changelog before the
Interlocking work is inside that noise, and `TODO.md` has said so for a while.
**`node src/sim/compare.ts 1600 trainCapSlack=1`** runs the current bot and one variant over the same
deals and reports the per-seed difference. Giving both sides the same seed takes the deal out of the
comparison: σ drops from ~9 on the level to **5.3 on the difference**, and 1600 seeds puts the
standard error at **±0.13** — in 1m45s. The noise floor moves from ±1.0 to about ±0.15, which makes
every heuristic in the training plan resolvable. No parallelism needed; 100 games take 4.8s.
**The report prints the better/worse/identical split beside the mean**, because they are different
claims. The first tweak measured is the case in point — `trainCapSlack=1` over 1600 paired seeds:
```
REVENUE DELTA +0.77 ± 0.13 (t = 5.97, σ of the paired difference 5.13)
seeds better 146 · worse 105 · identical 1349
best seeds: +75, +50, +46 worst seeds: -26, -20, -13
```
**84% of games are untouched.** It does not make the bot play better; it removes a rare catastrophe,
and the mean rides on a handful of rescued games. Reporting that as "revenue up 65%" would be
arithmetically true and misleading about what changed. At 400 seeds the same tweak read t = 2.57 —
"not proven" — which is exactly the verdict it deserved there.
The tweak is **measured but not adopted**: the flag stays off, so this commit changes no bot
behaviour. Turning it on is a change to how the bot plays and belongs in its own reviewed commit.
**And it prints the funnel for both sides**, because revenue can rise two ways: a channel started
working, or an expensive channel was abandoned for a cheap one. A strict train cap raises revenue
*and* cuts freight events nearly in half, and that has to be visible rather than inferred.
**The funnel is new** — `GameStats.funnel`, sampled live rather than recovered from the log, because
the interesting gates are conditions rather than occurrences. "Was a green box stocked while a car
was spotted" is not a thing that happens; it is true or false at a moment. It needed a per-decision
hook on `playGame` (`TurnObserver`), separate from the existing event observer, so a phase is
sampled once instead of once per event in the batch. What it says about the current bot:
- **Passengers:** of 7.4 arrivals a game, 70% reach a Passenger Facility, 25% carry the empty coach
boarding requires, 21% the loaded coach detraining requires.
- **Freight:** of 60 Cargo phases, a green box is stocked in **8%** and all three requirements meet
at one industry in **7%**. Freight is gated almost entirely on stocking.
- **Stuck:** 8.4 decisions a game are taken with a train's engine buried mid-consist, and on **3%**
of them is there a legal way to set the nose cars out — because it happens at the Office, where
Rolling Stock may not be left.
**`developerBot` is now `makeDeveloperBot({})`**, byte-identical to what came before (asserted on
three seeds by full event-stream fingerprint, and 1.4400 mean on 100 games either way). Variants
exist so two policies can be compared in one process rather than by editing the bot between runs —
which is how you end up comparing two things you cannot reproduce. Tweaks are temporary: a flag that
measures well becomes the default and is deleted in the same commit.
**The first tweak found the bug this tooling exists to find, twice over.** `trainCapSlack` caps
committed trains against the Office's A/D tracks. Written first as a gate on the two branches whose
comments say they exist to play a train card, it measured **exactly zero difference over 400 paired
seeds** — because `followThrough` ends with a generic "play what is in hand" fallback that played the
card anyway. A cap has to remove the option, not guard the branches that reach for it.
Then the *test* for it was wrong in the same shape: it asserted only that some decision changed
somewhere, and **passed against the broken bot**, because a tight cap does change which Local
Operations option gets chosen — it just fails to stop the card being played two steps later. It now
asserts the contract (committed trains never exceed what the Office can hold) and is verified to fail
against the broken version. "Something moved" is not the promise.
Also new: a determinism self-check. The same policy on the same seed must produce an identical event
stream, since the paired method rests on it and the failure mode is silent.
Tests 403 → 413.
## 0.1.0 — 2026-08-08
The first numbered build. Everything below the "second playtest pass" heading was made under