added harness to support bots being able to run / test strategies.
This commit is contained in:
@@ -21,6 +21,80 @@ page as `v0.1.0 · <sha> · <date>`, so what is deployed can always be identifie
|
||||
|
||||
## Unreleased
|
||||
|
||||
## 0.1.1 — 2026-08-08
|
||||
|
||||
### A way to tell whether a change to the bot helped
|
||||
|
||||
The bot averages **1.4 Revenue against a target of 20**, and the obvious next step is to teach it to
|
||||
play better. That step could not be taken honestly, because there was no way to tell whether a
|
||||
heuristic had helped: revenue has σ ≈ 9 across games, so two runs of the *identical* bot differ by
|
||||
about a point through nothing but the deal. Every single-change claim in this changelog before the
|
||||
Interlocking work is inside that noise, and `TODO.md` has said so for a while.
|
||||
|
||||
**`node src/sim/compare.ts 1600 trainCapSlack=1`** runs the current bot and one variant over the same
|
||||
deals and reports the per-seed difference. Giving both sides the same seed takes the deal out of the
|
||||
comparison: σ drops from ~9 on the level to **5.3 on the difference**, and 1600 seeds puts the
|
||||
standard error at **±0.13** — in 1m45s. The noise floor moves from ±1.0 to about ±0.15, which makes
|
||||
every heuristic in the training plan resolvable. No parallelism needed; 100 games take 4.8s.
|
||||
|
||||
**The report prints the better/worse/identical split beside the mean**, because they are different
|
||||
claims. The first tweak measured is the case in point — `trainCapSlack=1` over 1600 paired seeds:
|
||||
|
||||
```
|
||||
REVENUE DELTA +0.77 ± 0.13 (t = 5.97, σ of the paired difference 5.13)
|
||||
seeds better 146 · worse 105 · identical 1349
|
||||
best seeds: +75, +50, +46 worst seeds: -26, -20, -13
|
||||
```
|
||||
|
||||
**84% of games are untouched.** It does not make the bot play better; it removes a rare catastrophe,
|
||||
and the mean rides on a handful of rescued games. Reporting that as "revenue up 65%" would be
|
||||
arithmetically true and misleading about what changed. At 400 seeds the same tweak read t = 2.57 —
|
||||
"not proven" — which is exactly the verdict it deserved there.
|
||||
|
||||
The tweak is **measured but not adopted**: the flag stays off, so this commit changes no bot
|
||||
behaviour. Turning it on is a change to how the bot plays and belongs in its own reviewed commit.
|
||||
|
||||
**And it prints the funnel for both sides**, because revenue can rise two ways: a channel started
|
||||
working, or an expensive channel was abandoned for a cheap one. A strict train cap raises revenue
|
||||
*and* cuts freight events nearly in half, and that has to be visible rather than inferred.
|
||||
|
||||
**The funnel is new** — `GameStats.funnel`, sampled live rather than recovered from the log, because
|
||||
the interesting gates are conditions rather than occurrences. "Was a green box stocked while a car
|
||||
was spotted" is not a thing that happens; it is true or false at a moment. It needed a per-decision
|
||||
hook on `playGame` (`TurnObserver`), separate from the existing event observer, so a phase is
|
||||
sampled once instead of once per event in the batch. What it says about the current bot:
|
||||
|
||||
- **Passengers:** of 7.4 arrivals a game, 70% reach a Passenger Facility, 25% carry the empty coach
|
||||
boarding requires, 21% the loaded coach detraining requires.
|
||||
- **Freight:** of 60 Cargo phases, a green box is stocked in **8%** and all three requirements meet
|
||||
at one industry in **7%**. Freight is gated almost entirely on stocking.
|
||||
- **Stuck:** 8.4 decisions a game are taken with a train's engine buried mid-consist, and on **3%**
|
||||
of them is there a legal way to set the nose cars out — because it happens at the Office, where
|
||||
Rolling Stock may not be left.
|
||||
|
||||
**`developerBot` is now `makeDeveloperBot({})`**, byte-identical to what came before (asserted on
|
||||
three seeds by full event-stream fingerprint, and 1.4400 mean on 100 games either way). Variants
|
||||
exist so two policies can be compared in one process rather than by editing the bot between runs —
|
||||
which is how you end up comparing two things you cannot reproduce. Tweaks are temporary: a flag that
|
||||
measures well becomes the default and is deleted in the same commit.
|
||||
|
||||
**The first tweak found the bug this tooling exists to find, twice over.** `trainCapSlack` caps
|
||||
committed trains against the Office's A/D tracks. Written first as a gate on the two branches whose
|
||||
comments say they exist to play a train card, it measured **exactly zero difference over 400 paired
|
||||
seeds** — because `followThrough` ends with a generic "play what is in hand" fallback that played the
|
||||
card anyway. A cap has to remove the option, not guard the branches that reach for it.
|
||||
|
||||
Then the *test* for it was wrong in the same shape: it asserted only that some decision changed
|
||||
somewhere, and **passed against the broken bot**, because a tight cap does change which Local
|
||||
Operations option gets chosen — it just fails to stop the card being played two steps later. It now
|
||||
asserts the contract (committed trains never exceed what the Office can hold) and is verified to fail
|
||||
against the broken version. "Something moved" is not the promise.
|
||||
|
||||
Also new: a determinism self-check. The same policy on the same seed must produce an identical event
|
||||
stream, since the paired method rests on it and the failure mode is silent.
|
||||
|
||||
Tests 403 → 413.
|
||||
|
||||
## 0.1.0 — 2026-08-08
|
||||
|
||||
The first numbered build. Everything below the "second playtest pass" heading was made under
|
||||
|
||||
Reference in New Issue
Block a user