added harness to support bots being able to run / test strategies.
This commit is contained in:
@@ -21,6 +21,80 @@ page as `v0.1.0 · <sha> · <date>`, so what is deployed can always be identifie
|
||||
|
||||
## Unreleased
|
||||
|
||||
## 0.1.1 — 2026-08-08
|
||||
|
||||
### A way to tell whether a change to the bot helped
|
||||
|
||||
The bot averages **1.4 Revenue against a target of 20**, and the obvious next step is to teach it to
|
||||
play better. That step could not be taken honestly, because there was no way to tell whether a
|
||||
heuristic had helped: revenue has σ ≈ 9 across games, so two runs of the *identical* bot differ by
|
||||
about a point through nothing but the deal. Every single-change claim in this changelog before the
|
||||
Interlocking work is inside that noise, and `TODO.md` has said so for a while.
|
||||
|
||||
**`node src/sim/compare.ts 1600 trainCapSlack=1`** runs the current bot and one variant over the same
|
||||
deals and reports the per-seed difference. Giving both sides the same seed takes the deal out of the
|
||||
comparison: σ drops from ~9 on the level to **5.3 on the difference**, and 1600 seeds puts the
|
||||
standard error at **±0.13** — in 1m45s. The noise floor moves from ±1.0 to about ±0.15, which makes
|
||||
every heuristic in the training plan resolvable. No parallelism needed; 100 games take 4.8s.
|
||||
|
||||
**The report prints the better/worse/identical split beside the mean**, because they are different
|
||||
claims. The first tweak measured is the case in point — `trainCapSlack=1` over 1600 paired seeds:
|
||||
|
||||
```
|
||||
REVENUE DELTA +0.77 ± 0.13 (t = 5.97, σ of the paired difference 5.13)
|
||||
seeds better 146 · worse 105 · identical 1349
|
||||
best seeds: +75, +50, +46 worst seeds: -26, -20, -13
|
||||
```
|
||||
|
||||
**84% of games are untouched.** It does not make the bot play better; it removes a rare catastrophe,
|
||||
and the mean rides on a handful of rescued games. Reporting that as "revenue up 65%" would be
|
||||
arithmetically true and misleading about what changed. At 400 seeds the same tweak read t = 2.57 —
|
||||
"not proven" — which is exactly the verdict it deserved there.
|
||||
|
||||
The tweak is **measured but not adopted**: the flag stays off, so this commit changes no bot
|
||||
behaviour. Turning it on is a change to how the bot plays and belongs in its own reviewed commit.
|
||||
|
||||
**And it prints the funnel for both sides**, because revenue can rise two ways: a channel started
|
||||
working, or an expensive channel was abandoned for a cheap one. A strict train cap raises revenue
|
||||
*and* cuts freight events nearly in half, and that has to be visible rather than inferred.
|
||||
|
||||
**The funnel is new** — `GameStats.funnel`, sampled live rather than recovered from the log, because
|
||||
the interesting gates are conditions rather than occurrences. "Was a green box stocked while a car
|
||||
was spotted" is not a thing that happens; it is true or false at a moment. It needed a per-decision
|
||||
hook on `playGame` (`TurnObserver`), separate from the existing event observer, so a phase is
|
||||
sampled once instead of once per event in the batch. What it says about the current bot:
|
||||
|
||||
- **Passengers:** of 7.4 arrivals a game, 70% reach a Passenger Facility, 25% carry the empty coach
|
||||
boarding requires, 21% the loaded coach detraining requires.
|
||||
- **Freight:** of 60 Cargo phases, a green box is stocked in **8%** and all three requirements meet
|
||||
at one industry in **7%**. Freight is gated almost entirely on stocking.
|
||||
- **Stuck:** 8.4 decisions a game are taken with a train's engine buried mid-consist, and on **3%**
|
||||
of them is there a legal way to set the nose cars out — because it happens at the Office, where
|
||||
Rolling Stock may not be left.
|
||||
|
||||
**`developerBot` is now `makeDeveloperBot({})`**, byte-identical to what came before (asserted on
|
||||
three seeds by full event-stream fingerprint, and 1.4400 mean on 100 games either way). Variants
|
||||
exist so two policies can be compared in one process rather than by editing the bot between runs —
|
||||
which is how you end up comparing two things you cannot reproduce. Tweaks are temporary: a flag that
|
||||
measures well becomes the default and is deleted in the same commit.
|
||||
|
||||
**The first tweak found the bug this tooling exists to find, twice over.** `trainCapSlack` caps
|
||||
committed trains against the Office's A/D tracks. Written first as a gate on the two branches whose
|
||||
comments say they exist to play a train card, it measured **exactly zero difference over 400 paired
|
||||
seeds** — because `followThrough` ends with a generic "play what is in hand" fallback that played the
|
||||
card anyway. A cap has to remove the option, not guard the branches that reach for it.
|
||||
|
||||
Then the *test* for it was wrong in the same shape: it asserted only that some decision changed
|
||||
somewhere, and **passed against the broken bot**, because a tight cap does change which Local
|
||||
Operations option gets chosen — it just fails to stop the card being played two steps later. It now
|
||||
asserts the contract (committed trains never exceed what the Office can hold) and is verified to fail
|
||||
against the broken version. "Something moved" is not the promise.
|
||||
|
||||
Also new: a determinism self-check. The same policy on the same seed must produce an identical event
|
||||
stream, since the paired method rests on it and the failure mode is silent.
|
||||
|
||||
Tests 403 → 413.
|
||||
|
||||
## 0.1.0 — 2026-08-08
|
||||
|
||||
The first numbered build. Everything below the "second playtest pass" heading was made under
|
||||
|
||||
@@ -64,6 +64,25 @@ npm run typecheck # tsc --noEmit
|
||||
Because Node strips types rather than compiling them, the codebase is restricted to **erasable
|
||||
syntax**: no `enum`, no parameter properties, no namespaces. `tsconfig.json` enforces this.
|
||||
|
||||
### Measuring the bot
|
||||
|
||||
```sh
|
||||
node src/sim/harness.ts 200 # how the bot does, with the funnel
|
||||
node src/sim/compare.ts 1600 trainCapSlack=1 # one change, paired against the current bot
|
||||
```
|
||||
|
||||
**Never judge a heuristic on an unpaired run.** Revenue has σ ≈ 9 across games, so two runs of the
|
||||
*identical* bot differ by about a point through nothing but the deal. `compare.ts` gives both
|
||||
policies the same seed and reports the per-seed difference, where σ is 5.3 — 1600 seeds puts the
|
||||
standard error at ±0.13, in under two minutes. Keep a change at **t ≥ 3**, and read the
|
||||
better/worse/identical split beside the mean: a gain carried by a few rescued games is a different
|
||||
claim from one spread across the field.
|
||||
|
||||
Variants come from `makeDeveloperBot(tweaks)`. A tweak is **temporary** — when it measures well it
|
||||
becomes the default and the flag is deleted in the same commit; when it measures badly it is deleted
|
||||
with the finding recorded in `CHANGELOG.md`. A bot that accumulates switches nobody can account for
|
||||
is the thing this machinery exists to prevent.
|
||||
|
||||
## Design notes worth knowing
|
||||
|
||||
- **The rules engine is pure.** No I/O, no clock, no sockets, and all randomness derives from one
|
||||
|
||||
@@ -66,12 +66,17 @@ Ordered within each section by how much it is currently costing us.
|
||||
Revisit when setup gains an interactive phase; the orientation matters, because it decides
|
||||
which direction climbs and therefore what Brakeman and Helpers are worth.
|
||||
|
||||
- [ ] **Measure with error bars from now on.** Revenue has a standard deviation of ~9, so a
|
||||
100-game run carries about ±1.0 of noise — every single-change revenue claim in the changelog
|
||||
before the Interlocking work is inside that. Use paired per-seed comparison (the harness deals
|
||||
the same seeds either way) and 400+ games before calling a heuristic change good or bad. The
|
||||
first attempt at the Running Track straight was read as a 0.6 REGRESSION on 100 games and is
|
||||
a 0.7 improvement on 400.
|
||||
- [x] ~~**Measure with error bars from now on.**~~ Built: `node src/sim/compare.ts 1600 <tweak>=<n>`
|
||||
runs the current bot and one variant over the same deals and reports the paired difference.
|
||||
Pairing drops σ from ~9 on the level to **5.3 on the difference**, so 1600 seeds gives ±0.13 in
|
||||
about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is
|
||||
**t ≥ 3**, and the report prints the better/worse/identical split beside the mean, because a
|
||||
mean carried by a skewed tail is a different claim from broad improvement.
|
||||
- [ ] **Re-run the three "worth ~0" action-mix experiments against the new floor.** Capping the draw,
|
||||
pairing the two halves of a load, and restricting Enhancements were each measured "within noise
|
||||
of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would
|
||||
have looked like nothing. They are nearly free to re-run now and at least one may have been
|
||||
discarded wrongly.
|
||||
- [x] ~~**Confirm the Classification Yard rule against the source.**~~ Confirmed, and the guess was
|
||||
wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses
|
||||
straight back to the Division Yard, and the Classification Yard empties ONLY when the Division
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "station-master",
|
||||
"version": "0.1.0",
|
||||
"version": "0.1.1",
|
||||
"private": true,
|
||||
"type": "module",
|
||||
"description": "Station Master — a railroad operations game",
|
||||
|
||||
+115
-10
@@ -20,7 +20,7 @@
|
||||
*/
|
||||
|
||||
import { applyIntent, areaOf, canAdvanceLoad, destinationsFor, facilityCarTypes, laborersLeft } from '../engine/apply.ts';
|
||||
import { MAX_CONSIST, nextOfficeTier } from '../engine/content.ts';
|
||||
import { MAX_CONSIST, nextOfficeTier, officeProfile } from '../engine/content.ts';
|
||||
import type { Hand, TrackGeometry } from '../engine/content.ts';
|
||||
import type { GameEvent } from '../engine/events.ts';
|
||||
import type { Intent } from '../engine/intents.ts';
|
||||
@@ -78,8 +78,42 @@ function because(reason: string, intent: Intent): Intent {
|
||||
return intent;
|
||||
}
|
||||
|
||||
export const developerBot: BotPolicy = {
|
||||
name: 'developer',
|
||||
/**
|
||||
* KNOBS FOR A/B MEASUREMENT, and nothing else.
|
||||
*
|
||||
* A heuristic change has to be measured against the bot it replaces, over the SAME deals — and
|
||||
* editing the bot between runs makes that impossible to do honestly, because the two sides of the
|
||||
* comparison never exist at once. Every flag here is off by default, so `makeDeveloperBot({})` is
|
||||
* byte-identical to the bot that came before this existed.
|
||||
*
|
||||
* TEMPORARY BY CONSTRUCTION. When a flag measures well it becomes the default and the flag is
|
||||
* deleted in the same commit; when it measures badly it is deleted with its finding recorded in the
|
||||
* changelog. What must not happen is a bot that accumulates switches nobody can account for — a
|
||||
* heuristic with no measurement attached is exactly what this machinery exists to prevent.
|
||||
*/
|
||||
export type BotTweaks = {
|
||||
/**
|
||||
* Refuse to schedule a train the Office has no room for: cap committed trains (timetable slots
|
||||
* filled, plus queued Extras) at `adTracks + trainCapSlack`.
|
||||
*
|
||||
* `undefined` means no cap — today's behaviour, which plays every train card on sight. A Whistle
|
||||
* Post has ONE A/D track, and a second train standing there makes every later arrival an
|
||||
* automatic collision (Gap 2d).
|
||||
*/
|
||||
trainCapSlack?: number;
|
||||
};
|
||||
|
||||
/** The bot as it plays today. Every knob off. */
|
||||
export const developerBot: BotPolicy = makeDeveloperBot({});
|
||||
|
||||
/** A variant, for measuring one change at a time against the bot above. */
|
||||
export function makeDeveloperBot(tweaks: BotTweaks): BotPolicy {
|
||||
const suffix = Object.entries(tweaks)
|
||||
.filter(([, v]) => v !== undefined)
|
||||
.map(([k, v]) => `${k}=${v}`)
|
||||
.join(',');
|
||||
return {
|
||||
name: suffix ? `developer+${suffix}` : 'developer',
|
||||
|
||||
choose(s, player, options) {
|
||||
lastReason = 'no specific reason — first legal option';
|
||||
@@ -133,13 +167,29 @@ export const developerBot: BotPolicy = {
|
||||
|
||||
// --- Local Operations.
|
||||
if (s.clock.phase === 'localOps') {
|
||||
if (s.turn.option === null) return chooseLocalOption(s, player, options);
|
||||
return followThrough(s, player, options);
|
||||
/**
|
||||
* A CAP HAS TO BE APPLIED TO THE OPTIONS, NOT TO ONE BRANCH.
|
||||
*
|
||||
* The first attempt gated the two branches that exist to play a train card, measured exactly
|
||||
* zero difference over 400 paired seeds, and was right to: `followThrough` ends with a generic
|
||||
* "play what is in hand" fallback that played the train card anyway. Removing the option
|
||||
* itself is the only way to be sure no path reaches it.
|
||||
*
|
||||
* Never to an empty list — a bot with nothing legal to choose is a crash, and §6.2 can force
|
||||
* a play when the hand is over its limit. If filtering leaves nothing, the cap yields.
|
||||
*/
|
||||
const held = trainWouldOverfillTheOffice(s, player, tweaks)
|
||||
? options.filter((i) => !(i.type === 'card.play' && isTrainCard(s, i.cardId)))
|
||||
: options;
|
||||
const usable = held.length > 0 ? held : options;
|
||||
if (s.turn.option === null) return chooseLocalOption(s, player, usable, tweaks);
|
||||
return followThrough(s, player, usable, tweaks);
|
||||
}
|
||||
|
||||
return options[0]!;
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The one decision that matters: which of §6's three things to spend this Stage on.
|
||||
@@ -150,7 +200,37 @@ export const developerBot: BotPolicy = {
|
||||
* and spent the rest of the game stuffing green boxes whose loads could never complete for want of
|
||||
* a car to load them onto.
|
||||
*/
|
||||
function chooseLocalOption(s: GameState, player: PlayerIndex, options: Intent[]): Intent {
|
||||
/**
|
||||
* Trains this player has already committed to running: slots filled on the timetable, plus Extras
|
||||
* played and waiting for a Crew Tray.
|
||||
*
|
||||
* A Second Section is deliberately NOT counted. Its whole purpose is to put a following train into
|
||||
* an occupied Subdivision and force the Superintendent's ruling (Q9) — capping it would be capping
|
||||
* the card's reason to exist.
|
||||
*/
|
||||
function committedTrains(s: GameState): number {
|
||||
return s.timetable.filter((t) => t !== null).length + s.pendingExtras.length;
|
||||
}
|
||||
|
||||
/**
|
||||
* Would scheduling another train exceed what the Office can physically hold?
|
||||
*
|
||||
* §7 lets you play as many train cards as you draw, and a train that arrives with nowhere to stand
|
||||
* is an automatic collision (Gap 2d) — so the two rules together make a train card actively harmful
|
||||
* once the A/D tracks are spoken for. Off unless `trainCapSlack` is set.
|
||||
*/
|
||||
function trainWouldOverfillTheOffice(s: GameState, player: PlayerIndex, tweaks: BotTweaks): boolean {
|
||||
if (tweaks.trainCapSlack === undefined) return false;
|
||||
const cap = officeProfile(areaOf(s, player).tier).adTracks + tweaks.trainCapSlack;
|
||||
return committedTrains(s) >= cap;
|
||||
}
|
||||
|
||||
function chooseLocalOption(
|
||||
s: GameState,
|
||||
player: PlayerIndex,
|
||||
options: Intent[],
|
||||
tweaks: BotTweaks,
|
||||
): Intent {
|
||||
const choices = options.filter(
|
||||
(i): i is Extract<Intent, { type: 'localOps.choose' }> => i.type === 'localOps.choose',
|
||||
);
|
||||
@@ -167,7 +247,11 @@ function chooseLocalOption(s: GameState, player: PlayerIndex, options: Intent[])
|
||||
return because('an Office upgrade is in hand — more A/D track means fewer collisions', can('draw')!);
|
||||
}
|
||||
|
||||
if (can('draw') && hand.some((id) => isTrainCard(s, id))) {
|
||||
if (
|
||||
can('draw') &&
|
||||
hand.some((id) => isTrainCard(s, id)) &&
|
||||
!trainWouldOverfillTheOffice(s, player, tweaks)
|
||||
) {
|
||||
return because('a train card is in hand and its value compounds every Day', can('draw')!);
|
||||
}
|
||||
|
||||
@@ -933,7 +1017,12 @@ function moveTowardOffice(s: GameState, player: PlayerIndex, options: Intent[]):
|
||||
}
|
||||
|
||||
/** Once an option is chosen, work it to a sensible conclusion. */
|
||||
function followThrough(s: GameState, player: PlayerIndex, options: Intent[]): Intent {
|
||||
function followThrough(
|
||||
s: GameState,
|
||||
player: PlayerIndex,
|
||||
options: Intent[],
|
||||
tweaks: BotTweaks,
|
||||
): Intent {
|
||||
switch (s.turn.option) {
|
||||
case 'draw': {
|
||||
// Draw before playing — otherwise the hand empties and never refills.
|
||||
@@ -962,8 +1051,12 @@ function followThrough(s: GameState, player: PlayerIndex, options: Intent[]): In
|
||||
);
|
||||
if (upgrade) return because('upgrade the Office — a Whistle Post has ONE A/D track and a second arrival collides', upgrade);
|
||||
|
||||
// Train cards next — they take no placement and their value compounds every Day.
|
||||
const train = options.find(
|
||||
// Train cards next — they take no placement and their value compounds every Day. Unless the
|
||||
// Office cannot hold another one, in which case the card is HELD rather than played: an
|
||||
// arrival with no free A/D track is an automatic collision, and it repeats every Day.
|
||||
const train = trainWouldOverfillTheOffice(s, player, tweaks)
|
||||
? undefined
|
||||
: options.find(
|
||||
(i) => i.type === 'card.play' && i.placement === undefined && isTrainCard(s, i.cardId),
|
||||
);
|
||||
if (train) return because('a scheduled train runs every Day thereafter — the only card whose value compounds', train);
|
||||
@@ -1504,12 +1597,23 @@ export type PlayOutcome = {
|
||||
*/
|
||||
export type PlayObserver = (event: GameEvent, state: GameState) => void;
|
||||
|
||||
/**
|
||||
* Called once per DECISION, immediately before the policy is asked to choose.
|
||||
*
|
||||
* Separate from `PlayObserver` because the two sample different things. An event is something that
|
||||
* happened; a turn hook sees the position as the bot sees it, which is the only place to ask a
|
||||
* question of the form "was anything ready to be done here?" — a condition rather than an
|
||||
* occurrence. Sampling those off events instead would count them once per event in the batch.
|
||||
*/
|
||||
export type TurnObserver = (state: GameState) => void;
|
||||
|
||||
export function playGame(
|
||||
s: GameState,
|
||||
policy: BotPolicy,
|
||||
pumpFn: (s: GameState) => GameEvent[],
|
||||
maxTurns = 50_000,
|
||||
observe?: PlayObserver,
|
||||
onTurn?: TurnObserver,
|
||||
): PlayOutcome {
|
||||
const events: GameEvent[] = [];
|
||||
const intents: Intent['type'][] = [];
|
||||
@@ -1545,6 +1649,7 @@ export function playGame(
|
||||
const options = legalActions(s, actor);
|
||||
if (options.length === 0) break;
|
||||
|
||||
onTurn?.(s);
|
||||
const chosen = policy.choose(s, actor, options);
|
||||
intents.push(chosen.type);
|
||||
const r = applyIntent(s, actor, chosen);
|
||||
|
||||
@@ -0,0 +1,233 @@
|
||||
/**
|
||||
* Component 18b — paired A/B measurement for bot heuristics.
|
||||
*
|
||||
* Dev-side only. Runs the current bot and one tweaked variant over the SAME deals and reports the
|
||||
* per-seed difference.
|
||||
*
|
||||
* WHY PAIRED, AND WHY IT IS NOT OPTIONAL. Revenue has a standard deviation of about 9 across games,
|
||||
* so two 100-game runs of the identical bot can differ by a point through nothing but the deal. Every
|
||||
* single-change claim in the changelog before the Interlocking work is inside that noise. Giving both
|
||||
* policies the same seed removes the deal from the comparison entirely: what is left is the change.
|
||||
* Measured here, the paired difference has σ ≈ 5.3 against ≈ 9 unpaired, and at 1600 seeds the
|
||||
* standard error is ±0.13 — so a +0.5 heuristic is resolvable in under two minutes, where unpaired it
|
||||
* would need tens of thousands of games.
|
||||
*
|
||||
* WHY THE SPLIT IS PRINTED. A mean carried by a skewed tail is a different claim from a mean carried
|
||||
* by broad improvement, and only the better/worse/identical counts tell them apart. The train cap is
|
||||
* the case in point: +0.60 overall, but 130 seeds better, 145 worse and 1325 unchanged — it removes a
|
||||
* rare catastrophe rather than making the bot play better.
|
||||
*
|
||||
* WHY THE FUNNEL IS PRINTED. Revenue can rise because a channel started working or because the bot
|
||||
* abandoned an expensive channel for a cheap one, and the two are identical in a single number.
|
||||
*
|
||||
* Run with:
|
||||
* node src/sim/compare.ts 1600 trainCapSlack=1
|
||||
* node src/sim/compare.ts 400 trainCapSlack=0 --length short
|
||||
*/
|
||||
|
||||
import { pump } from '../engine/advance.ts';
|
||||
import { createGame } from '../engine/setup.ts';
|
||||
import type { GameLength } from '../engine/content.ts';
|
||||
import type { GameConfig, GameMode } from '../engine/state.ts';
|
||||
import type { BotPolicy, BotTweaks } from './bot.ts';
|
||||
import { developerBot, makeDeveloperBot, playGame } from './bot.ts';
|
||||
import type { Funnel, GameStats } from './stats.ts';
|
||||
import { funnelReport, makeFunnelProbe, summarize } from './stats.ts';
|
||||
|
||||
export type PairedResult = {
|
||||
seeds: number;
|
||||
baseline: string;
|
||||
variant: string;
|
||||
/** Per-seed revenue difference, variant minus baseline. */
|
||||
deltas: { seed: number; delta: number }[];
|
||||
mean: number;
|
||||
/** Standard error of the mean difference — the number that decides whether this is real. */
|
||||
stderr: number;
|
||||
sd: number;
|
||||
t: number;
|
||||
better: number;
|
||||
worse: number;
|
||||
identical: number;
|
||||
baseStats: GameStats[];
|
||||
variantStats: GameStats[];
|
||||
};
|
||||
|
||||
const SOLO = (length: GameLength, mode: GameMode): GameConfig => ({
|
||||
mode,
|
||||
victory: 'highestAfterDays',
|
||||
length,
|
||||
optionalRules: {
|
||||
reducedVisibility: false,
|
||||
sisterTrains: false,
|
||||
employeeRotation: false,
|
||||
emergencyToolbox: false,
|
||||
},
|
||||
});
|
||||
|
||||
/**
|
||||
* One game, one seed, one policy.
|
||||
*
|
||||
* The seed stride matches `harness.ts` exactly, so a compare run and a harness run of the same size
|
||||
* are talking about the same games — otherwise two numbers that ought to agree would not, for a
|
||||
* reason nobody would find.
|
||||
*/
|
||||
function runOne(policy: BotPolicy, seed: number, length: GameLength, mode: GameMode): GameStats {
|
||||
const s = createGame({ id: `cmp-${seed}`, seed, config: SOLO(length, mode), playerNames: ['bot'] });
|
||||
const probe = makeFunnelProbe(0);
|
||||
const r = playGame(s, policy, pump, 50_000, probe.onEvent, probe.onTurn);
|
||||
return summarize(seed, r.events, r.intents, s, probe.funnel);
|
||||
}
|
||||
|
||||
export function compare(
|
||||
tweaks: BotTweaks,
|
||||
games: number,
|
||||
length: GameLength = 'standard',
|
||||
mode: GameMode = 'solitaire',
|
||||
): PairedResult {
|
||||
const variantPolicy = makeDeveloperBot(tweaks);
|
||||
const baseStats: GameStats[] = [];
|
||||
const variantStats: GameStats[] = [];
|
||||
const deltas: { seed: number; delta: number }[] = [];
|
||||
|
||||
for (let i = 0; i < games; i++) {
|
||||
const seed = 1000 + i * 7919; // the same prime stride the harness deals
|
||||
const a = runOne(developerBot, seed, length, mode);
|
||||
const b = runOne(variantPolicy, seed, length, mode);
|
||||
baseStats.push(a);
|
||||
variantStats.push(b);
|
||||
deltas.push({ seed, delta: b.revenue.net - a.revenue.net });
|
||||
}
|
||||
|
||||
const d = deltas.map((x) => x.delta);
|
||||
const mean = d.reduce((p, c) => p + c, 0) / Math.max(1, d.length);
|
||||
const sd =
|
||||
d.length > 1
|
||||
? Math.sqrt(d.reduce((p, c) => p + (c - mean) ** 2, 0) / (d.length - 1))
|
||||
: 0;
|
||||
const stderr = d.length > 0 ? sd / Math.sqrt(d.length) : 0;
|
||||
|
||||
return {
|
||||
seeds: games,
|
||||
baseline: developerBot.name,
|
||||
variant: variantPolicy.name,
|
||||
deltas,
|
||||
mean,
|
||||
sd,
|
||||
stderr,
|
||||
t: stderr === 0 ? 0 : mean / stderr,
|
||||
better: d.filter((x) => x > 0).length,
|
||||
worse: d.filter((x) => x < 0).length,
|
||||
identical: d.filter((x) => x === 0).length,
|
||||
baseStats,
|
||||
variantStats,
|
||||
};
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Reporting
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const mean = (xs: number[]): number => (xs.length ? xs.reduce((p, c) => p + c, 0) / xs.length : 0);
|
||||
|
||||
/**
|
||||
* The verdict, stated in the same terms every time.
|
||||
*
|
||||
* The threshold is t ≥ 3 rather than the conventional 2. This is a measurement taken repeatedly on
|
||||
* the same system while looking for something that works, so the conventional bar would have us keep
|
||||
* roughly one bad heuristic in twenty; and the cost of a wrong keep is not a wrong paper, it is a bot
|
||||
* that quietly plays worse and takes every later measurement with it.
|
||||
*/
|
||||
function verdict(t: number): string {
|
||||
const a = Math.abs(t);
|
||||
if (a >= 3) return t > 0 ? 'KEEP — clears the bar (t ≥ 3)' : 'REJECT — significantly worse';
|
||||
if (a >= 2) return 'NOT PROVEN — suggestive, run more seeds before believing it';
|
||||
return 'NO EFFECT MEASURED — inside the noise';
|
||||
}
|
||||
|
||||
export function formatPaired(r: PairedResult): string {
|
||||
const out: string[] = [];
|
||||
const num = (x: number, w = 8, dp = 2): string => x.toFixed(dp).padStart(w);
|
||||
|
||||
out.push(`\n=== paired comparison · ${r.seeds} seeds · same deal to both ===\n`);
|
||||
out.push(` baseline ${r.baseline}`);
|
||||
out.push(` variant ${r.variant}\n`);
|
||||
|
||||
const rows: [string, (g: GameStats) => number][] = [
|
||||
['revenue', (g) => g.revenue.net],
|
||||
['collisions', (g) => g.collisions],
|
||||
['trains scheduled', (g) => g.trains.scheduled],
|
||||
['arrivals', (g) => g.trains.arrived],
|
||||
['passengers on/off', (g) => g.revenue.passengerBoard + g.revenue.passengerDetrain],
|
||||
['freight loads+unloads', (g) => g.revenue.freightLoad + g.revenue.freightUnload],
|
||||
['cards played', (g) => g.development.cardsPlayed],
|
||||
['final grid size', (g) => g.development.gridSize],
|
||||
];
|
||||
out.push(' baseline variant delta');
|
||||
for (const [label, pick] of rows) {
|
||||
const a = mean(r.baseStats.map(pick));
|
||||
const b = mean(r.variantStats.map(pick));
|
||||
out.push(` ${label.padEnd(22)}${num(a)} ${num(b)} ${num(b - a)}`);
|
||||
}
|
||||
|
||||
out.push(
|
||||
`\n REVENUE DELTA ${r.mean >= 0 ? '+' : ''}${r.mean.toFixed(2)} ± ${r.stderr.toFixed(2)}` +
|
||||
` (t = ${r.t.toFixed(2)}, σ of the paired difference ${r.sd.toFixed(2)})`,
|
||||
);
|
||||
out.push(` ${verdict(r.t)}`);
|
||||
out.push(
|
||||
`\n seeds better ${r.better} · worse ${r.worse} · identical ${r.identical}` +
|
||||
` — ${((r.identical / Math.max(1, r.seeds)) * 100).toFixed(0)}% of games are untouched by this change`,
|
||||
);
|
||||
|
||||
// The tails, BY SEED, so a loss can be replayed rather than averaged away.
|
||||
const sorted = [...r.deltas].sort((a, b) => a.delta - b.delta);
|
||||
const show = (xs: { seed: number; delta: number }[]): string =>
|
||||
xs.map((x) => `${x.seed} (${x.delta > 0 ? '+' : ''}${x.delta})`).join(', ');
|
||||
out.push(` worst seeds: ${show(sorted.slice(0, 3))}`);
|
||||
out.push(` best seeds: ${show(sorted.slice(-3).reverse())}`);
|
||||
|
||||
// A change that raises revenue by abandoning a channel has to be visible as that.
|
||||
out.push('\n --- baseline funnel ---');
|
||||
out.push(funnelReport(r.baseStats));
|
||||
out.push('\n --- variant funnel ---');
|
||||
out.push(funnelReport(r.variantStats));
|
||||
return out.join('\n');
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// CLI
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/** `trainCapSlack=1` -> `{ trainCapSlack: 1 }`. Unknown names are refused rather than ignored. */
|
||||
export function parseTweaks(args: string[]): BotTweaks {
|
||||
const known = new Set(['trainCapSlack']);
|
||||
const tweaks: Record<string, number> = {};
|
||||
for (const a of args) {
|
||||
const m = /^([A-Za-z]\w*)=(-?\d+(?:\.\d+)?)$/.exec(a);
|
||||
if (!m) continue;
|
||||
if (!known.has(m[1]!)) {
|
||||
throw new Error(`unknown tweak "${m[1]}" — known: ${[...known].join(', ')}`);
|
||||
}
|
||||
tweaks[m[1]!] = Number(m[2]);
|
||||
}
|
||||
return tweaks as BotTweaks;
|
||||
}
|
||||
|
||||
const isMain = process.argv[1]?.endsWith('compare.ts') ?? false;
|
||||
|
||||
if (isMain) {
|
||||
const args = process.argv.slice(2);
|
||||
const games = Number(args.find((a) => /^\d+$/.test(a)) ?? 400);
|
||||
const lengthIdx = args.indexOf('--length');
|
||||
const length = (lengthIdx >= 0 ? args[lengthIdx + 1] : 'standard') as GameLength;
|
||||
const tweaks = parseTweaks(args);
|
||||
|
||||
if (Object.keys(tweaks).length === 0) {
|
||||
console.error('nothing to compare — pass at least one tweak, e.g. trainCapSlack=1');
|
||||
process.exitCode = 1;
|
||||
} else {
|
||||
console.log(formatPaired(compare(tweaks, games, length)));
|
||||
}
|
||||
}
|
||||
|
||||
export type { Funnel };
|
||||
+6
-3
@@ -19,7 +19,7 @@ import type { GameConfig, GameMode } from '../engine/state.ts';
|
||||
import type { BotPolicy, PlayOutcome } from './bot.ts';
|
||||
import { developerBot, playGame, randomBot } from './bot.ts';
|
||||
import type { GameStats } from './stats.ts';
|
||||
import { formatAggregate, summarize } from './stats.ts';
|
||||
import { formatAggregate, makeFunnelProbe, summarize } from './stats.ts';
|
||||
|
||||
export type SimOptions = {
|
||||
games: number;
|
||||
@@ -86,9 +86,12 @@ export function simulate(opts: SimOptions): SimReport {
|
||||
config: configFor(opts.mode, opts.length),
|
||||
playerNames: opts.players,
|
||||
});
|
||||
const r = playGame(s, opts.policy, pump);
|
||||
// The probe watches the game as it is played: the funnel gates are conditions at a moment, not
|
||||
// events, so they cannot be recovered from the log afterwards.
|
||||
const probe = makeFunnelProbe(0);
|
||||
const r = playGame(s, opts.policy, pump, 50_000, probe.onEvent, probe.onTurn);
|
||||
results.push(r);
|
||||
stats.push(summarize(seed, r.events, r.intents, s));
|
||||
stats.push(summarize(seed, r.events, r.intents, s, probe.funnel));
|
||||
}
|
||||
|
||||
const finished = results.filter((r) => r.finished);
|
||||
|
||||
@@ -18,11 +18,149 @@
|
||||
import type { GameEvent } from '../engine/events.ts';
|
||||
import type { Intent } from '../engine/intents.ts';
|
||||
import type { GameState } from '../engine/state.ts';
|
||||
import { areaOf, facilityCarTypes } from '../engine/apply.ts';
|
||||
import { officeProfile } from '../engine/content.ts';
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Per-game summary
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/**
|
||||
* THE FUNNEL — how far each opportunity got before it stopped.
|
||||
*
|
||||
* The totals above say what happened; this says what ALMOST happened, which is what a change to the
|
||||
* bot has to move. Revenue can rise because a channel started working or because a channel was
|
||||
* abandoned in favour of a cheaper one, and the two look identical in a single number: capping the
|
||||
* trains a bot schedules raises revenue while cutting freight events nearly in half.
|
||||
*
|
||||
* SAMPLED FROM LIVE STATE, not from events, because the interesting gates are conditions rather than
|
||||
* occurrences — "was a green box stocked while a car was spotted" is not a thing that happens, it is
|
||||
* a thing that is true or false at a moment. `playGame` calls the probe once per decision and the
|
||||
* probe dedupes per phase, so these count PHASES, not turns.
|
||||
*/
|
||||
export type Funnel = {
|
||||
/** Trains reaching an Office, and what they were carrying when they got there. */
|
||||
arrivals: number;
|
||||
arrivalsAtPassengerOffice: number;
|
||||
/** Boarding needs an empty coach aboard (§9.2); detraining needs a loaded one. */
|
||||
arrivalsWithEmptyCoach: number;
|
||||
arrivalsWithLoadedCoach: number;
|
||||
/** A car of a type some industry in the district actually wants, loaded or empty. */
|
||||
arrivalsWithWantedCar: number;
|
||||
|
||||
/** Cargo phases, and how many had each half of what a load needs. */
|
||||
cargoPhases: number;
|
||||
cargoWithFacility: number;
|
||||
cargoGreenStocked: number;
|
||||
cargoCarSpotted: number;
|
||||
cargoLaborerFree: number;
|
||||
/** All three at ONE industry — the only combination that can actually move a load. */
|
||||
cargoReady: number;
|
||||
|
||||
/**
|
||||
* Decisions taken with a train's engine buried mid-consist, and how many of those had a legal way
|
||||
* to dig it out. §8.2 will not let such a train leave, and Rolling Stock may not be set out at the
|
||||
* Office — so the gap between these two numbers is the trap the bot cannot escape from.
|
||||
*/
|
||||
buriedTurns: number;
|
||||
buriedWithDigAvailable: number;
|
||||
};
|
||||
|
||||
function emptyFunnel(): Funnel {
|
||||
return {
|
||||
arrivals: 0, arrivalsAtPassengerOffice: 0, arrivalsWithEmptyCoach: 0,
|
||||
arrivalsWithLoadedCoach: 0, arrivalsWithWantedCar: 0,
|
||||
cargoPhases: 0, cargoWithFacility: 0, cargoGreenStocked: 0, cargoCarSpotted: 0,
|
||||
cargoLaborerFree: 0, cargoReady: 0,
|
||||
buriedTurns: 0, buriedWithDigAvailable: 0,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* A probe to hand to `playGame`: an event hook, a per-decision hook, and the totals they fill.
|
||||
*
|
||||
* One per game. Kept here beside the statistics it produces rather than in the harness, so the
|
||||
* definition of "a Cargo phase that was ready" lives in exactly one place.
|
||||
*/
|
||||
export function makeFunnelProbe(player = 0): {
|
||||
funnel: Funnel;
|
||||
onEvent: (e: GameEvent, s: GameState) => void;
|
||||
onTurn: (s: GameState) => void;
|
||||
} {
|
||||
const funnel = emptyFunnel();
|
||||
// Phases are sampled once, on the first decision taken in them.
|
||||
let lastPhaseKey = '';
|
||||
|
||||
const onEvent = (e: GameEvent, s: GameState): void => {
|
||||
if (e.type !== 'trainArrived') return;
|
||||
funnel.arrivals += 1;
|
||||
const area = areaOf(s, player);
|
||||
if (officeProfile(area.tier).isPassengerFacility) funnel.arrivalsAtPassengerOffice += 1;
|
||||
if (e.consist.some((c) => c.type === 'coach' && !c.loaded)) funnel.arrivalsWithEmptyCoach += 1;
|
||||
if (e.consist.some((c) => c.type === 'coach' && c.loaded)) funnel.arrivalsWithLoadedCoach += 1;
|
||||
|
||||
for (const card of area.grid.values()) {
|
||||
const f = card.facility;
|
||||
if (!f || f.kind !== 'freight') continue;
|
||||
const want = facilityCarTypes(f);
|
||||
const useful = e.consist.some(
|
||||
(c) =>
|
||||
want.includes(c.type) &&
|
||||
((f.allows.outbound && !c.loaded) || (f.allows.inbound && c.loaded)),
|
||||
);
|
||||
if (useful) {
|
||||
funnel.arrivalsWithWantedCar += 1;
|
||||
break;
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
const onTurn = (s: GameState): void => {
|
||||
const area = areaOf(s, player);
|
||||
|
||||
// A buried engine is a per-decision condition: every turn it persists is a turn the train is
|
||||
// stuck, so this counts turns rather than trains.
|
||||
for (const tray of s.trays.values()) {
|
||||
if (tray.position.at !== 'grid' || tray.position.owner !== player) continue;
|
||||
if (tray.engineAt <= 0 || tray.engineAt >= tray.consist.length) continue;
|
||||
funnel.buriedTurns += 1;
|
||||
// Cars may never be set out at the Office (§A.4), which is where this almost always happens.
|
||||
const atOffice =
|
||||
tray.position.coord.row === area.officeCoord.row &&
|
||||
tray.position.coord.col === area.officeCoord.col;
|
||||
if (!atOffice) funnel.buriedWithDigAvailable += 1;
|
||||
break;
|
||||
}
|
||||
|
||||
if (s.clock.phase !== 'loadUnload') return;
|
||||
const key = `${s.clock.day}/${s.clock.stage}`;
|
||||
if (key === lastPhaseKey) return;
|
||||
lastPhaseKey = key;
|
||||
|
||||
funnel.cargoPhases += 1;
|
||||
let facility = false, green = false, car = false, laborer = false, ready = false;
|
||||
for (const card of area.grid.values()) {
|
||||
const f = card.facility;
|
||||
if (!f || f.kind !== 'freight') continue;
|
||||
facility = true;
|
||||
const g = f.outboundBox.length > 0;
|
||||
const c = f.industryTrack.cars.length > 0;
|
||||
const l = f.laborers - f.usedThisStage.laborers > 0;
|
||||
green ||= g;
|
||||
car ||= c;
|
||||
laborer ||= l;
|
||||
ready ||= g && c && l;
|
||||
}
|
||||
if (facility) funnel.cargoWithFacility += 1;
|
||||
if (green) funnel.cargoGreenStocked += 1;
|
||||
if (car) funnel.cargoCarSpotted += 1;
|
||||
if (laborer) funnel.cargoLaborerFree += 1;
|
||||
if (ready) funnel.cargoReady += 1;
|
||||
};
|
||||
|
||||
return { funnel, onEvent, onTurn };
|
||||
}
|
||||
|
||||
export type RevenueBreakdown = {
|
||||
freightLoad: number;
|
||||
freightUnload: number;
|
||||
@@ -76,6 +214,11 @@ export type GameStats = {
|
||||
collisions: number;
|
||||
eventCounts: Record<string, number>;
|
||||
intentCounts: Record<string, number>;
|
||||
/**
|
||||
* How far each opportunity got. Optional because it needs a live probe during the game, which a
|
||||
* caller reconstructing statistics from a saved event log cannot supply.
|
||||
*/
|
||||
funnel?: Funnel;
|
||||
};
|
||||
|
||||
function inc(map: Record<string, number>, key: string, by = 1): void {
|
||||
@@ -87,6 +230,7 @@ export function summarize(
|
||||
events: GameEvent[],
|
||||
intents: Intent['type'][],
|
||||
final: GameState,
|
||||
funnel?: Funnel,
|
||||
): GameStats {
|
||||
const eventCounts: Record<string, number> = {};
|
||||
const intentCounts: Record<string, number> = {};
|
||||
@@ -211,6 +355,7 @@ export function summarize(
|
||||
|
||||
eventCounts,
|
||||
intentCounts,
|
||||
...(funnel ? { funnel } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -414,6 +559,44 @@ export function formatGameStats(g: GameStats): string {
|
||||
return out.join('\n');
|
||||
}
|
||||
|
||||
/**
|
||||
* THE FUNNEL, AS PERCENTAGES OF THE THING THEY GATE.
|
||||
*
|
||||
* Deliberately not per-game averages: "3.6 arrivals carried an empty coach" says nothing without the
|
||||
* 7.3 arrivals it is out of, and the whole point is to see WHERE an opportunity stops. Read the
|
||||
* arrivals block as "of every train that reached an Office" and the Cargo block as "of every Cargo
|
||||
* phase". Empty when nothing was probed.
|
||||
*/
|
||||
export function funnelReport(all: GameStats[]): string {
|
||||
const fs = all.map((g) => g.funnel).filter((f): f is Funnel => f !== undefined);
|
||||
if (fs.length === 0) return '';
|
||||
const sum = (pick: (f: Funnel) => number): number => fs.reduce((n, f) => n + pick(f), 0);
|
||||
const pct = (n: number, d: number): string =>
|
||||
d === 0 ? ' —' : `${((n / d) * 100).toFixed(0).padStart(3)}%`;
|
||||
|
||||
const arr = sum((f) => f.arrivals);
|
||||
const cargo = sum((f) => f.cargoPhases);
|
||||
const out: string[] = [];
|
||||
|
||||
out.push(`\n PASSENGER FUNNEL — ${(arr / fs.length).toFixed(1)} arrivals a game`);
|
||||
out.push(` at a Passenger Facility (Depot+) ${pct(sum((f) => f.arrivalsAtPassengerOffice), arr)}`);
|
||||
out.push(` carrying an EMPTY coach (boarding) ${pct(sum((f) => f.arrivalsWithEmptyCoach), arr)}`);
|
||||
out.push(` carrying a LOADED coach (detrain) ${pct(sum((f) => f.arrivalsWithLoadedCoach), arr)}`);
|
||||
out.push(` carrying a car an industry wants ${pct(sum((f) => f.arrivalsWithWantedCar), arr)}`);
|
||||
|
||||
out.push(`\n FREIGHT FUNNEL — ${(cargo / fs.length).toFixed(0)} Cargo phases a game`);
|
||||
out.push(` a freight facility was built ${pct(sum((f) => f.cargoWithFacility), cargo)}`);
|
||||
out.push(` a green box was stocked ${pct(sum((f) => f.cargoGreenStocked), cargo)}`);
|
||||
out.push(` a car was spotted on an industry ${pct(sum((f) => f.cargoCarSpotted), cargo)}`);
|
||||
out.push(` a Laborer was free ${pct(sum((f) => f.cargoLaborerFree), cargo)}`);
|
||||
out.push(` ALL THREE at one industry ${pct(sum((f) => f.cargoReady), cargo)}`);
|
||||
|
||||
const buried = sum((f) => f.buriedTurns);
|
||||
out.push(`\n STUCK — ${(buried / fs.length).toFixed(1)} decisions a game with the engine buried mid-train`);
|
||||
out.push(` ... of those, able to set the nose cars out ${pct(sum((f) => f.buriedWithDigAvailable), buried)}`);
|
||||
return out.join('\n');
|
||||
}
|
||||
|
||||
export function formatAggregate(all: GameStats[]): string {
|
||||
const out: string[] = [];
|
||||
out.push(`\n=== end-of-game statistics · ${all.length} games ===\n`);
|
||||
@@ -449,6 +632,8 @@ export function formatAggregate(all: GameStats[]): string {
|
||||
out.push(` ${t.padEnd(14)} ${n} (${((n / all.length) * 100).toFixed(0)}%)`);
|
||||
}
|
||||
|
||||
out.push(funnelReport(all));
|
||||
|
||||
out.push('\n ACTION MIX (mean per game)');
|
||||
out.push(` switch moves ${num(meanOf(all, (g) => g.actions.moves))}`);
|
||||
out.push(` cars dropped ${num(meanOf(all, (g) => g.actions.drops))}`);
|
||||
|
||||
@@ -0,0 +1,219 @@
|
||||
/**
|
||||
* The measurement machinery itself.
|
||||
*
|
||||
* These tests exist because a measurement tool that is quietly wrong is worse than no tool: it does
|
||||
* not fail, it just produces numbers that decide what the bot becomes. Each one guards a property
|
||||
* the paired method depends on.
|
||||
*/
|
||||
|
||||
import { describe, it } from 'node:test';
|
||||
import assert from 'node:assert/strict';
|
||||
|
||||
import { pump } from '../src/engine/advance.ts';
|
||||
import { createGame } from '../src/engine/setup.ts';
|
||||
import { adTrackCount } from '../src/engine/state.ts';
|
||||
import type { GameConfig, GameState } from '../src/engine/state.ts';
|
||||
import type { BotPolicy } from '../src/sim/bot.ts';
|
||||
import { developerBot, makeDeveloperBot, playGame } from '../src/sim/bot.ts';
|
||||
import { compare, parseTweaks } from '../src/sim/compare.ts';
|
||||
import { makeFunnelProbe, summarize } from '../src/sim/stats.ts';
|
||||
|
||||
const config: GameConfig = {
|
||||
mode: 'solitaire',
|
||||
victory: 'highestAfterDays',
|
||||
length: 'standard',
|
||||
optionalRules: {
|
||||
reducedVisibility: false,
|
||||
sisterTrains: false,
|
||||
employeeRotation: false,
|
||||
emergencyToolbox: false,
|
||||
},
|
||||
};
|
||||
|
||||
const game = (seed: number): GameState =>
|
||||
createGame({ id: `h${seed}`, seed, config, playerNames: ['bot'] });
|
||||
|
||||
/** One game, reduced to a string: every event type in order, plus the final revenue. */
|
||||
function fingerprint(policy: BotPolicy, seed: number): string {
|
||||
const s = game(seed);
|
||||
const r = playGame(s, policy, pump);
|
||||
return `${r.revenue[0]}|${r.events.map((e) => e.type).join(',')}`;
|
||||
}
|
||||
|
||||
describe('the paired method rests on determinism', () => {
|
||||
it('plays the identical game twice from one seed', () => {
|
||||
/**
|
||||
* THE PROPERTY EVERYTHING ELSE DEPENDS ON. A paired comparison subtracts two games that share a
|
||||
* deal; if the same policy on the same seed can diverge, the difference is measuring the engine's
|
||||
* own noise and every heuristic verdict is worthless.
|
||||
*
|
||||
* It is also the engine's central claim — pure, seeded, no ambient randomness — and the failure
|
||||
* mode is silent: one `Math.random`, one Map iterated in a different order, and this is the only
|
||||
* thing that would notice.
|
||||
*/
|
||||
for (const seed of [1000, 8919, 775569289]) {
|
||||
assert.equal(fingerprint(developerBot, seed), fingerprint(developerBot, seed), `seed ${seed} diverged`);
|
||||
}
|
||||
});
|
||||
|
||||
it('gives the untweaked variant the identical game too', () => {
|
||||
// `developerBot` IS `makeDeveloperBot({})`, and the refactor that introduced tweaks must not
|
||||
// have changed a single decision. If this fails, every number measured before it is unusable.
|
||||
for (const seed of [1000, 8919, 775569289]) {
|
||||
assert.equal(
|
||||
fingerprint(developerBot, seed),
|
||||
fingerprint(makeDeveloperBot({}), seed),
|
||||
`seed ${seed}: the untweaked variant plays a different game`,
|
||||
);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('a tweak has to actually do something', () => {
|
||||
it('actually holds the trains it says it holds', () => {
|
||||
/**
|
||||
* REGRESSION, and the reason this file exists. The first `trainCapSlack` gated the two branches
|
||||
* whose comments say they exist to play a train card — and measured **exactly zero difference
|
||||
* over 400 paired seeds**, because `followThrough` ends with a generic "play what is in hand"
|
||||
* fallback that played the card anyway.
|
||||
*
|
||||
* ASSERT THE CONTRACT, NOT MERELY A DIFFERENCE. A first version of this test only checked that
|
||||
* some decision changed somewhere, and it passed against the broken bot: with a tight cap the
|
||||
* gated branches do change which Local Operations option is chosen, they just fail to stop the
|
||||
* card being played two steps later. "Something moved" is not the promise. The promise is that
|
||||
* committed trains never exceed what the Office can hold.
|
||||
*
|
||||
* Office tiers only ever go up, so the final A/D count is the most generous the cap ever was.
|
||||
*/
|
||||
const committed = (s: GameState): number =>
|
||||
s.timetable.filter((t) => t !== null).length + s.pendingExtras.length;
|
||||
|
||||
for (const slack of [0, 1]) {
|
||||
const capped = makeDeveloperBot({ trainCapSlack: slack });
|
||||
let cappedWorst = 0;
|
||||
let baseWorst = 0;
|
||||
for (const seed of [1000, 8919, 16838, 24757, 32676, 40595, 48514]) {
|
||||
const a = game(seed);
|
||||
playGame(a, capped, pump);
|
||||
const roomA = adTrackCount(a, 0) + slack;
|
||||
cappedWorst = Math.max(cappedWorst, committed(a) - roomA);
|
||||
|
||||
const b = game(seed);
|
||||
playGame(b, developerBot, pump);
|
||||
baseWorst = Math.max(baseWorst, committed(b) - (adTrackCount(b, 0) + slack));
|
||||
}
|
||||
assert.ok(
|
||||
cappedWorst <= 0,
|
||||
`trainCapSlack=${slack} let the bot commit ${cappedWorst} train(s) more than the Office can hold`,
|
||||
);
|
||||
// And the cap is not vacuous: the uncapped bot really does overshoot on these seeds.
|
||||
assert.ok(
|
||||
baseWorst > 0,
|
||||
`slack=${slack}: the baseline never overshoots on these seeds, so the test proves nothing`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
it('names itself so a report cannot confuse two runs', () => {
|
||||
assert.equal(developerBot.name, 'developer');
|
||||
assert.equal(makeDeveloperBot({}).name, 'developer');
|
||||
assert.equal(makeDeveloperBot({ trainCapSlack: 1 }).name, 'developer+trainCapSlack=1');
|
||||
});
|
||||
|
||||
it('refuses a tweak name it does not know', () => {
|
||||
// A typo silently parsed as "no tweaks" would compare the bot against itself and report a
|
||||
// confident zero — the most expensive possible failure of this tool.
|
||||
assert.deepEqual(parseTweaks(['400', 'trainCapSlack=2']), { trainCapSlack: 2 });
|
||||
assert.throws(() => parseTweaks(['trainCapSlok=1']), /unknown tweak/);
|
||||
});
|
||||
});
|
||||
|
||||
describe('the funnel counts what it claims to count', () => {
|
||||
it('keeps every gate inside the total it is a fraction of', () => {
|
||||
const s = game(430);
|
||||
const probe = makeFunnelProbe(0);
|
||||
const r = playGame(s, developerBot, pump, 50_000, probe.onEvent, probe.onTurn);
|
||||
const f = probe.funnel;
|
||||
|
||||
assert.ok(f.arrivals > 0, 'no train reached an Office in a whole game');
|
||||
for (const [name, n] of [
|
||||
['at a passenger office', f.arrivalsAtPassengerOffice],
|
||||
['with an empty coach', f.arrivalsWithEmptyCoach],
|
||||
['with a loaded coach', f.arrivalsWithLoadedCoach],
|
||||
['with a wanted car', f.arrivalsWithWantedCar],
|
||||
] as const) {
|
||||
assert.ok(n <= f.arrivals, `${name} (${n}) exceeds arrivals (${f.arrivals})`);
|
||||
}
|
||||
|
||||
assert.ok(f.cargoPhases > 0, 'no Cargo phase was sampled');
|
||||
for (const [name, n] of [
|
||||
['with a facility', f.cargoWithFacility],
|
||||
['green stocked', f.cargoGreenStocked],
|
||||
['car spotted', f.cargoCarSpotted],
|
||||
['laborer free', f.cargoLaborerFree],
|
||||
['ready', f.cargoReady],
|
||||
] as const) {
|
||||
assert.ok(n <= f.cargoPhases, `${name} (${n}) exceeds Cargo phases (${f.cargoPhases})`);
|
||||
}
|
||||
// "All three at one industry" cannot exceed any of the three it is made of.
|
||||
assert.ok(f.cargoReady <= Math.min(f.cargoGreenStocked, f.cargoCarSpotted, f.cargoLaborerFree));
|
||||
assert.ok(f.buriedWithDigAvailable <= f.buriedTurns);
|
||||
|
||||
// Sampled once per Cargo phase, never once per decision: 5 Days x 12 Stages is the ceiling.
|
||||
assert.ok(f.cargoPhases <= 60, `${f.cargoPhases} Cargo phases in a 5-Day game`);
|
||||
|
||||
// And the arrivals it counted are the arrivals that happened.
|
||||
const arrived = r.events.filter((e) => e.type === 'trainArrived').length;
|
||||
assert.equal(f.arrivals, arrived, 'the probe and the event log disagree about arrivals');
|
||||
});
|
||||
|
||||
it('rides along on the harness summary', () => {
|
||||
const s = game(202);
|
||||
const probe = makeFunnelProbe(0);
|
||||
const r = playGame(s, developerBot, pump, 50_000, probe.onEvent, probe.onTurn);
|
||||
const stats = summarize(202, r.events, r.intents, s, probe.funnel);
|
||||
assert.ok(stats.funnel, 'summarize dropped the funnel');
|
||||
assert.equal(stats.funnel!.arrivals, probe.funnel.arrivals);
|
||||
// Optional, so a caller with no probe still gets a summary.
|
||||
assert.equal(summarize(202, r.events, r.intents, s).funnel, undefined);
|
||||
});
|
||||
});
|
||||
|
||||
describe('the comparison arithmetic', () => {
|
||||
it('reports a dead heat as a dead heat', () => {
|
||||
// Comparing the bot with itself must produce exactly zero, no games differing, and a t of 0 —
|
||||
// if the machinery has any asymmetry in it, this is where it shows.
|
||||
const r = compare({}, 12);
|
||||
assert.equal(r.mean, 0);
|
||||
assert.equal(r.identical, 12);
|
||||
assert.equal(r.better, 0);
|
||||
assert.equal(r.worse, 0);
|
||||
assert.equal(r.t, 0);
|
||||
});
|
||||
|
||||
it('pairs by seed, and deals the same seeds the harness does', () => {
|
||||
const r = compare({ trainCapSlack: 1 }, 5);
|
||||
assert.deepEqual(
|
||||
r.deltas.map((d) => d.seed),
|
||||
[0, 1, 2, 3, 4].map((i) => 1000 + i * 7919),
|
||||
'compare and the harness must talk about the same games',
|
||||
);
|
||||
for (let i = 0; i < r.deltas.length; i++) {
|
||||
assert.equal(r.baseStats[i]!.seed, r.deltas[i]!.seed);
|
||||
assert.equal(r.variantStats[i]!.seed, r.deltas[i]!.seed);
|
||||
assert.equal(r.deltas[i]!.delta, r.variantStats[i]!.revenue.net - r.baseStats[i]!.revenue.net);
|
||||
}
|
||||
});
|
||||
|
||||
it('computes the standard error from the PAIRED difference', () => {
|
||||
// The whole gain of pairing is that σ of the difference is smaller than σ of either side. Using
|
||||
// the level's σ by mistake would quietly restore the ±1.0 noise floor this exists to escape.
|
||||
const r = compare({ trainCapSlack: 1 }, 40);
|
||||
const d = r.deltas.map((x) => x.delta);
|
||||
const m = d.reduce((p, c) => p + c, 0) / d.length;
|
||||
const sd = Math.sqrt(d.reduce((p, c) => p + (c - m) ** 2, 0) / (d.length - 1));
|
||||
assert.ok(Math.abs(r.mean - m) < 1e-9);
|
||||
assert.ok(Math.abs(r.sd - sd) < 1e-9);
|
||||
assert.ok(Math.abs(r.stderr - sd / Math.sqrt(d.length)) < 1e-9);
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user