added harness to support bots being able to run / test strategies.

This commit is contained in:
Jesse
2026-08-08 15:07:42 -04:00
parent 5d825b97d2
commit 35575147cd
9 changed files with 866 additions and 23 deletions
+74
View File
@@ -21,6 +21,80 @@ page as `v0.1.0 · <sha> · <date>`, so what is deployed can always be identifie
## Unreleased
## 0.1.1 — 2026-08-08
### A way to tell whether a change to the bot helped
The bot averages **1.4 Revenue against a target of 20**, and the obvious next step is to teach it to
play better. That step could not be taken honestly, because there was no way to tell whether a
heuristic had helped: revenue has σ ≈ 9 across games, so two runs of the *identical* bot differ by
about a point through nothing but the deal. Every single-change claim in this changelog before the
Interlocking work is inside that noise, and `TODO.md` has said so for a while.
**`node src/sim/compare.ts 1600 trainCapSlack=1`** runs the current bot and one variant over the same
deals and reports the per-seed difference. Giving both sides the same seed takes the deal out of the
comparison: σ drops from ~9 on the level to **5.3 on the difference**, and 1600 seeds puts the
standard error at **±0.13** — in 1m45s. The noise floor moves from ±1.0 to about ±0.15, which makes
every heuristic in the training plan resolvable. No parallelism needed; 100 games take 4.8s.
**The report prints the better/worse/identical split beside the mean**, because they are different
claims. The first tweak measured is the case in point — `trainCapSlack=1` over 1600 paired seeds:
```
REVENUE DELTA +0.77 ± 0.13 (t = 5.97, σ of the paired difference 5.13)
seeds better 146 · worse 105 · identical 1349
best seeds: +75, +50, +46 worst seeds: -26, -20, -13
```
**84% of games are untouched.** It does not make the bot play better; it removes a rare catastrophe,
and the mean rides on a handful of rescued games. Reporting that as "revenue up 65%" would be
arithmetically true and misleading about what changed. At 400 seeds the same tweak read t = 2.57 —
"not proven" — which is exactly the verdict it deserved there.
The tweak is **measured but not adopted**: the flag stays off, so this commit changes no bot
behaviour. Turning it on is a change to how the bot plays and belongs in its own reviewed commit.
**And it prints the funnel for both sides**, because revenue can rise two ways: a channel started
working, or an expensive channel was abandoned for a cheap one. A strict train cap raises revenue
*and* cuts freight events nearly in half, and that has to be visible rather than inferred.
**The funnel is new** — `GameStats.funnel`, sampled live rather than recovered from the log, because
the interesting gates are conditions rather than occurrences. "Was a green box stocked while a car
was spotted" is not a thing that happens; it is true or false at a moment. It needed a per-decision
hook on `playGame` (`TurnObserver`), separate from the existing event observer, so a phase is
sampled once instead of once per event in the batch. What it says about the current bot:
- **Passengers:** of 7.4 arrivals a game, 70% reach a Passenger Facility, 25% carry the empty coach
boarding requires, 21% the loaded coach detraining requires.
- **Freight:** of 60 Cargo phases, a green box is stocked in **8%** and all three requirements meet
at one industry in **7%**. Freight is gated almost entirely on stocking.
- **Stuck:** 8.4 decisions a game are taken with a train's engine buried mid-consist, and on **3%**
of them is there a legal way to set the nose cars out — because it happens at the Office, where
Rolling Stock may not be left.
**`developerBot` is now `makeDeveloperBot({})`**, byte-identical to what came before (asserted on
three seeds by full event-stream fingerprint, and 1.4400 mean on 100 games either way). Variants
exist so two policies can be compared in one process rather than by editing the bot between runs —
which is how you end up comparing two things you cannot reproduce. Tweaks are temporary: a flag that
measures well becomes the default and is deleted in the same commit.
**The first tweak found the bug this tooling exists to find, twice over.** `trainCapSlack` caps
committed trains against the Office's A/D tracks. Written first as a gate on the two branches whose
comments say they exist to play a train card, it measured **exactly zero difference over 400 paired
seeds** — because `followThrough` ends with a generic "play what is in hand" fallback that played the
card anyway. A cap has to remove the option, not guard the branches that reach for it.
Then the *test* for it was wrong in the same shape: it asserted only that some decision changed
somewhere, and **passed against the broken bot**, because a tight cap does change which Local
Operations option gets chosen — it just fails to stop the card being played two steps later. It now
asserts the contract (committed trains never exceed what the Office can hold) and is verified to fail
against the broken version. "Something moved" is not the promise.
Also new: a determinism self-check. The same policy on the same seed must produce an identical event
stream, since the paired method rests on it and the failure mode is silent.
Tests 403 → 413.
## 0.1.0 — 2026-08-08
The first numbered build. Everything below the "second playtest pass" heading was made under
+19
View File
@@ -64,6 +64,25 @@ npm run typecheck # tsc --noEmit
Because Node strips types rather than compiling them, the codebase is restricted to **erasable
syntax**: no `enum`, no parameter properties, no namespaces. `tsconfig.json` enforces this.
### Measuring the bot
```sh
node src/sim/harness.ts 200 # how the bot does, with the funnel
node src/sim/compare.ts 1600 trainCapSlack=1 # one change, paired against the current bot
```
**Never judge a heuristic on an unpaired run.** Revenue has σ ≈ 9 across games, so two runs of the
*identical* bot differ by about a point through nothing but the deal. `compare.ts` gives both
policies the same seed and reports the per-seed difference, where σ is 5.3 — 1600 seeds puts the
standard error at ±0.13, in under two minutes. Keep a change at **t ≥ 3**, and read the
better/worse/identical split beside the mean: a gain carried by a few rescued games is a different
claim from one spread across the field.
Variants come from `makeDeveloperBot(tweaks)`. A tweak is **temporary** — when it measures well it
becomes the default and the flag is deleted in the same commit; when it measures badly it is deleted
with the finding recorded in `CHANGELOG.md`. A bot that accumulates switches nobody can account for
is the thing this machinery exists to prevent.
## Design notes worth knowing
- **The rules engine is pure.** No I/O, no clock, no sockets, and all randomness derives from one
+11 -6
View File
@@ -66,12 +66,17 @@ Ordered within each section by how much it is currently costing us.
Revisit when setup gains an interactive phase; the orientation matters, because it decides
which direction climbs and therefore what Brakeman and Helpers are worth.
- [ ] **Measure with error bars from now on.** Revenue has a standard deviation of ~9, so a
100-game run carries about ±1.0 of noise — every single-change revenue claim in the changelog
before the Interlocking work is inside that. Use paired per-seed comparison (the harness deals
the same seeds either way) and 400+ games before calling a heuristic change good or bad. The
first attempt at the Running Track straight was read as a 0.6 REGRESSION on 100 games and is
a 0.7 improvement on 400.
- [x] ~~**Measure with error bars from now on.**~~ Built: `node src/sim/compare.ts 1600 <tweak>=<n>`
runs the current bot and one variant over the same deals and reports the paired difference.
Pairing drops σ from ~9 on the level to **5.3 on the difference**, so 1600 seeds gives ±0.13 in
about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is
**t ≥ 3**, and the report prints the better/worse/identical split beside the mean, because a
mean carried by a skewed tail is a different claim from broad improvement.
- [ ] **Re-run the three "worth ~0" action-mix experiments against the new floor.** Capping the draw,
pairing the two halves of a load, and restricting Enhancements were each measured "within noise
of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would
have looked like nothing. They are nearly free to re-run now and at least one may have been
discarded wrongly.
- [x] ~~**Confirm the Classification Yard rule against the source.**~~ Confirmed, and the guess was
wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses
straight back to the Division Yard, and the Classification Yard empties ONLY when the Division
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "station-master",
"version": "0.1.0",
"version": "0.1.1",
"private": true,
"type": "module",
"description": "Station Master — a railroad operations game",
+115 -10
View File
@@ -20,7 +20,7 @@
*/
import { applyIntent, areaOf, canAdvanceLoad, destinationsFor, facilityCarTypes, laborersLeft } from '../engine/apply.ts';
import { MAX_CONSIST, nextOfficeTier } from '../engine/content.ts';
import { MAX_CONSIST, nextOfficeTier, officeProfile } from '../engine/content.ts';
import type { Hand, TrackGeometry } from '../engine/content.ts';
import type { GameEvent } from '../engine/events.ts';
import type { Intent } from '../engine/intents.ts';
@@ -78,8 +78,42 @@ function because(reason: string, intent: Intent): Intent {
return intent;
}
export const developerBot: BotPolicy = {
name: 'developer',
/**
* KNOBS FOR A/B MEASUREMENT, and nothing else.
*
* A heuristic change has to be measured against the bot it replaces, over the SAME deals — and
* editing the bot between runs makes that impossible to do honestly, because the two sides of the
* comparison never exist at once. Every flag here is off by default, so `makeDeveloperBot({})` is
* byte-identical to the bot that came before this existed.
*
* TEMPORARY BY CONSTRUCTION. When a flag measures well it becomes the default and the flag is
* deleted in the same commit; when it measures badly it is deleted with its finding recorded in the
* changelog. What must not happen is a bot that accumulates switches nobody can account for — a
* heuristic with no measurement attached is exactly what this machinery exists to prevent.
*/
export type BotTweaks = {
/**
* Refuse to schedule a train the Office has no room for: cap committed trains (timetable slots
* filled, plus queued Extras) at `adTracks + trainCapSlack`.
*
* `undefined` means no cap — today's behaviour, which plays every train card on sight. A Whistle
* Post has ONE A/D track, and a second train standing there makes every later arrival an
* automatic collision (Gap 2d).
*/
trainCapSlack?: number;
};
/** The bot as it plays today. Every knob off. */
export const developerBot: BotPolicy = makeDeveloperBot({});
/** A variant, for measuring one change at a time against the bot above. */
export function makeDeveloperBot(tweaks: BotTweaks): BotPolicy {
const suffix = Object.entries(tweaks)
.filter(([, v]) => v !== undefined)
.map(([k, v]) => `${k}=${v}`)
.join(',');
return {
name: suffix ? `developer+${suffix}` : 'developer',
choose(s, player, options) {
lastReason = 'no specific reason — first legal option';
@@ -133,13 +167,29 @@ export const developerBot: BotPolicy = {
// --- Local Operations.
if (s.clock.phase === 'localOps') {
if (s.turn.option === null) return chooseLocalOption(s, player, options);
return followThrough(s, player, options);
/**
* A CAP HAS TO BE APPLIED TO THE OPTIONS, NOT TO ONE BRANCH.
*
* The first attempt gated the two branches that exist to play a train card, measured exactly
* zero difference over 400 paired seeds, and was right to: `followThrough` ends with a generic
* "play what is in hand" fallback that played the train card anyway. Removing the option
* itself is the only way to be sure no path reaches it.
*
* Never to an empty list — a bot with nothing legal to choose is a crash, and §6.2 can force
* a play when the hand is over its limit. If filtering leaves nothing, the cap yields.
*/
const held = trainWouldOverfillTheOffice(s, player, tweaks)
? options.filter((i) => !(i.type === 'card.play' && isTrainCard(s, i.cardId)))
: options;
const usable = held.length > 0 ? held : options;
if (s.turn.option === null) return chooseLocalOption(s, player, usable, tweaks);
return followThrough(s, player, usable, tweaks);
}
return options[0]!;
},
};
}
/**
* The one decision that matters: which of §6's three things to spend this Stage on.
@@ -150,7 +200,37 @@ export const developerBot: BotPolicy = {
* and spent the rest of the game stuffing green boxes whose loads could never complete for want of
* a car to load them onto.
*/
function chooseLocalOption(s: GameState, player: PlayerIndex, options: Intent[]): Intent {
/**
* Trains this player has already committed to running: slots filled on the timetable, plus Extras
* played and waiting for a Crew Tray.
*
* A Second Section is deliberately NOT counted. Its whole purpose is to put a following train into
* an occupied Subdivision and force the Superintendent's ruling (Q9) — capping it would be capping
* the card's reason to exist.
*/
function committedTrains(s: GameState): number {
return s.timetable.filter((t) => t !== null).length + s.pendingExtras.length;
}
/**
* Would scheduling another train exceed what the Office can physically hold?
*
* §7 lets you play as many train cards as you draw, and a train that arrives with nowhere to stand
* is an automatic collision (Gap 2d) — so the two rules together make a train card actively harmful
* once the A/D tracks are spoken for. Off unless `trainCapSlack` is set.
*/
function trainWouldOverfillTheOffice(s: GameState, player: PlayerIndex, tweaks: BotTweaks): boolean {
if (tweaks.trainCapSlack === undefined) return false;
const cap = officeProfile(areaOf(s, player).tier).adTracks + tweaks.trainCapSlack;
return committedTrains(s) >= cap;
}
function chooseLocalOption(
s: GameState,
player: PlayerIndex,
options: Intent[],
tweaks: BotTweaks,
): Intent {
const choices = options.filter(
(i): i is Extract<Intent, { type: 'localOps.choose' }> => i.type === 'localOps.choose',
);
@@ -167,7 +247,11 @@ function chooseLocalOption(s: GameState, player: PlayerIndex, options: Intent[])
return because('an Office upgrade is in hand — more A/D track means fewer collisions', can('draw')!);
}
if (can('draw') && hand.some((id) => isTrainCard(s, id))) {
if (
can('draw') &&
hand.some((id) => isTrainCard(s, id)) &&
!trainWouldOverfillTheOffice(s, player, tweaks)
) {
return because('a train card is in hand and its value compounds every Day', can('draw')!);
}
@@ -933,7 +1017,12 @@ function moveTowardOffice(s: GameState, player: PlayerIndex, options: Intent[]):
}
/** Once an option is chosen, work it to a sensible conclusion. */
function followThrough(s: GameState, player: PlayerIndex, options: Intent[]): Intent {
function followThrough(
s: GameState,
player: PlayerIndex,
options: Intent[],
tweaks: BotTweaks,
): Intent {
switch (s.turn.option) {
case 'draw': {
// Draw before playing — otherwise the hand empties and never refills.
@@ -962,8 +1051,12 @@ function followThrough(s: GameState, player: PlayerIndex, options: Intent[]): In
);
if (upgrade) return because('upgrade the Office — a Whistle Post has ONE A/D track and a second arrival collides', upgrade);
// Train cards next — they take no placement and their value compounds every Day.
const train = options.find(
// Train cards next — they take no placement and their value compounds every Day. Unless the
// Office cannot hold another one, in which case the card is HELD rather than played: an
// arrival with no free A/D track is an automatic collision, and it repeats every Day.
const train = trainWouldOverfillTheOffice(s, player, tweaks)
? undefined
: options.find(
(i) => i.type === 'card.play' && i.placement === undefined && isTrainCard(s, i.cardId),
);
if (train) return because('a scheduled train runs every Day thereafter — the only card whose value compounds', train);
@@ -1504,12 +1597,23 @@ export type PlayOutcome = {
*/
export type PlayObserver = (event: GameEvent, state: GameState) => void;
/**
* Called once per DECISION, immediately before the policy is asked to choose.
*
* Separate from `PlayObserver` because the two sample different things. An event is something that
* happened; a turn hook sees the position as the bot sees it, which is the only place to ask a
* question of the form "was anything ready to be done here?" — a condition rather than an
* occurrence. Sampling those off events instead would count them once per event in the batch.
*/
export type TurnObserver = (state: GameState) => void;
export function playGame(
s: GameState,
policy: BotPolicy,
pumpFn: (s: GameState) => GameEvent[],
maxTurns = 50_000,
observe?: PlayObserver,
onTurn?: TurnObserver,
): PlayOutcome {
const events: GameEvent[] = [];
const intents: Intent['type'][] = [];
@@ -1545,6 +1649,7 @@ export function playGame(
const options = legalActions(s, actor);
if (options.length === 0) break;
onTurn?.(s);
const chosen = policy.choose(s, actor, options);
intents.push(chosen.type);
const r = applyIntent(s, actor, chosen);
+233
View File
@@ -0,0 +1,233 @@
/**
* Component 18b — paired A/B measurement for bot heuristics.
*
* Dev-side only. Runs the current bot and one tweaked variant over the SAME deals and reports the
* per-seed difference.
*
* WHY PAIRED, AND WHY IT IS NOT OPTIONAL. Revenue has a standard deviation of about 9 across games,
* so two 100-game runs of the identical bot can differ by a point through nothing but the deal. Every
* single-change claim in the changelog before the Interlocking work is inside that noise. Giving both
* policies the same seed removes the deal from the comparison entirely: what is left is the change.
* Measured here, the paired difference has σ ≈ 5.3 against ≈ 9 unpaired, and at 1600 seeds the
* standard error is ±0.13 — so a +0.5 heuristic is resolvable in under two minutes, where unpaired it
* would need tens of thousands of games.
*
* WHY THE SPLIT IS PRINTED. A mean carried by a skewed tail is a different claim from a mean carried
* by broad improvement, and only the better/worse/identical counts tell them apart. The train cap is
* the case in point: +0.60 overall, but 130 seeds better, 145 worse and 1325 unchanged — it removes a
* rare catastrophe rather than making the bot play better.
*
* WHY THE FUNNEL IS PRINTED. Revenue can rise because a channel started working or because the bot
* abandoned an expensive channel for a cheap one, and the two are identical in a single number.
*
* Run with:
* node src/sim/compare.ts 1600 trainCapSlack=1
* node src/sim/compare.ts 400 trainCapSlack=0 --length short
*/
import { pump } from '../engine/advance.ts';
import { createGame } from '../engine/setup.ts';
import type { GameLength } from '../engine/content.ts';
import type { GameConfig, GameMode } from '../engine/state.ts';
import type { BotPolicy, BotTweaks } from './bot.ts';
import { developerBot, makeDeveloperBot, playGame } from './bot.ts';
import type { Funnel, GameStats } from './stats.ts';
import { funnelReport, makeFunnelProbe, summarize } from './stats.ts';
export type PairedResult = {
seeds: number;
baseline: string;
variant: string;
/** Per-seed revenue difference, variant minus baseline. */
deltas: { seed: number; delta: number }[];
mean: number;
/** Standard error of the mean difference — the number that decides whether this is real. */
stderr: number;
sd: number;
t: number;
better: number;
worse: number;
identical: number;
baseStats: GameStats[];
variantStats: GameStats[];
};
const SOLO = (length: GameLength, mode: GameMode): GameConfig => ({
mode,
victory: 'highestAfterDays',
length,
optionalRules: {
reducedVisibility: false,
sisterTrains: false,
employeeRotation: false,
emergencyToolbox: false,
},
});
/**
* One game, one seed, one policy.
*
* The seed stride matches `harness.ts` exactly, so a compare run and a harness run of the same size
* are talking about the same games — otherwise two numbers that ought to agree would not, for a
* reason nobody would find.
*/
function runOne(policy: BotPolicy, seed: number, length: GameLength, mode: GameMode): GameStats {
const s = createGame({ id: `cmp-${seed}`, seed, config: SOLO(length, mode), playerNames: ['bot'] });
const probe = makeFunnelProbe(0);
const r = playGame(s, policy, pump, 50_000, probe.onEvent, probe.onTurn);
return summarize(seed, r.events, r.intents, s, probe.funnel);
}
export function compare(
tweaks: BotTweaks,
games: number,
length: GameLength = 'standard',
mode: GameMode = 'solitaire',
): PairedResult {
const variantPolicy = makeDeveloperBot(tweaks);
const baseStats: GameStats[] = [];
const variantStats: GameStats[] = [];
const deltas: { seed: number; delta: number }[] = [];
for (let i = 0; i < games; i++) {
const seed = 1000 + i * 7919; // the same prime stride the harness deals
const a = runOne(developerBot, seed, length, mode);
const b = runOne(variantPolicy, seed, length, mode);
baseStats.push(a);
variantStats.push(b);
deltas.push({ seed, delta: b.revenue.net - a.revenue.net });
}
const d = deltas.map((x) => x.delta);
const mean = d.reduce((p, c) => p + c, 0) / Math.max(1, d.length);
const sd =
d.length > 1
? Math.sqrt(d.reduce((p, c) => p + (c - mean) ** 2, 0) / (d.length - 1))
: 0;
const stderr = d.length > 0 ? sd / Math.sqrt(d.length) : 0;
return {
seeds: games,
baseline: developerBot.name,
variant: variantPolicy.name,
deltas,
mean,
sd,
stderr,
t: stderr === 0 ? 0 : mean / stderr,
better: d.filter((x) => x > 0).length,
worse: d.filter((x) => x < 0).length,
identical: d.filter((x) => x === 0).length,
baseStats,
variantStats,
};
}
// ---------------------------------------------------------------------------
// Reporting
// ---------------------------------------------------------------------------
const mean = (xs: number[]): number => (xs.length ? xs.reduce((p, c) => p + c, 0) / xs.length : 0);
/**
* The verdict, stated in the same terms every time.
*
* The threshold is t ≥ 3 rather than the conventional 2. This is a measurement taken repeatedly on
* the same system while looking for something that works, so the conventional bar would have us keep
* roughly one bad heuristic in twenty; and the cost of a wrong keep is not a wrong paper, it is a bot
* that quietly plays worse and takes every later measurement with it.
*/
function verdict(t: number): string {
const a = Math.abs(t);
if (a >= 3) return t > 0 ? 'KEEP — clears the bar (t ≥ 3)' : 'REJECT — significantly worse';
if (a >= 2) return 'NOT PROVEN — suggestive, run more seeds before believing it';
return 'NO EFFECT MEASURED — inside the noise';
}
export function formatPaired(r: PairedResult): string {
const out: string[] = [];
const num = (x: number, w = 8, dp = 2): string => x.toFixed(dp).padStart(w);
out.push(`\n=== paired comparison · ${r.seeds} seeds · same deal to both ===\n`);
out.push(` baseline ${r.baseline}`);
out.push(` variant ${r.variant}\n`);
const rows: [string, (g: GameStats) => number][] = [
['revenue', (g) => g.revenue.net],
['collisions', (g) => g.collisions],
['trains scheduled', (g) => g.trains.scheduled],
['arrivals', (g) => g.trains.arrived],
['passengers on/off', (g) => g.revenue.passengerBoard + g.revenue.passengerDetrain],
['freight loads+unloads', (g) => g.revenue.freightLoad + g.revenue.freightUnload],
['cards played', (g) => g.development.cardsPlayed],
['final grid size', (g) => g.development.gridSize],
];
out.push(' baseline variant delta');
for (const [label, pick] of rows) {
const a = mean(r.baseStats.map(pick));
const b = mean(r.variantStats.map(pick));
out.push(` ${label.padEnd(22)}${num(a)} ${num(b)} ${num(b - a)}`);
}
out.push(
`\n REVENUE DELTA ${r.mean >= 0 ? '+' : ''}${r.mean.toFixed(2)} ± ${r.stderr.toFixed(2)}` +
` (t = ${r.t.toFixed(2)}, σ of the paired difference ${r.sd.toFixed(2)})`,
);
out.push(` ${verdict(r.t)}`);
out.push(
`\n seeds better ${r.better} · worse ${r.worse} · identical ${r.identical}` +
` — ${((r.identical / Math.max(1, r.seeds)) * 100).toFixed(0)}% of games are untouched by this change`,
);
// The tails, BY SEED, so a loss can be replayed rather than averaged away.
const sorted = [...r.deltas].sort((a, b) => a.delta - b.delta);
const show = (xs: { seed: number; delta: number }[]): string =>
xs.map((x) => `${x.seed} (${x.delta > 0 ? '+' : ''}${x.delta})`).join(', ');
out.push(` worst seeds: ${show(sorted.slice(0, 3))}`);
out.push(` best seeds: ${show(sorted.slice(-3).reverse())}`);
// A change that raises revenue by abandoning a channel has to be visible as that.
out.push('\n --- baseline funnel ---');
out.push(funnelReport(r.baseStats));
out.push('\n --- variant funnel ---');
out.push(funnelReport(r.variantStats));
return out.join('\n');
}
// ---------------------------------------------------------------------------
// CLI
// ---------------------------------------------------------------------------
/** `trainCapSlack=1` -> `{ trainCapSlack: 1 }`. Unknown names are refused rather than ignored. */
export function parseTweaks(args: string[]): BotTweaks {
const known = new Set(['trainCapSlack']);
const tweaks: Record<string, number> = {};
for (const a of args) {
const m = /^([A-Za-z]\w*)=(-?\d+(?:\.\d+)?)$/.exec(a);
if (!m) continue;
if (!known.has(m[1]!)) {
throw new Error(`unknown tweak "${m[1]}" — known: ${[...known].join(', ')}`);
}
tweaks[m[1]!] = Number(m[2]);
}
return tweaks as BotTweaks;
}
const isMain = process.argv[1]?.endsWith('compare.ts') ?? false;
if (isMain) {
const args = process.argv.slice(2);
const games = Number(args.find((a) => /^\d+$/.test(a)) ?? 400);
const lengthIdx = args.indexOf('--length');
const length = (lengthIdx >= 0 ? args[lengthIdx + 1] : 'standard') as GameLength;
const tweaks = parseTweaks(args);
if (Object.keys(tweaks).length === 0) {
console.error('nothing to compare — pass at least one tweak, e.g. trainCapSlack=1');
process.exitCode = 1;
} else {
console.log(formatPaired(compare(tweaks, games, length)));
}
}
export type { Funnel };
+6 -3
View File
@@ -19,7 +19,7 @@ import type { GameConfig, GameMode } from '../engine/state.ts';
import type { BotPolicy, PlayOutcome } from './bot.ts';
import { developerBot, playGame, randomBot } from './bot.ts';
import type { GameStats } from './stats.ts';
import { formatAggregate, summarize } from './stats.ts';
import { formatAggregate, makeFunnelProbe, summarize } from './stats.ts';
export type SimOptions = {
games: number;
@@ -86,9 +86,12 @@ export function simulate(opts: SimOptions): SimReport {
config: configFor(opts.mode, opts.length),
playerNames: opts.players,
});
const r = playGame(s, opts.policy, pump);
// The probe watches the game as it is played: the funnel gates are conditions at a moment, not
// events, so they cannot be recovered from the log afterwards.
const probe = makeFunnelProbe(0);
const r = playGame(s, opts.policy, pump, 50_000, probe.onEvent, probe.onTurn);
results.push(r);
stats.push(summarize(seed, r.events, r.intents, s));
stats.push(summarize(seed, r.events, r.intents, s, probe.funnel));
}
const finished = results.filter((r) => r.finished);
+185
View File
@@ -18,11 +18,149 @@
import type { GameEvent } from '../engine/events.ts';
import type { Intent } from '../engine/intents.ts';
import type { GameState } from '../engine/state.ts';
import { areaOf, facilityCarTypes } from '../engine/apply.ts';
import { officeProfile } from '../engine/content.ts';
// ---------------------------------------------------------------------------
// Per-game summary
// ---------------------------------------------------------------------------
/**
* THE FUNNEL — how far each opportunity got before it stopped.
*
* The totals above say what happened; this says what ALMOST happened, which is what a change to the
* bot has to move. Revenue can rise because a channel started working or because a channel was
* abandoned in favour of a cheaper one, and the two look identical in a single number: capping the
* trains a bot schedules raises revenue while cutting freight events nearly in half.
*
* SAMPLED FROM LIVE STATE, not from events, because the interesting gates are conditions rather than
* occurrences — "was a green box stocked while a car was spotted" is not a thing that happens, it is
* a thing that is true or false at a moment. `playGame` calls the probe once per decision and the
* probe dedupes per phase, so these count PHASES, not turns.
*/
export type Funnel = {
/** Trains reaching an Office, and what they were carrying when they got there. */
arrivals: number;
arrivalsAtPassengerOffice: number;
/** Boarding needs an empty coach aboard (§9.2); detraining needs a loaded one. */
arrivalsWithEmptyCoach: number;
arrivalsWithLoadedCoach: number;
/** A car of a type some industry in the district actually wants, loaded or empty. */
arrivalsWithWantedCar: number;
/** Cargo phases, and how many had each half of what a load needs. */
cargoPhases: number;
cargoWithFacility: number;
cargoGreenStocked: number;
cargoCarSpotted: number;
cargoLaborerFree: number;
/** All three at ONE industry — the only combination that can actually move a load. */
cargoReady: number;
/**
* Decisions taken with a train's engine buried mid-consist, and how many of those had a legal way
* to dig it out. §8.2 will not let such a train leave, and Rolling Stock may not be set out at the
* Office — so the gap between these two numbers is the trap the bot cannot escape from.
*/
buriedTurns: number;
buriedWithDigAvailable: number;
};
function emptyFunnel(): Funnel {
return {
arrivals: 0, arrivalsAtPassengerOffice: 0, arrivalsWithEmptyCoach: 0,
arrivalsWithLoadedCoach: 0, arrivalsWithWantedCar: 0,
cargoPhases: 0, cargoWithFacility: 0, cargoGreenStocked: 0, cargoCarSpotted: 0,
cargoLaborerFree: 0, cargoReady: 0,
buriedTurns: 0, buriedWithDigAvailable: 0,
};
}
/**
* A probe to hand to `playGame`: an event hook, a per-decision hook, and the totals they fill.
*
* One per game. Kept here beside the statistics it produces rather than in the harness, so the
* definition of "a Cargo phase that was ready" lives in exactly one place.
*/
export function makeFunnelProbe(player = 0): {
funnel: Funnel;
onEvent: (e: GameEvent, s: GameState) => void;
onTurn: (s: GameState) => void;
} {
const funnel = emptyFunnel();
// Phases are sampled once, on the first decision taken in them.
let lastPhaseKey = '';
const onEvent = (e: GameEvent, s: GameState): void => {
if (e.type !== 'trainArrived') return;
funnel.arrivals += 1;
const area = areaOf(s, player);
if (officeProfile(area.tier).isPassengerFacility) funnel.arrivalsAtPassengerOffice += 1;
if (e.consist.some((c) => c.type === 'coach' && !c.loaded)) funnel.arrivalsWithEmptyCoach += 1;
if (e.consist.some((c) => c.type === 'coach' && c.loaded)) funnel.arrivalsWithLoadedCoach += 1;
for (const card of area.grid.values()) {
const f = card.facility;
if (!f || f.kind !== 'freight') continue;
const want = facilityCarTypes(f);
const useful = e.consist.some(
(c) =>
want.includes(c.type) &&
((f.allows.outbound && !c.loaded) || (f.allows.inbound && c.loaded)),
);
if (useful) {
funnel.arrivalsWithWantedCar += 1;
break;
}
}
};
const onTurn = (s: GameState): void => {
const area = areaOf(s, player);
// A buried engine is a per-decision condition: every turn it persists is a turn the train is
// stuck, so this counts turns rather than trains.
for (const tray of s.trays.values()) {
if (tray.position.at !== 'grid' || tray.position.owner !== player) continue;
if (tray.engineAt <= 0 || tray.engineAt >= tray.consist.length) continue;
funnel.buriedTurns += 1;
// Cars may never be set out at the Office (§A.4), which is where this almost always happens.
const atOffice =
tray.position.coord.row === area.officeCoord.row &&
tray.position.coord.col === area.officeCoord.col;
if (!atOffice) funnel.buriedWithDigAvailable += 1;
break;
}
if (s.clock.phase !== 'loadUnload') return;
const key = `${s.clock.day}/${s.clock.stage}`;
if (key === lastPhaseKey) return;
lastPhaseKey = key;
funnel.cargoPhases += 1;
let facility = false, green = false, car = false, laborer = false, ready = false;
for (const card of area.grid.values()) {
const f = card.facility;
if (!f || f.kind !== 'freight') continue;
facility = true;
const g = f.outboundBox.length > 0;
const c = f.industryTrack.cars.length > 0;
const l = f.laborers - f.usedThisStage.laborers > 0;
green ||= g;
car ||= c;
laborer ||= l;
ready ||= g && c && l;
}
if (facility) funnel.cargoWithFacility += 1;
if (green) funnel.cargoGreenStocked += 1;
if (car) funnel.cargoCarSpotted += 1;
if (laborer) funnel.cargoLaborerFree += 1;
if (ready) funnel.cargoReady += 1;
};
return { funnel, onEvent, onTurn };
}
export type RevenueBreakdown = {
freightLoad: number;
freightUnload: number;
@@ -76,6 +214,11 @@ export type GameStats = {
collisions: number;
eventCounts: Record<string, number>;
intentCounts: Record<string, number>;
/**
* How far each opportunity got. Optional because it needs a live probe during the game, which a
* caller reconstructing statistics from a saved event log cannot supply.
*/
funnel?: Funnel;
};
function inc(map: Record<string, number>, key: string, by = 1): void {
@@ -87,6 +230,7 @@ export function summarize(
events: GameEvent[],
intents: Intent['type'][],
final: GameState,
funnel?: Funnel,
): GameStats {
const eventCounts: Record<string, number> = {};
const intentCounts: Record<string, number> = {};
@@ -211,6 +355,7 @@ export function summarize(
eventCounts,
intentCounts,
...(funnel ? { funnel } : {}),
};
}
@@ -414,6 +559,44 @@ export function formatGameStats(g: GameStats): string {
return out.join('\n');
}
/**
* THE FUNNEL, AS PERCENTAGES OF THE THING THEY GATE.
*
* Deliberately not per-game averages: "3.6 arrivals carried an empty coach" says nothing without the
* 7.3 arrivals it is out of, and the whole point is to see WHERE an opportunity stops. Read the
* arrivals block as "of every train that reached an Office" and the Cargo block as "of every Cargo
* phase". Empty when nothing was probed.
*/
export function funnelReport(all: GameStats[]): string {
const fs = all.map((g) => g.funnel).filter((f): f is Funnel => f !== undefined);
if (fs.length === 0) return '';
const sum = (pick: (f: Funnel) => number): number => fs.reduce((n, f) => n + pick(f), 0);
const pct = (n: number, d: number): string =>
d === 0 ? ' —' : `${((n / d) * 100).toFixed(0).padStart(3)}%`;
const arr = sum((f) => f.arrivals);
const cargo = sum((f) => f.cargoPhases);
const out: string[] = [];
out.push(`\n PASSENGER FUNNEL — ${(arr / fs.length).toFixed(1)} arrivals a game`);
out.push(` at a Passenger Facility (Depot+) ${pct(sum((f) => f.arrivalsAtPassengerOffice), arr)}`);
out.push(` carrying an EMPTY coach (boarding) ${pct(sum((f) => f.arrivalsWithEmptyCoach), arr)}`);
out.push(` carrying a LOADED coach (detrain) ${pct(sum((f) => f.arrivalsWithLoadedCoach), arr)}`);
out.push(` carrying a car an industry wants ${pct(sum((f) => f.arrivalsWithWantedCar), arr)}`);
out.push(`\n FREIGHT FUNNEL — ${(cargo / fs.length).toFixed(0)} Cargo phases a game`);
out.push(` a freight facility was built ${pct(sum((f) => f.cargoWithFacility), cargo)}`);
out.push(` a green box was stocked ${pct(sum((f) => f.cargoGreenStocked), cargo)}`);
out.push(` a car was spotted on an industry ${pct(sum((f) => f.cargoCarSpotted), cargo)}`);
out.push(` a Laborer was free ${pct(sum((f) => f.cargoLaborerFree), cargo)}`);
out.push(` ALL THREE at one industry ${pct(sum((f) => f.cargoReady), cargo)}`);
const buried = sum((f) => f.buriedTurns);
out.push(`\n STUCK — ${(buried / fs.length).toFixed(1)} decisions a game with the engine buried mid-train`);
out.push(` ... of those, able to set the nose cars out ${pct(sum((f) => f.buriedWithDigAvailable), buried)}`);
return out.join('\n');
}
export function formatAggregate(all: GameStats[]): string {
const out: string[] = [];
out.push(`\n=== end-of-game statistics · ${all.length} games ===\n`);
@@ -449,6 +632,8 @@ export function formatAggregate(all: GameStats[]): string {
out.push(` ${t.padEnd(14)} ${n} (${((n / all.length) * 100).toFixed(0)}%)`);
}
out.push(funnelReport(all));
out.push('\n ACTION MIX (mean per game)');
out.push(` switch moves ${num(meanOf(all, (g) => g.actions.moves))}`);
out.push(` cars dropped ${num(meanOf(all, (g) => g.actions.drops))}`);
+219
View File
@@ -0,0 +1,219 @@
/**
* The measurement machinery itself.
*
* These tests exist because a measurement tool that is quietly wrong is worse than no tool: it does
* not fail, it just produces numbers that decide what the bot becomes. Each one guards a property
* the paired method depends on.
*/
import { describe, it } from 'node:test';
import assert from 'node:assert/strict';
import { pump } from '../src/engine/advance.ts';
import { createGame } from '../src/engine/setup.ts';
import { adTrackCount } from '../src/engine/state.ts';
import type { GameConfig, GameState } from '../src/engine/state.ts';
import type { BotPolicy } from '../src/sim/bot.ts';
import { developerBot, makeDeveloperBot, playGame } from '../src/sim/bot.ts';
import { compare, parseTweaks } from '../src/sim/compare.ts';
import { makeFunnelProbe, summarize } from '../src/sim/stats.ts';
const config: GameConfig = {
mode: 'solitaire',
victory: 'highestAfterDays',
length: 'standard',
optionalRules: {
reducedVisibility: false,
sisterTrains: false,
employeeRotation: false,
emergencyToolbox: false,
},
};
const game = (seed: number): GameState =>
createGame({ id: `h${seed}`, seed, config, playerNames: ['bot'] });
/** One game, reduced to a string: every event type in order, plus the final revenue. */
function fingerprint(policy: BotPolicy, seed: number): string {
const s = game(seed);
const r = playGame(s, policy, pump);
return `${r.revenue[0]}|${r.events.map((e) => e.type).join(',')}`;
}
describe('the paired method rests on determinism', () => {
it('plays the identical game twice from one seed', () => {
/**
* THE PROPERTY EVERYTHING ELSE DEPENDS ON. A paired comparison subtracts two games that share a
* deal; if the same policy on the same seed can diverge, the difference is measuring the engine's
* own noise and every heuristic verdict is worthless.
*
* It is also the engine's central claim — pure, seeded, no ambient randomness — and the failure
* mode is silent: one `Math.random`, one Map iterated in a different order, and this is the only
* thing that would notice.
*/
for (const seed of [1000, 8919, 775569289]) {
assert.equal(fingerprint(developerBot, seed), fingerprint(developerBot, seed), `seed ${seed} diverged`);
}
});
it('gives the untweaked variant the identical game too', () => {
// `developerBot` IS `makeDeveloperBot({})`, and the refactor that introduced tweaks must not
// have changed a single decision. If this fails, every number measured before it is unusable.
for (const seed of [1000, 8919, 775569289]) {
assert.equal(
fingerprint(developerBot, seed),
fingerprint(makeDeveloperBot({}), seed),
`seed ${seed}: the untweaked variant plays a different game`,
);
}
});
});
describe('a tweak has to actually do something', () => {
it('actually holds the trains it says it holds', () => {
/**
* REGRESSION, and the reason this file exists. The first `trainCapSlack` gated the two branches
* whose comments say they exist to play a train card — and measured **exactly zero difference
* over 400 paired seeds**, because `followThrough` ends with a generic "play what is in hand"
* fallback that played the card anyway.
*
* ASSERT THE CONTRACT, NOT MERELY A DIFFERENCE. A first version of this test only checked that
* some decision changed somewhere, and it passed against the broken bot: with a tight cap the
* gated branches do change which Local Operations option is chosen, they just fail to stop the
* card being played two steps later. "Something moved" is not the promise. The promise is that
* committed trains never exceed what the Office can hold.
*
* Office tiers only ever go up, so the final A/D count is the most generous the cap ever was.
*/
const committed = (s: GameState): number =>
s.timetable.filter((t) => t !== null).length + s.pendingExtras.length;
for (const slack of [0, 1]) {
const capped = makeDeveloperBot({ trainCapSlack: slack });
let cappedWorst = 0;
let baseWorst = 0;
for (const seed of [1000, 8919, 16838, 24757, 32676, 40595, 48514]) {
const a = game(seed);
playGame(a, capped, pump);
const roomA = adTrackCount(a, 0) + slack;
cappedWorst = Math.max(cappedWorst, committed(a) - roomA);
const b = game(seed);
playGame(b, developerBot, pump);
baseWorst = Math.max(baseWorst, committed(b) - (adTrackCount(b, 0) + slack));
}
assert.ok(
cappedWorst <= 0,
`trainCapSlack=${slack} let the bot commit ${cappedWorst} train(s) more than the Office can hold`,
);
// And the cap is not vacuous: the uncapped bot really does overshoot on these seeds.
assert.ok(
baseWorst > 0,
`slack=${slack}: the baseline never overshoots on these seeds, so the test proves nothing`,
);
}
});
it('names itself so a report cannot confuse two runs', () => {
assert.equal(developerBot.name, 'developer');
assert.equal(makeDeveloperBot({}).name, 'developer');
assert.equal(makeDeveloperBot({ trainCapSlack: 1 }).name, 'developer+trainCapSlack=1');
});
it('refuses a tweak name it does not know', () => {
// A typo silently parsed as "no tweaks" would compare the bot against itself and report a
// confident zero — the most expensive possible failure of this tool.
assert.deepEqual(parseTweaks(['400', 'trainCapSlack=2']), { trainCapSlack: 2 });
assert.throws(() => parseTweaks(['trainCapSlok=1']), /unknown tweak/);
});
});
describe('the funnel counts what it claims to count', () => {
it('keeps every gate inside the total it is a fraction of', () => {
const s = game(430);
const probe = makeFunnelProbe(0);
const r = playGame(s, developerBot, pump, 50_000, probe.onEvent, probe.onTurn);
const f = probe.funnel;
assert.ok(f.arrivals > 0, 'no train reached an Office in a whole game');
for (const [name, n] of [
['at a passenger office', f.arrivalsAtPassengerOffice],
['with an empty coach', f.arrivalsWithEmptyCoach],
['with a loaded coach', f.arrivalsWithLoadedCoach],
['with a wanted car', f.arrivalsWithWantedCar],
] as const) {
assert.ok(n <= f.arrivals, `${name} (${n}) exceeds arrivals (${f.arrivals})`);
}
assert.ok(f.cargoPhases > 0, 'no Cargo phase was sampled');
for (const [name, n] of [
['with a facility', f.cargoWithFacility],
['green stocked', f.cargoGreenStocked],
['car spotted', f.cargoCarSpotted],
['laborer free', f.cargoLaborerFree],
['ready', f.cargoReady],
] as const) {
assert.ok(n <= f.cargoPhases, `${name} (${n}) exceeds Cargo phases (${f.cargoPhases})`);
}
// "All three at one industry" cannot exceed any of the three it is made of.
assert.ok(f.cargoReady <= Math.min(f.cargoGreenStocked, f.cargoCarSpotted, f.cargoLaborerFree));
assert.ok(f.buriedWithDigAvailable <= f.buriedTurns);
// Sampled once per Cargo phase, never once per decision: 5 Days x 12 Stages is the ceiling.
assert.ok(f.cargoPhases <= 60, `${f.cargoPhases} Cargo phases in a 5-Day game`);
// And the arrivals it counted are the arrivals that happened.
const arrived = r.events.filter((e) => e.type === 'trainArrived').length;
assert.equal(f.arrivals, arrived, 'the probe and the event log disagree about arrivals');
});
it('rides along on the harness summary', () => {
const s = game(202);
const probe = makeFunnelProbe(0);
const r = playGame(s, developerBot, pump, 50_000, probe.onEvent, probe.onTurn);
const stats = summarize(202, r.events, r.intents, s, probe.funnel);
assert.ok(stats.funnel, 'summarize dropped the funnel');
assert.equal(stats.funnel!.arrivals, probe.funnel.arrivals);
// Optional, so a caller with no probe still gets a summary.
assert.equal(summarize(202, r.events, r.intents, s).funnel, undefined);
});
});
describe('the comparison arithmetic', () => {
it('reports a dead heat as a dead heat', () => {
// Comparing the bot with itself must produce exactly zero, no games differing, and a t of 0 —
// if the machinery has any asymmetry in it, this is where it shows.
const r = compare({}, 12);
assert.equal(r.mean, 0);
assert.equal(r.identical, 12);
assert.equal(r.better, 0);
assert.equal(r.worse, 0);
assert.equal(r.t, 0);
});
it('pairs by seed, and deals the same seeds the harness does', () => {
const r = compare({ trainCapSlack: 1 }, 5);
assert.deepEqual(
r.deltas.map((d) => d.seed),
[0, 1, 2, 3, 4].map((i) => 1000 + i * 7919),
'compare and the harness must talk about the same games',
);
for (let i = 0; i < r.deltas.length; i++) {
assert.equal(r.baseStats[i]!.seed, r.deltas[i]!.seed);
assert.equal(r.variantStats[i]!.seed, r.deltas[i]!.seed);
assert.equal(r.deltas[i]!.delta, r.variantStats[i]!.revenue.net - r.baseStats[i]!.revenue.net);
}
});
it('computes the standard error from the PAIRED difference', () => {
// The whole gain of pairing is that σ of the difference is smaller than σ of either side. Using
// the level's σ by mistake would quietly restore the ±1.0 noise floor this exists to escape.
const r = compare({ trainCapSlack: 1 }, 40);
const d = r.deltas.map((x) => x.delta);
const m = d.reduce((p, c) => p + c, 0) / d.length;
const sd = Math.sqrt(d.reduce((p, c) => p + (c - m) ** 2, 0) / (d.length - 1));
assert.ok(Math.abs(r.mean - m) < 1e-9);
assert.ok(Math.abs(r.sd - sd) < 1e-9);
assert.ok(Math.abs(r.stderr - sd / Math.sqrt(d.length)) < 1e-9);
});
});