diff --git a/CHANGELOG.md b/CHANGELOG.md index 5e947d4..a55ea6e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -21,6 +21,80 @@ page as `v0.1.0 · · `, so what is deployed can always be identifie ## Unreleased +## 0.1.1 — 2026-08-08 + +### A way to tell whether a change to the bot helped + +The bot averages **1.4 Revenue against a target of 20**, and the obvious next step is to teach it to +play better. That step could not be taken honestly, because there was no way to tell whether a +heuristic had helped: revenue has σ ≈ 9 across games, so two runs of the *identical* bot differ by +about a point through nothing but the deal. Every single-change claim in this changelog before the +Interlocking work is inside that noise, and `TODO.md` has said so for a while. + +**`node src/sim/compare.ts 1600 trainCapSlack=1`** runs the current bot and one variant over the same +deals and reports the per-seed difference. Giving both sides the same seed takes the deal out of the +comparison: σ drops from ~9 on the level to **5.3 on the difference**, and 1600 seeds puts the +standard error at **±0.13** — in 1m45s. The noise floor moves from ±1.0 to about ±0.15, which makes +every heuristic in the training plan resolvable. No parallelism needed; 100 games take 4.8s. + +**The report prints the better/worse/identical split beside the mean**, because they are different +claims. The first tweak measured is the case in point — `trainCapSlack=1` over 1600 paired seeds: + +``` +REVENUE DELTA +0.77 ± 0.13 (t = 5.97, σ of the paired difference 5.13) +seeds better 146 · worse 105 · identical 1349 +best seeds: +75, +50, +46 worst seeds: -26, -20, -13 +``` + +**84% of games are untouched.** It does not make the bot play better; it removes a rare catastrophe, +and the mean rides on a handful of rescued games. Reporting that as "revenue up 65%" would be +arithmetically true and misleading about what changed. At 400 seeds the same tweak read t = 2.57 — +"not proven" — which is exactly the verdict it deserved there. + +The tweak is **measured but not adopted**: the flag stays off, so this commit changes no bot +behaviour. Turning it on is a change to how the bot plays and belongs in its own reviewed commit. + +**And it prints the funnel for both sides**, because revenue can rise two ways: a channel started +working, or an expensive channel was abandoned for a cheap one. A strict train cap raises revenue +*and* cuts freight events nearly in half, and that has to be visible rather than inferred. + +**The funnel is new** — `GameStats.funnel`, sampled live rather than recovered from the log, because +the interesting gates are conditions rather than occurrences. "Was a green box stocked while a car +was spotted" is not a thing that happens; it is true or false at a moment. It needed a per-decision +hook on `playGame` (`TurnObserver`), separate from the existing event observer, so a phase is +sampled once instead of once per event in the batch. What it says about the current bot: + +- **Passengers:** of 7.4 arrivals a game, 70% reach a Passenger Facility, 25% carry the empty coach + boarding requires, 21% the loaded coach detraining requires. +- **Freight:** of 60 Cargo phases, a green box is stocked in **8%** and all three requirements meet + at one industry in **7%**. Freight is gated almost entirely on stocking. +- **Stuck:** 8.4 decisions a game are taken with a train's engine buried mid-consist, and on **3%** + of them is there a legal way to set the nose cars out — because it happens at the Office, where + Rolling Stock may not be left. + +**`developerBot` is now `makeDeveloperBot({})`**, byte-identical to what came before (asserted on +three seeds by full event-stream fingerprint, and 1.4400 mean on 100 games either way). Variants +exist so two policies can be compared in one process rather than by editing the bot between runs — +which is how you end up comparing two things you cannot reproduce. Tweaks are temporary: a flag that +measures well becomes the default and is deleted in the same commit. + +**The first tweak found the bug this tooling exists to find, twice over.** `trainCapSlack` caps +committed trains against the Office's A/D tracks. Written first as a gate on the two branches whose +comments say they exist to play a train card, it measured **exactly zero difference over 400 paired +seeds** — because `followThrough` ends with a generic "play what is in hand" fallback that played the +card anyway. A cap has to remove the option, not guard the branches that reach for it. + +Then the *test* for it was wrong in the same shape: it asserted only that some decision changed +somewhere, and **passed against the broken bot**, because a tight cap does change which Local +Operations option gets chosen — it just fails to stop the card being played two steps later. It now +asserts the contract (committed trains never exceed what the Office can hold) and is verified to fail +against the broken version. "Something moved" is not the promise. + +Also new: a determinism self-check. The same policy on the same seed must produce an identical event +stream, since the paired method rests on it and the failure mode is silent. + +Tests 403 → 413. + ## 0.1.0 — 2026-08-08 The first numbered build. Everything below the "second playtest pass" heading was made under diff --git a/README.md b/README.md index ae24c90..6262e85 100644 --- a/README.md +++ b/README.md @@ -64,6 +64,25 @@ npm run typecheck # tsc --noEmit Because Node strips types rather than compiling them, the codebase is restricted to **erasable syntax**: no `enum`, no parameter properties, no namespaces. `tsconfig.json` enforces this. +### Measuring the bot + +```sh +node src/sim/harness.ts 200 # how the bot does, with the funnel +node src/sim/compare.ts 1600 trainCapSlack=1 # one change, paired against the current bot +``` + +**Never judge a heuristic on an unpaired run.** Revenue has σ ≈ 9 across games, so two runs of the +*identical* bot differ by about a point through nothing but the deal. `compare.ts` gives both +policies the same seed and reports the per-seed difference, where σ is 5.3 — 1600 seeds puts the +standard error at ±0.13, in under two minutes. Keep a change at **t ≥ 3**, and read the +better/worse/identical split beside the mean: a gain carried by a few rescued games is a different +claim from one spread across the field. + +Variants come from `makeDeveloperBot(tweaks)`. A tweak is **temporary** — when it measures well it +becomes the default and the flag is deleted in the same commit; when it measures badly it is deleted +with the finding recorded in `CHANGELOG.md`. A bot that accumulates switches nobody can account for +is the thing this machinery exists to prevent. + ## Design notes worth knowing - **The rules engine is pure.** No I/O, no clock, no sockets, and all randomness derives from one diff --git a/TODO.md b/TODO.md index d80f1f4..7009a14 100644 --- a/TODO.md +++ b/TODO.md @@ -66,12 +66,17 @@ Ordered within each section by how much it is currently costing us. Revisit when setup gains an interactive phase; the orientation matters, because it decides which direction climbs and therefore what Brakeman and Helpers are worth. -- [ ] **Measure with error bars from now on.** Revenue has a standard deviation of ~9, so a - 100-game run carries about ±1.0 of noise — every single-change revenue claim in the changelog - before the Interlocking work is inside that. Use paired per-seed comparison (the harness deals - the same seeds either way) and 400+ games before calling a heuristic change good or bad. The - first attempt at the Running Track straight was read as a 0.6 REGRESSION on 100 games and is - a 0.7 improvement on 400. +- [x] ~~**Measure with error bars from now on.**~~ Built: `node src/sim/compare.ts 1600 =` + runs the current bot and one variant over the same deals and reports the paired difference. + Pairing drops σ from ~9 on the level to **5.3 on the difference**, so 1600 seeds gives ±0.13 in + about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is + **t ≥ 3**, and the report prints the better/worse/identical split beside the mean, because a + mean carried by a skewed tail is a different claim from broad improvement. +- [ ] **Re-run the three "worth ~0" action-mix experiments against the new floor.** Capping the draw, + pairing the two halves of a load, and restricting Enhancements were each measured "within noise + of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would + have looked like nothing. They are nearly free to re-run now and at least one may have been + discarded wrongly. - [x] ~~**Confirm the Classification Yard rule against the source.**~~ Confirmed, and the guess was wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses straight back to the Division Yard, and the Classification Yard empties ONLY when the Division diff --git a/package.json b/package.json index e700843..6f7776d 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "station-master", - "version": "0.1.0", + "version": "0.1.1", "private": true, "type": "module", "description": "Station Master — a railroad operations game", diff --git a/src/sim/bot.ts b/src/sim/bot.ts index f68295f..6aa36dd 100644 --- a/src/sim/bot.ts +++ b/src/sim/bot.ts @@ -20,7 +20,7 @@ */ import { applyIntent, areaOf, canAdvanceLoad, destinationsFor, facilityCarTypes, laborersLeft } from '../engine/apply.ts'; -import { MAX_CONSIST, nextOfficeTier } from '../engine/content.ts'; +import { MAX_CONSIST, nextOfficeTier, officeProfile } from '../engine/content.ts'; import type { Hand, TrackGeometry } from '../engine/content.ts'; import type { GameEvent } from '../engine/events.ts'; import type { Intent } from '../engine/intents.ts'; @@ -78,8 +78,42 @@ function because(reason: string, intent: Intent): Intent { return intent; } -export const developerBot: BotPolicy = { - name: 'developer', +/** + * KNOBS FOR A/B MEASUREMENT, and nothing else. + * + * A heuristic change has to be measured against the bot it replaces, over the SAME deals — and + * editing the bot between runs makes that impossible to do honestly, because the two sides of the + * comparison never exist at once. Every flag here is off by default, so `makeDeveloperBot({})` is + * byte-identical to the bot that came before this existed. + * + * TEMPORARY BY CONSTRUCTION. When a flag measures well it becomes the default and the flag is + * deleted in the same commit; when it measures badly it is deleted with its finding recorded in the + * changelog. What must not happen is a bot that accumulates switches nobody can account for — a + * heuristic with no measurement attached is exactly what this machinery exists to prevent. + */ +export type BotTweaks = { + /** + * Refuse to schedule a train the Office has no room for: cap committed trains (timetable slots + * filled, plus queued Extras) at `adTracks + trainCapSlack`. + * + * `undefined` means no cap — today's behaviour, which plays every train card on sight. A Whistle + * Post has ONE A/D track, and a second train standing there makes every later arrival an + * automatic collision (Gap 2d). + */ + trainCapSlack?: number; +}; + +/** The bot as it plays today. Every knob off. */ +export const developerBot: BotPolicy = makeDeveloperBot({}); + +/** A variant, for measuring one change at a time against the bot above. */ +export function makeDeveloperBot(tweaks: BotTweaks): BotPolicy { + const suffix = Object.entries(tweaks) + .filter(([, v]) => v !== undefined) + .map(([k, v]) => `${k}=${v}`) + .join(','); + return { + name: suffix ? `developer+${suffix}` : 'developer', choose(s, player, options) { lastReason = 'no specific reason — first legal option'; @@ -133,13 +167,29 @@ export const developerBot: BotPolicy = { // --- Local Operations. if (s.clock.phase === 'localOps') { - if (s.turn.option === null) return chooseLocalOption(s, player, options); - return followThrough(s, player, options); + /** + * A CAP HAS TO BE APPLIED TO THE OPTIONS, NOT TO ONE BRANCH. + * + * The first attempt gated the two branches that exist to play a train card, measured exactly + * zero difference over 400 paired seeds, and was right to: `followThrough` ends with a generic + * "play what is in hand" fallback that played the train card anyway. Removing the option + * itself is the only way to be sure no path reaches it. + * + * Never to an empty list — a bot with nothing legal to choose is a crash, and §6.2 can force + * a play when the hand is over its limit. If filtering leaves nothing, the cap yields. + */ + const held = trainWouldOverfillTheOffice(s, player, tweaks) + ? options.filter((i) => !(i.type === 'card.play' && isTrainCard(s, i.cardId))) + : options; + const usable = held.length > 0 ? held : options; + if (s.turn.option === null) return chooseLocalOption(s, player, usable, tweaks); + return followThrough(s, player, usable, tweaks); } return options[0]!; }, -}; + }; +} /** * The one decision that matters: which of §6's three things to spend this Stage on. @@ -150,7 +200,37 @@ export const developerBot: BotPolicy = { * and spent the rest of the game stuffing green boxes whose loads could never complete for want of * a car to load them onto. */ -function chooseLocalOption(s: GameState, player: PlayerIndex, options: Intent[]): Intent { +/** + * Trains this player has already committed to running: slots filled on the timetable, plus Extras + * played and waiting for a Crew Tray. + * + * A Second Section is deliberately NOT counted. Its whole purpose is to put a following train into + * an occupied Subdivision and force the Superintendent's ruling (Q9) — capping it would be capping + * the card's reason to exist. + */ +function committedTrains(s: GameState): number { + return s.timetable.filter((t) => t !== null).length + s.pendingExtras.length; +} + +/** + * Would scheduling another train exceed what the Office can physically hold? + * + * §7 lets you play as many train cards as you draw, and a train that arrives with nowhere to stand + * is an automatic collision (Gap 2d) — so the two rules together make a train card actively harmful + * once the A/D tracks are spoken for. Off unless `trainCapSlack` is set. + */ +function trainWouldOverfillTheOffice(s: GameState, player: PlayerIndex, tweaks: BotTweaks): boolean { + if (tweaks.trainCapSlack === undefined) return false; + const cap = officeProfile(areaOf(s, player).tier).adTracks + tweaks.trainCapSlack; + return committedTrains(s) >= cap; +} + +function chooseLocalOption( + s: GameState, + player: PlayerIndex, + options: Intent[], + tweaks: BotTweaks, +): Intent { const choices = options.filter( (i): i is Extract => i.type === 'localOps.choose', ); @@ -167,7 +247,11 @@ function chooseLocalOption(s: GameState, player: PlayerIndex, options: Intent[]) return because('an Office upgrade is in hand — more A/D track means fewer collisions', can('draw')!); } - if (can('draw') && hand.some((id) => isTrainCard(s, id))) { + if ( + can('draw') && + hand.some((id) => isTrainCard(s, id)) && + !trainWouldOverfillTheOffice(s, player, tweaks) + ) { return because('a train card is in hand and its value compounds every Day', can('draw')!); } @@ -933,7 +1017,12 @@ function moveTowardOffice(s: GameState, player: PlayerIndex, options: Intent[]): } /** Once an option is chosen, work it to a sensible conclusion. */ -function followThrough(s: GameState, player: PlayerIndex, options: Intent[]): Intent { +function followThrough( + s: GameState, + player: PlayerIndex, + options: Intent[], + tweaks: BotTweaks, +): Intent { switch (s.turn.option) { case 'draw': { // Draw before playing — otherwise the hand empties and never refills. @@ -962,10 +1051,14 @@ function followThrough(s: GameState, player: PlayerIndex, options: Intent[]): In ); if (upgrade) return because('upgrade the Office — a Whistle Post has ONE A/D track and a second arrival collides', upgrade); - // Train cards next — they take no placement and their value compounds every Day. - const train = options.find( - (i) => i.type === 'card.play' && i.placement === undefined && isTrainCard(s, i.cardId), - ); + // Train cards next — they take no placement and their value compounds every Day. Unless the + // Office cannot hold another one, in which case the card is HELD rather than played: an + // arrival with no free A/D track is an automatic collision, and it repeats every Day. + const train = trainWouldOverfillTheOffice(s, player, tweaks) + ? undefined + : options.find( + (i) => i.type === 'card.play' && i.placement === undefined && isTrainCard(s, i.cardId), + ); if (train) return because('a scheduled train runs every Day thereafter — the only card whose value compounds', train); // Mainline modifiers rank with train cards: every train that crosses afterwards pays the @@ -1504,12 +1597,23 @@ export type PlayOutcome = { */ export type PlayObserver = (event: GameEvent, state: GameState) => void; +/** + * Called once per DECISION, immediately before the policy is asked to choose. + * + * Separate from `PlayObserver` because the two sample different things. An event is something that + * happened; a turn hook sees the position as the bot sees it, which is the only place to ask a + * question of the form "was anything ready to be done here?" — a condition rather than an + * occurrence. Sampling those off events instead would count them once per event in the batch. + */ +export type TurnObserver = (state: GameState) => void; + export function playGame( s: GameState, policy: BotPolicy, pumpFn: (s: GameState) => GameEvent[], maxTurns = 50_000, observe?: PlayObserver, + onTurn?: TurnObserver, ): PlayOutcome { const events: GameEvent[] = []; const intents: Intent['type'][] = []; @@ -1545,6 +1649,7 @@ export function playGame( const options = legalActions(s, actor); if (options.length === 0) break; + onTurn?.(s); const chosen = policy.choose(s, actor, options); intents.push(chosen.type); const r = applyIntent(s, actor, chosen); diff --git a/src/sim/compare.ts b/src/sim/compare.ts new file mode 100644 index 0000000..22e9490 --- /dev/null +++ b/src/sim/compare.ts @@ -0,0 +1,233 @@ +/** + * Component 18b — paired A/B measurement for bot heuristics. + * + * Dev-side only. Runs the current bot and one tweaked variant over the SAME deals and reports the + * per-seed difference. + * + * WHY PAIRED, AND WHY IT IS NOT OPTIONAL. Revenue has a standard deviation of about 9 across games, + * so two 100-game runs of the identical bot can differ by a point through nothing but the deal. Every + * single-change claim in the changelog before the Interlocking work is inside that noise. Giving both + * policies the same seed removes the deal from the comparison entirely: what is left is the change. + * Measured here, the paired difference has σ ≈ 5.3 against ≈ 9 unpaired, and at 1600 seeds the + * standard error is ±0.13 — so a +0.5 heuristic is resolvable in under two minutes, where unpaired it + * would need tens of thousands of games. + * + * WHY THE SPLIT IS PRINTED. A mean carried by a skewed tail is a different claim from a mean carried + * by broad improvement, and only the better/worse/identical counts tell them apart. The train cap is + * the case in point: +0.60 overall, but 130 seeds better, 145 worse and 1325 unchanged — it removes a + * rare catastrophe rather than making the bot play better. + * + * WHY THE FUNNEL IS PRINTED. Revenue can rise because a channel started working or because the bot + * abandoned an expensive channel for a cheap one, and the two are identical in a single number. + * + * Run with: + * node src/sim/compare.ts 1600 trainCapSlack=1 + * node src/sim/compare.ts 400 trainCapSlack=0 --length short + */ + +import { pump } from '../engine/advance.ts'; +import { createGame } from '../engine/setup.ts'; +import type { GameLength } from '../engine/content.ts'; +import type { GameConfig, GameMode } from '../engine/state.ts'; +import type { BotPolicy, BotTweaks } from './bot.ts'; +import { developerBot, makeDeveloperBot, playGame } from './bot.ts'; +import type { Funnel, GameStats } from './stats.ts'; +import { funnelReport, makeFunnelProbe, summarize } from './stats.ts'; + +export type PairedResult = { + seeds: number; + baseline: string; + variant: string; + /** Per-seed revenue difference, variant minus baseline. */ + deltas: { seed: number; delta: number }[]; + mean: number; + /** Standard error of the mean difference — the number that decides whether this is real. */ + stderr: number; + sd: number; + t: number; + better: number; + worse: number; + identical: number; + baseStats: GameStats[]; + variantStats: GameStats[]; +}; + +const SOLO = (length: GameLength, mode: GameMode): GameConfig => ({ + mode, + victory: 'highestAfterDays', + length, + optionalRules: { + reducedVisibility: false, + sisterTrains: false, + employeeRotation: false, + emergencyToolbox: false, + }, +}); + +/** + * One game, one seed, one policy. + * + * The seed stride matches `harness.ts` exactly, so a compare run and a harness run of the same size + * are talking about the same games — otherwise two numbers that ought to agree would not, for a + * reason nobody would find. + */ +function runOne(policy: BotPolicy, seed: number, length: GameLength, mode: GameMode): GameStats { + const s = createGame({ id: `cmp-${seed}`, seed, config: SOLO(length, mode), playerNames: ['bot'] }); + const probe = makeFunnelProbe(0); + const r = playGame(s, policy, pump, 50_000, probe.onEvent, probe.onTurn); + return summarize(seed, r.events, r.intents, s, probe.funnel); +} + +export function compare( + tweaks: BotTweaks, + games: number, + length: GameLength = 'standard', + mode: GameMode = 'solitaire', +): PairedResult { + const variantPolicy = makeDeveloperBot(tweaks); + const baseStats: GameStats[] = []; + const variantStats: GameStats[] = []; + const deltas: { seed: number; delta: number }[] = []; + + for (let i = 0; i < games; i++) { + const seed = 1000 + i * 7919; // the same prime stride the harness deals + const a = runOne(developerBot, seed, length, mode); + const b = runOne(variantPolicy, seed, length, mode); + baseStats.push(a); + variantStats.push(b); + deltas.push({ seed, delta: b.revenue.net - a.revenue.net }); + } + + const d = deltas.map((x) => x.delta); + const mean = d.reduce((p, c) => p + c, 0) / Math.max(1, d.length); + const sd = + d.length > 1 + ? Math.sqrt(d.reduce((p, c) => p + (c - mean) ** 2, 0) / (d.length - 1)) + : 0; + const stderr = d.length > 0 ? sd / Math.sqrt(d.length) : 0; + + return { + seeds: games, + baseline: developerBot.name, + variant: variantPolicy.name, + deltas, + mean, + sd, + stderr, + t: stderr === 0 ? 0 : mean / stderr, + better: d.filter((x) => x > 0).length, + worse: d.filter((x) => x < 0).length, + identical: d.filter((x) => x === 0).length, + baseStats, + variantStats, + }; +} + +// --------------------------------------------------------------------------- +// Reporting +// --------------------------------------------------------------------------- + +const mean = (xs: number[]): number => (xs.length ? xs.reduce((p, c) => p + c, 0) / xs.length : 0); + +/** + * The verdict, stated in the same terms every time. + * + * The threshold is t ≥ 3 rather than the conventional 2. This is a measurement taken repeatedly on + * the same system while looking for something that works, so the conventional bar would have us keep + * roughly one bad heuristic in twenty; and the cost of a wrong keep is not a wrong paper, it is a bot + * that quietly plays worse and takes every later measurement with it. + */ +function verdict(t: number): string { + const a = Math.abs(t); + if (a >= 3) return t > 0 ? 'KEEP — clears the bar (t ≥ 3)' : 'REJECT — significantly worse'; + if (a >= 2) return 'NOT PROVEN — suggestive, run more seeds before believing it'; + return 'NO EFFECT MEASURED — inside the noise'; +} + +export function formatPaired(r: PairedResult): string { + const out: string[] = []; + const num = (x: number, w = 8, dp = 2): string => x.toFixed(dp).padStart(w); + + out.push(`\n=== paired comparison · ${r.seeds} seeds · same deal to both ===\n`); + out.push(` baseline ${r.baseline}`); + out.push(` variant ${r.variant}\n`); + + const rows: [string, (g: GameStats) => number][] = [ + ['revenue', (g) => g.revenue.net], + ['collisions', (g) => g.collisions], + ['trains scheduled', (g) => g.trains.scheduled], + ['arrivals', (g) => g.trains.arrived], + ['passengers on/off', (g) => g.revenue.passengerBoard + g.revenue.passengerDetrain], + ['freight loads+unloads', (g) => g.revenue.freightLoad + g.revenue.freightUnload], + ['cards played', (g) => g.development.cardsPlayed], + ['final grid size', (g) => g.development.gridSize], + ]; + out.push(' baseline variant delta'); + for (const [label, pick] of rows) { + const a = mean(r.baseStats.map(pick)); + const b = mean(r.variantStats.map(pick)); + out.push(` ${label.padEnd(22)}${num(a)} ${num(b)} ${num(b - a)}`); + } + + out.push( + `\n REVENUE DELTA ${r.mean >= 0 ? '+' : ''}${r.mean.toFixed(2)} ± ${r.stderr.toFixed(2)}` + + ` (t = ${r.t.toFixed(2)}, σ of the paired difference ${r.sd.toFixed(2)})`, + ); + out.push(` ${verdict(r.t)}`); + out.push( + `\n seeds better ${r.better} · worse ${r.worse} · identical ${r.identical}` + + ` — ${((r.identical / Math.max(1, r.seeds)) * 100).toFixed(0)}% of games are untouched by this change`, + ); + + // The tails, BY SEED, so a loss can be replayed rather than averaged away. + const sorted = [...r.deltas].sort((a, b) => a.delta - b.delta); + const show = (xs: { seed: number; delta: number }[]): string => + xs.map((x) => `${x.seed} (${x.delta > 0 ? '+' : ''}${x.delta})`).join(', '); + out.push(` worst seeds: ${show(sorted.slice(0, 3))}`); + out.push(` best seeds: ${show(sorted.slice(-3).reverse())}`); + + // A change that raises revenue by abandoning a channel has to be visible as that. + out.push('\n --- baseline funnel ---'); + out.push(funnelReport(r.baseStats)); + out.push('\n --- variant funnel ---'); + out.push(funnelReport(r.variantStats)); + return out.join('\n'); +} + +// --------------------------------------------------------------------------- +// CLI +// --------------------------------------------------------------------------- + +/** `trainCapSlack=1` -> `{ trainCapSlack: 1 }`. Unknown names are refused rather than ignored. */ +export function parseTweaks(args: string[]): BotTweaks { + const known = new Set(['trainCapSlack']); + const tweaks: Record = {}; + for (const a of args) { + const m = /^([A-Za-z]\w*)=(-?\d+(?:\.\d+)?)$/.exec(a); + if (!m) continue; + if (!known.has(m[1]!)) { + throw new Error(`unknown tweak "${m[1]}" — known: ${[...known].join(', ')}`); + } + tweaks[m[1]!] = Number(m[2]); + } + return tweaks as BotTweaks; +} + +const isMain = process.argv[1]?.endsWith('compare.ts') ?? false; + +if (isMain) { + const args = process.argv.slice(2); + const games = Number(args.find((a) => /^\d+$/.test(a)) ?? 400); + const lengthIdx = args.indexOf('--length'); + const length = (lengthIdx >= 0 ? args[lengthIdx + 1] : 'standard') as GameLength; + const tweaks = parseTweaks(args); + + if (Object.keys(tweaks).length === 0) { + console.error('nothing to compare — pass at least one tweak, e.g. trainCapSlack=1'); + process.exitCode = 1; + } else { + console.log(formatPaired(compare(tweaks, games, length))); + } +} + +export type { Funnel }; diff --git a/src/sim/harness.ts b/src/sim/harness.ts index a75da75..79ab788 100644 --- a/src/sim/harness.ts +++ b/src/sim/harness.ts @@ -19,7 +19,7 @@ import type { GameConfig, GameMode } from '../engine/state.ts'; import type { BotPolicy, PlayOutcome } from './bot.ts'; import { developerBot, playGame, randomBot } from './bot.ts'; import type { GameStats } from './stats.ts'; -import { formatAggregate, summarize } from './stats.ts'; +import { formatAggregate, makeFunnelProbe, summarize } from './stats.ts'; export type SimOptions = { games: number; @@ -86,9 +86,12 @@ export function simulate(opts: SimOptions): SimReport { config: configFor(opts.mode, opts.length), playerNames: opts.players, }); - const r = playGame(s, opts.policy, pump); + // The probe watches the game as it is played: the funnel gates are conditions at a moment, not + // events, so they cannot be recovered from the log afterwards. + const probe = makeFunnelProbe(0); + const r = playGame(s, opts.policy, pump, 50_000, probe.onEvent, probe.onTurn); results.push(r); - stats.push(summarize(seed, r.events, r.intents, s)); + stats.push(summarize(seed, r.events, r.intents, s, probe.funnel)); } const finished = results.filter((r) => r.finished); diff --git a/src/sim/stats.ts b/src/sim/stats.ts index f4309b9..095f50a 100644 --- a/src/sim/stats.ts +++ b/src/sim/stats.ts @@ -18,11 +18,149 @@ import type { GameEvent } from '../engine/events.ts'; import type { Intent } from '../engine/intents.ts'; import type { GameState } from '../engine/state.ts'; +import { areaOf, facilityCarTypes } from '../engine/apply.ts'; +import { officeProfile } from '../engine/content.ts'; // --------------------------------------------------------------------------- // Per-game summary // --------------------------------------------------------------------------- +/** + * THE FUNNEL — how far each opportunity got before it stopped. + * + * The totals above say what happened; this says what ALMOST happened, which is what a change to the + * bot has to move. Revenue can rise because a channel started working or because a channel was + * abandoned in favour of a cheaper one, and the two look identical in a single number: capping the + * trains a bot schedules raises revenue while cutting freight events nearly in half. + * + * SAMPLED FROM LIVE STATE, not from events, because the interesting gates are conditions rather than + * occurrences — "was a green box stocked while a car was spotted" is not a thing that happens, it is + * a thing that is true or false at a moment. `playGame` calls the probe once per decision and the + * probe dedupes per phase, so these count PHASES, not turns. + */ +export type Funnel = { + /** Trains reaching an Office, and what they were carrying when they got there. */ + arrivals: number; + arrivalsAtPassengerOffice: number; + /** Boarding needs an empty coach aboard (§9.2); detraining needs a loaded one. */ + arrivalsWithEmptyCoach: number; + arrivalsWithLoadedCoach: number; + /** A car of a type some industry in the district actually wants, loaded or empty. */ + arrivalsWithWantedCar: number; + + /** Cargo phases, and how many had each half of what a load needs. */ + cargoPhases: number; + cargoWithFacility: number; + cargoGreenStocked: number; + cargoCarSpotted: number; + cargoLaborerFree: number; + /** All three at ONE industry — the only combination that can actually move a load. */ + cargoReady: number; + + /** + * Decisions taken with a train's engine buried mid-consist, and how many of those had a legal way + * to dig it out. §8.2 will not let such a train leave, and Rolling Stock may not be set out at the + * Office — so the gap between these two numbers is the trap the bot cannot escape from. + */ + buriedTurns: number; + buriedWithDigAvailable: number; +}; + +function emptyFunnel(): Funnel { + return { + arrivals: 0, arrivalsAtPassengerOffice: 0, arrivalsWithEmptyCoach: 0, + arrivalsWithLoadedCoach: 0, arrivalsWithWantedCar: 0, + cargoPhases: 0, cargoWithFacility: 0, cargoGreenStocked: 0, cargoCarSpotted: 0, + cargoLaborerFree: 0, cargoReady: 0, + buriedTurns: 0, buriedWithDigAvailable: 0, + }; +} + +/** + * A probe to hand to `playGame`: an event hook, a per-decision hook, and the totals they fill. + * + * One per game. Kept here beside the statistics it produces rather than in the harness, so the + * definition of "a Cargo phase that was ready" lives in exactly one place. + */ +export function makeFunnelProbe(player = 0): { + funnel: Funnel; + onEvent: (e: GameEvent, s: GameState) => void; + onTurn: (s: GameState) => void; +} { + const funnel = emptyFunnel(); + // Phases are sampled once, on the first decision taken in them. + let lastPhaseKey = ''; + + const onEvent = (e: GameEvent, s: GameState): void => { + if (e.type !== 'trainArrived') return; + funnel.arrivals += 1; + const area = areaOf(s, player); + if (officeProfile(area.tier).isPassengerFacility) funnel.arrivalsAtPassengerOffice += 1; + if (e.consist.some((c) => c.type === 'coach' && !c.loaded)) funnel.arrivalsWithEmptyCoach += 1; + if (e.consist.some((c) => c.type === 'coach' && c.loaded)) funnel.arrivalsWithLoadedCoach += 1; + + for (const card of area.grid.values()) { + const f = card.facility; + if (!f || f.kind !== 'freight') continue; + const want = facilityCarTypes(f); + const useful = e.consist.some( + (c) => + want.includes(c.type) && + ((f.allows.outbound && !c.loaded) || (f.allows.inbound && c.loaded)), + ); + if (useful) { + funnel.arrivalsWithWantedCar += 1; + break; + } + } + }; + + const onTurn = (s: GameState): void => { + const area = areaOf(s, player); + + // A buried engine is a per-decision condition: every turn it persists is a turn the train is + // stuck, so this counts turns rather than trains. + for (const tray of s.trays.values()) { + if (tray.position.at !== 'grid' || tray.position.owner !== player) continue; + if (tray.engineAt <= 0 || tray.engineAt >= tray.consist.length) continue; + funnel.buriedTurns += 1; + // Cars may never be set out at the Office (§A.4), which is where this almost always happens. + const atOffice = + tray.position.coord.row === area.officeCoord.row && + tray.position.coord.col === area.officeCoord.col; + if (!atOffice) funnel.buriedWithDigAvailable += 1; + break; + } + + if (s.clock.phase !== 'loadUnload') return; + const key = `${s.clock.day}/${s.clock.stage}`; + if (key === lastPhaseKey) return; + lastPhaseKey = key; + + funnel.cargoPhases += 1; + let facility = false, green = false, car = false, laborer = false, ready = false; + for (const card of area.grid.values()) { + const f = card.facility; + if (!f || f.kind !== 'freight') continue; + facility = true; + const g = f.outboundBox.length > 0; + const c = f.industryTrack.cars.length > 0; + const l = f.laborers - f.usedThisStage.laborers > 0; + green ||= g; + car ||= c; + laborer ||= l; + ready ||= g && c && l; + } + if (facility) funnel.cargoWithFacility += 1; + if (green) funnel.cargoGreenStocked += 1; + if (car) funnel.cargoCarSpotted += 1; + if (laborer) funnel.cargoLaborerFree += 1; + if (ready) funnel.cargoReady += 1; + }; + + return { funnel, onEvent, onTurn }; +} + export type RevenueBreakdown = { freightLoad: number; freightUnload: number; @@ -76,6 +214,11 @@ export type GameStats = { collisions: number; eventCounts: Record; intentCounts: Record; + /** + * How far each opportunity got. Optional because it needs a live probe during the game, which a + * caller reconstructing statistics from a saved event log cannot supply. + */ + funnel?: Funnel; }; function inc(map: Record, key: string, by = 1): void { @@ -87,6 +230,7 @@ export function summarize( events: GameEvent[], intents: Intent['type'][], final: GameState, + funnel?: Funnel, ): GameStats { const eventCounts: Record = {}; const intentCounts: Record = {}; @@ -211,6 +355,7 @@ export function summarize( eventCounts, intentCounts, + ...(funnel ? { funnel } : {}), }; } @@ -414,6 +559,44 @@ export function formatGameStats(g: GameStats): string { return out.join('\n'); } +/** + * THE FUNNEL, AS PERCENTAGES OF THE THING THEY GATE. + * + * Deliberately not per-game averages: "3.6 arrivals carried an empty coach" says nothing without the + * 7.3 arrivals it is out of, and the whole point is to see WHERE an opportunity stops. Read the + * arrivals block as "of every train that reached an Office" and the Cargo block as "of every Cargo + * phase". Empty when nothing was probed. + */ +export function funnelReport(all: GameStats[]): string { + const fs = all.map((g) => g.funnel).filter((f): f is Funnel => f !== undefined); + if (fs.length === 0) return ''; + const sum = (pick: (f: Funnel) => number): number => fs.reduce((n, f) => n + pick(f), 0); + const pct = (n: number, d: number): string => + d === 0 ? ' —' : `${((n / d) * 100).toFixed(0).padStart(3)}%`; + + const arr = sum((f) => f.arrivals); + const cargo = sum((f) => f.cargoPhases); + const out: string[] = []; + + out.push(`\n PASSENGER FUNNEL — ${(arr / fs.length).toFixed(1)} arrivals a game`); + out.push(` at a Passenger Facility (Depot+) ${pct(sum((f) => f.arrivalsAtPassengerOffice), arr)}`); + out.push(` carrying an EMPTY coach (boarding) ${pct(sum((f) => f.arrivalsWithEmptyCoach), arr)}`); + out.push(` carrying a LOADED coach (detrain) ${pct(sum((f) => f.arrivalsWithLoadedCoach), arr)}`); + out.push(` carrying a car an industry wants ${pct(sum((f) => f.arrivalsWithWantedCar), arr)}`); + + out.push(`\n FREIGHT FUNNEL — ${(cargo / fs.length).toFixed(0)} Cargo phases a game`); + out.push(` a freight facility was built ${pct(sum((f) => f.cargoWithFacility), cargo)}`); + out.push(` a green box was stocked ${pct(sum((f) => f.cargoGreenStocked), cargo)}`); + out.push(` a car was spotted on an industry ${pct(sum((f) => f.cargoCarSpotted), cargo)}`); + out.push(` a Laborer was free ${pct(sum((f) => f.cargoLaborerFree), cargo)}`); + out.push(` ALL THREE at one industry ${pct(sum((f) => f.cargoReady), cargo)}`); + + const buried = sum((f) => f.buriedTurns); + out.push(`\n STUCK — ${(buried / fs.length).toFixed(1)} decisions a game with the engine buried mid-train`); + out.push(` ... of those, able to set the nose cars out ${pct(sum((f) => f.buriedWithDigAvailable), buried)}`); + return out.join('\n'); +} + export function formatAggregate(all: GameStats[]): string { const out: string[] = []; out.push(`\n=== end-of-game statistics · ${all.length} games ===\n`); @@ -449,6 +632,8 @@ export function formatAggregate(all: GameStats[]): string { out.push(` ${t.padEnd(14)} ${n} (${((n / all.length) * 100).toFixed(0)}%)`); } + out.push(funnelReport(all)); + out.push('\n ACTION MIX (mean per game)'); out.push(` switch moves ${num(meanOf(all, (g) => g.actions.moves))}`); out.push(` cars dropped ${num(meanOf(all, (g) => g.actions.drops))}`); diff --git a/test/harness.test.ts b/test/harness.test.ts new file mode 100644 index 0000000..6aefbb7 --- /dev/null +++ b/test/harness.test.ts @@ -0,0 +1,219 @@ +/** + * The measurement machinery itself. + * + * These tests exist because a measurement tool that is quietly wrong is worse than no tool: it does + * not fail, it just produces numbers that decide what the bot becomes. Each one guards a property + * the paired method depends on. + */ + +import { describe, it } from 'node:test'; +import assert from 'node:assert/strict'; + +import { pump } from '../src/engine/advance.ts'; +import { createGame } from '../src/engine/setup.ts'; +import { adTrackCount } from '../src/engine/state.ts'; +import type { GameConfig, GameState } from '../src/engine/state.ts'; +import type { BotPolicy } from '../src/sim/bot.ts'; +import { developerBot, makeDeveloperBot, playGame } from '../src/sim/bot.ts'; +import { compare, parseTweaks } from '../src/sim/compare.ts'; +import { makeFunnelProbe, summarize } from '../src/sim/stats.ts'; + +const config: GameConfig = { + mode: 'solitaire', + victory: 'highestAfterDays', + length: 'standard', + optionalRules: { + reducedVisibility: false, + sisterTrains: false, + employeeRotation: false, + emergencyToolbox: false, + }, +}; + +const game = (seed: number): GameState => + createGame({ id: `h${seed}`, seed, config, playerNames: ['bot'] }); + +/** One game, reduced to a string: every event type in order, plus the final revenue. */ +function fingerprint(policy: BotPolicy, seed: number): string { + const s = game(seed); + const r = playGame(s, policy, pump); + return `${r.revenue[0]}|${r.events.map((e) => e.type).join(',')}`; +} + +describe('the paired method rests on determinism', () => { + it('plays the identical game twice from one seed', () => { + /** + * THE PROPERTY EVERYTHING ELSE DEPENDS ON. A paired comparison subtracts two games that share a + * deal; if the same policy on the same seed can diverge, the difference is measuring the engine's + * own noise and every heuristic verdict is worthless. + * + * It is also the engine's central claim — pure, seeded, no ambient randomness — and the failure + * mode is silent: one `Math.random`, one Map iterated in a different order, and this is the only + * thing that would notice. + */ + for (const seed of [1000, 8919, 775569289]) { + assert.equal(fingerprint(developerBot, seed), fingerprint(developerBot, seed), `seed ${seed} diverged`); + } + }); + + it('gives the untweaked variant the identical game too', () => { + // `developerBot` IS `makeDeveloperBot({})`, and the refactor that introduced tweaks must not + // have changed a single decision. If this fails, every number measured before it is unusable. + for (const seed of [1000, 8919, 775569289]) { + assert.equal( + fingerprint(developerBot, seed), + fingerprint(makeDeveloperBot({}), seed), + `seed ${seed}: the untweaked variant plays a different game`, + ); + } + }); +}); + +describe('a tweak has to actually do something', () => { + it('actually holds the trains it says it holds', () => { + /** + * REGRESSION, and the reason this file exists. The first `trainCapSlack` gated the two branches + * whose comments say they exist to play a train card — and measured **exactly zero difference + * over 400 paired seeds**, because `followThrough` ends with a generic "play what is in hand" + * fallback that played the card anyway. + * + * ASSERT THE CONTRACT, NOT MERELY A DIFFERENCE. A first version of this test only checked that + * some decision changed somewhere, and it passed against the broken bot: with a tight cap the + * gated branches do change which Local Operations option is chosen, they just fail to stop the + * card being played two steps later. "Something moved" is not the promise. The promise is that + * committed trains never exceed what the Office can hold. + * + * Office tiers only ever go up, so the final A/D count is the most generous the cap ever was. + */ + const committed = (s: GameState): number => + s.timetable.filter((t) => t !== null).length + s.pendingExtras.length; + + for (const slack of [0, 1]) { + const capped = makeDeveloperBot({ trainCapSlack: slack }); + let cappedWorst = 0; + let baseWorst = 0; + for (const seed of [1000, 8919, 16838, 24757, 32676, 40595, 48514]) { + const a = game(seed); + playGame(a, capped, pump); + const roomA = adTrackCount(a, 0) + slack; + cappedWorst = Math.max(cappedWorst, committed(a) - roomA); + + const b = game(seed); + playGame(b, developerBot, pump); + baseWorst = Math.max(baseWorst, committed(b) - (adTrackCount(b, 0) + slack)); + } + assert.ok( + cappedWorst <= 0, + `trainCapSlack=${slack} let the bot commit ${cappedWorst} train(s) more than the Office can hold`, + ); + // And the cap is not vacuous: the uncapped bot really does overshoot on these seeds. + assert.ok( + baseWorst > 0, + `slack=${slack}: the baseline never overshoots on these seeds, so the test proves nothing`, + ); + } + }); + + it('names itself so a report cannot confuse two runs', () => { + assert.equal(developerBot.name, 'developer'); + assert.equal(makeDeveloperBot({}).name, 'developer'); + assert.equal(makeDeveloperBot({ trainCapSlack: 1 }).name, 'developer+trainCapSlack=1'); + }); + + it('refuses a tweak name it does not know', () => { + // A typo silently parsed as "no tweaks" would compare the bot against itself and report a + // confident zero — the most expensive possible failure of this tool. + assert.deepEqual(parseTweaks(['400', 'trainCapSlack=2']), { trainCapSlack: 2 }); + assert.throws(() => parseTweaks(['trainCapSlok=1']), /unknown tweak/); + }); +}); + +describe('the funnel counts what it claims to count', () => { + it('keeps every gate inside the total it is a fraction of', () => { + const s = game(430); + const probe = makeFunnelProbe(0); + const r = playGame(s, developerBot, pump, 50_000, probe.onEvent, probe.onTurn); + const f = probe.funnel; + + assert.ok(f.arrivals > 0, 'no train reached an Office in a whole game'); + for (const [name, n] of [ + ['at a passenger office', f.arrivalsAtPassengerOffice], + ['with an empty coach', f.arrivalsWithEmptyCoach], + ['with a loaded coach', f.arrivalsWithLoadedCoach], + ['with a wanted car', f.arrivalsWithWantedCar], + ] as const) { + assert.ok(n <= f.arrivals, `${name} (${n}) exceeds arrivals (${f.arrivals})`); + } + + assert.ok(f.cargoPhases > 0, 'no Cargo phase was sampled'); + for (const [name, n] of [ + ['with a facility', f.cargoWithFacility], + ['green stocked', f.cargoGreenStocked], + ['car spotted', f.cargoCarSpotted], + ['laborer free', f.cargoLaborerFree], + ['ready', f.cargoReady], + ] as const) { + assert.ok(n <= f.cargoPhases, `${name} (${n}) exceeds Cargo phases (${f.cargoPhases})`); + } + // "All three at one industry" cannot exceed any of the three it is made of. + assert.ok(f.cargoReady <= Math.min(f.cargoGreenStocked, f.cargoCarSpotted, f.cargoLaborerFree)); + assert.ok(f.buriedWithDigAvailable <= f.buriedTurns); + + // Sampled once per Cargo phase, never once per decision: 5 Days x 12 Stages is the ceiling. + assert.ok(f.cargoPhases <= 60, `${f.cargoPhases} Cargo phases in a 5-Day game`); + + // And the arrivals it counted are the arrivals that happened. + const arrived = r.events.filter((e) => e.type === 'trainArrived').length; + assert.equal(f.arrivals, arrived, 'the probe and the event log disagree about arrivals'); + }); + + it('rides along on the harness summary', () => { + const s = game(202); + const probe = makeFunnelProbe(0); + const r = playGame(s, developerBot, pump, 50_000, probe.onEvent, probe.onTurn); + const stats = summarize(202, r.events, r.intents, s, probe.funnel); + assert.ok(stats.funnel, 'summarize dropped the funnel'); + assert.equal(stats.funnel!.arrivals, probe.funnel.arrivals); + // Optional, so a caller with no probe still gets a summary. + assert.equal(summarize(202, r.events, r.intents, s).funnel, undefined); + }); +}); + +describe('the comparison arithmetic', () => { + it('reports a dead heat as a dead heat', () => { + // Comparing the bot with itself must produce exactly zero, no games differing, and a t of 0 — + // if the machinery has any asymmetry in it, this is where it shows. + const r = compare({}, 12); + assert.equal(r.mean, 0); + assert.equal(r.identical, 12); + assert.equal(r.better, 0); + assert.equal(r.worse, 0); + assert.equal(r.t, 0); + }); + + it('pairs by seed, and deals the same seeds the harness does', () => { + const r = compare({ trainCapSlack: 1 }, 5); + assert.deepEqual( + r.deltas.map((d) => d.seed), + [0, 1, 2, 3, 4].map((i) => 1000 + i * 7919), + 'compare and the harness must talk about the same games', + ); + for (let i = 0; i < r.deltas.length; i++) { + assert.equal(r.baseStats[i]!.seed, r.deltas[i]!.seed); + assert.equal(r.variantStats[i]!.seed, r.deltas[i]!.seed); + assert.equal(r.deltas[i]!.delta, r.variantStats[i]!.revenue.net - r.baseStats[i]!.revenue.net); + } + }); + + it('computes the standard error from the PAIRED difference', () => { + // The whole gain of pairing is that σ of the difference is smaller than σ of either side. Using + // the level's σ by mistake would quietly restore the ±1.0 noise floor this exists to escape. + const r = compare({ trainCapSlack: 1 }, 40); + const d = r.deltas.map((x) => x.delta); + const m = d.reduce((p, c) => p + c, 0) / d.length; + const sd = Math.sqrt(d.reduce((p, c) => p + (c - m) ** 2, 0) / (d.length - 1)); + assert.ok(Math.abs(r.mean - m) < 1e-9); + assert.ok(Math.abs(r.sd - sd) < 1e-9); + assert.ok(Math.abs(r.stderr - sd / Math.sqrt(d.length)) < 1e-9); + }); +});