Files
station-master/TODO.md
T

12 KiB
Raw Blame History

To do

Things worth coming back to. Anything noted here should either get done or get an explicit decision not to — the point is that nothing quietly evaporates.

Ordered within each section by how much it is currently costing us.


Next

  • Measure with error bars from now on. Revenue has a standard deviation of ~9, so a 100-game run carries about ±1.0 of noise — every single-change revenue claim in the changelog before the Interlocking work is inside that. Use paired per-seed comparison (the harness deals the same seeds either way) and 400+ games before calling a heuristic change good or bad. The first attempt at the Running Track straight was read as a 0.6 REGRESSION on 100 games and is a 0.7 improvement on 400.
  • Confirm the Classification Yard rule against the source. Cars now return to the Division Yard at the Day boundary — an ASSUMPTION, not a recovered rule. Gap 2c says everything but cabooses goes to Classification and never says how it empties. It is worth +2.32 revenue a game, so if the real rule differs the balance numbers move with it.
  • Enhancements are placed but mostly do nothing. Measured: forbidding every Enhancement except Interlocking is worth -0.01 ± 0.41 (t = -0.04) over 400 paired seeds. They neither pay nor cost. Left alone. Unlocking the Running Track straight put nine kinds on the board (telegraph 0.73, waterColumn 0.57 …), but only Interlocking has a measured effect. The Telegraph/Telephone/Radio chain adds to the other train's number when dispatching facing trains, which may be worth nothing in solitaire; Water Column removes a Watertower; Facing Point Locks prevents Derail, which is multiplayer-only. Worth measuring what each is actually worth before the bot spends actions on them.
  • The marginal Local Operations action is worth ~0, and that is the real ceiling. Three separate attempts to spend the 60 actions better — capping the draw, pairing the two halves of a load, restricting Enhancements — each measured within noise of zero over 400 paired seeds. 76% of the time an outbound industry has neither a stocked green box nor a spotted car, and only 5% of Stages have a single workable facility anywhere, yet redirecting actions at that does nothing. Something upstream limits how much work exists to do at all; find out what before spending more effort on the option mix.
  • Freight was stuck at ~2.7 loads a game and three fixes have not moved it. Sidings, facility placement, car selection and the discarded-load leak all raised revenue (3.2 → 6.5) without raising loadStarted past 2.7. The chain is not leaking and the cars are arriving correctly (57% of drops land on a facility that wants them, 0% on one that does not). The binding constraint is now upstream of routing: 60 Local Operations actions a game, and a load needs a stocked green box AND a spotted car AND a free Laborer to line up in the same Stage. Measure how many Stages have all three before changing any heuristic — the answer may be that the economy, not the bot, is what caps freight.
  • stats.ts undercounts freight. rev.freightUnload is assigned eventCounts['unloadBegan'] — unloads begun, not revenue earned — and grossFreight uses freightLoad alone, so freightShare omits unload revenue entirely. The comment justifying it ("an unload scores through the same event as a load completion") is wrong: apply.ts:947 emits a distinct freightUnload reason. This is why freight was recorded at 13–18% of gross.

Open questions for Jesse

Blocked on a decision, not on work.

  • Q13 — rear-end collisions on a Mainline card. §10 says a Mainline collision is the Superintendent's fault and removes both trains, and ABS Signals exists to prevent rear-enders — but §8.3's trigger list does not include one, and none is implemented. Granting clearance is currently free: verified, both trains survive, no penalty. Three candidates: collide on entry (clearance becomes a gamble), collide on catching up (rewards judging the gap), or accept that clearance is safe and ABS Signals is worth less than it reads. The buttons currently describe only what the engine does, so nothing promises a consequence that cannot happen.
  • Poling. The only card in the deck with no defined behaviour — the sheet records its effect as "TBD in the source". A test asserts it stays TBD so nobody invents one.
  • Heavy Grade orientation at setup. The card says "Player sets orientation", but createGame is synchronous and has no decision point, so it is currently rolled from the seed. Should become a real choice when setup gains an interactive phase.

Balance, provisional

Numbers chosen to fix a measured problem rather than taken from the design. Revisit once the victory target is settled and freight carries its intended share.

  • Office card density (Depot 4→8, Station 2→4, Terminal 1→2). Chosen to remove a 25% chance of an unwinnable opening deal. Blunt: it lifts the whole ladder and dilutes every other category. The better answer may be fewer Terminals, a cheaper first upgrade, or more A/D capacity at the Whistle Post itself.
  • Industry density (9 → 27, Gap 12). Restored roughly the prototype ratio. The "freight is only 13–18% of gross" figure that motivated this was partly a measurement bug (see the stats.ts item) and partly the car-selection bug; freight now runs at 37%. Worth re-deciding whether 27 is still the right number now that the industries are actually served.
  • Train density. Left alone by decision, but noted: 22 train cards in 140 are drawn less often than 22 in 115 were, and trains scheduled fell 2.9 → 2.1 as a side effect of the other density changes.
  • The victory target itself (20 over 5 Days). 5 wins in 100, up from 1, and the bot now does exploit sidings — so that unclaimed gain has been claimed and the target is still missed by a wide margin (mean 6.0 against 20). This is the next real balance question.

Not yet built

  • Action cards (10) and Space-use cards (12). Genuinely multiplayer-only — they are played AT an opponent. Rejected with NOT_IMPLEMENTED.
  • Multiplayer proper. The engine runs 2–5 player games and the bot plays them, but there is no server, no turn submission, and no per-player view.

Smaller things

  • Carry links forward in replay frames. Static per card and never changes once laid, but re-sent whenever anything on the board changes. The replay is 3.4 MB, mostly board state.
  • The 5 MB replay size limit is arbitrary. Invented, not a browser constraint. It has earned its place — it caught a 5.2 MB payload that turned out to be the whole grid re-serialised every frame — but the number itself deserves a reason.
  • Curves are drawn as two straight segments meeting, not true arcs. Fine at this size, angular close up.
  • Wide boards scroll. A 40-card district and a 13-section Division both need horizontal scrolling. Legible, not compact.
  • No way to start a fresh game from inside the page. ?seed= gives a reproducible deal and "new game" only appears once a game has ended, so abandoning a bad opening means editing the URL.
  • Save/restore is not version-aware. A save from an older ruleset stops replaying rather than failing loudly, which is the safe direction but says little about what changed.

Done, kept for the reasoning

  • Put rolling stock back into circulation. The Classification Yard was write-only — seven writers, no readers — so 37% of all rolling stock left the game by Day 5. Returning it at the Day boundary is +2.32 ± 0.52 (t = 8.79), the largest single change measured on this bot, and it was ranked THIRD and predicted not to matter because the Division Yard never runs dry. The aggregate was the wrong measure; having the right commodity at the right moment is what counts.
  • Make Enhancements reachable at all. The bot never laid a straight on the Running Track (0.00 in 100 games) because two-arc run-arounds do not need one — so 13 of the 18 Enhancement cards had nowhere to go, including Interlocking, the only cure for the only penalty in the game (no free A/D track, 27% of gross). One scored straight fixed it: enhancements placed 0.64 → 3.01, collision cost 2.70 → 1.91, worst game −47 → −24. Revenue +0.70 ± 0.74 paired over 400 seeds — real but not significant alone; the variance reduction is the clearer win.
  • Stop the bot discarding its own freight. canStockProductively did not check the Division Yard while the engine's stockOutbound does, so Freight Agent was chosen when nothing could be stocked and the follow-through fell through to an unjam that threw a waiting load out of the green box — 3.10 a game against 2.71 started. Now 0.00. Revenue 6.0 → 6.5, wins 5 → 8 in 100. Also confirmed routing was never the problem: 0% of drops land on a facility that does not want the car.
  • Why switching work did not become Revenue. Answered: it was the freight the crew shuffled, not the shuffling. The chain never leaked — 95% of started loads finished — it was barely entered, because a load needs a matching empty car spotted and half the industries never asked for one. Three fixes later (sidings, facility placement, car selection) revenue is 3.2 → 6.0 and freight 26% → 37% of gross.
  • Fix car selection. Three of six industries were invisible to wantedCars — a hand-written industry→car map naming two industries that do not exist and omitting three that do — so tank cars were dropped 0 times in 100 games. Derived from INDUSTRY_PROFILES now, and the second commodity of the two-commodity industries is reachable. Revenue 5.0 → 6.0, freight share 25% → 37%, wins 1 → 5 in 100.
  • Put the industries on the run-around. Facility placement was unscored — the first legal square — so 0.00 facilities a game sat on a loop; now 1.08. The instructive part was the second bug: scoring facilities onto the siding row dropped run-arounds 91→36, because the anchor test asked a card's KIND rather than its PORTS and an industry in the line read as a dead end. Revenue 4.1 → 5.0. Freight did not follow, which is the item above.
  • Make the bot build sidings that are sidings. 0 run-arounds in 100 games → 91. Three bugs, all scoring on local shape without checking it reached anything; the decisive one was that bestTrackLay never declined a piece, so it spent the track supply on whatever was legal.
  • Teach the bot what a siding is for. Nose coupling (§A.3) implemented, so approach direction decides which car is droppable; the bot runs around rather than setting out, when the drop can follow. Switching activity transformed, revenue unchanged.
  • Curve geometry. Curves were topologically identical duplicates of turnouts, and nothing reached north, so a district could only be a vertical column. Now two-port rotatable arcs.
  • Q10 — when track may be laid. During the "draw a card" option, one piece a turn. Track was a card when §6.2 was written; a 26-piece supply has no hand to bound it.
  • Q11 — which way a Heavy Grade climbs. Answered from the card: it prints "(Up)" and "Player sets orientation", so it is a property of the placed card, not a compass constant.
  • Q12 — Whistle Post lock-in. Players always start at a Whistle Post; office density doubled instead.