Files
station-master/TODO.md
T

29 KiB
Raw Blame History

To do

Things worth coming back to. Anything noted here should either get done or get an explicit decision not to — the point is that nothing quietly evaporates.

Ordered within each section by how much it is currently costing us.


Next

  • REBALANCE, once the rules are right — deliberately deferred. Card counts, industry counts and the track mix all need a pass together, and none of them should move until the rules stop moving. Standing distortions to account for when it happens: offices are doubled (Q12) and industries tripled (Gap 12), both tuned when the deck held 139 cards and no track; it now holds 243 of which 104 are track, so every draw is diluted by 43% — precisely the pressure those multipliers exist to relieve. Until then, read no balance conclusion from the revenue numbers; they are a functionality signal only.

  • Review the standalone replay against the site's replay viewer. node src/sim/replay.ts --seed 1234 --out replay.html writes a self-contained HTML file; the site instead reads JSON saves from public/replays/. Nothing links to the standalone one and its output is gitignored, so it is a developer tool that happens to look like a product feature. It carries two panels the site viewer does not — the bot's decision trace ("what it chose, why, and what it passed over") and the timetable — which is debugging material rather than something a player wants. Decide: fold the decision trace into the JSON viewer and delete the standalone, or keep it and accept that it is a tool. No action for now.

  • The bot was partly living off an illegal placement. Barring curves from the Running Track (they have no east-west road and dead-end the main) cost it districts 28.0 → 19.7 cards and revenue ~2.0 → 0.8. It has no plan for where a curve should go once the easy square is gone. Same root cause as the two items below; fix them together, after the rebalance.

  • THE BOT'S PRIORITIES ARE NOT THE PROBLEM — measured. Ten heuristic variations, each paired over 400+ seeds. Every reordering of what the bot prefers came out inside the noise; the only thing that moved revenue was refusing to schedule a train the Office cannot hold (+1.09 ± 0.16, t = 6.79 at 1600 seeds, revenue 1.19 → 2.28). Notable failures, all instructive: - Refusing to bury the engine costs more than it saves (−0.35, t = −2.98). It works — burial falls from 8.4 decisions a game to 0.03 — and freight halves with it, because coupling is mandatory (§A.4): the moves that bury the engine ARE the moves that pick cars up. Burial is the price of collecting, not a mistake. - Reserving Moves to get home costs 0.55 (t = −2.32), though 62 of 120 trains left on the board at game end were stranded in the district. The switching work is worth more than the departures. - Granting clearance when the train ahead has one Stage left is −0.98 (t = −5.24). Trains move in numeric order, so a follower can enter the region the leader still occupies before the leader moves. "About to leave" is not "gone". - Preferring coaches at make-up, stocking the platform first, playing Interlocking earlier, hunting the Depot in the Departments: all within noise, and three of them were exact no-ops — Interlocking sits in hand alongside a train card 0.04 decisions a game. The funnel says why: only 8% of Cargo phases have a stocked green box, and the bot already takes 42% of the turns where stocking is productive. The opportunities are not there to be prioritised better. What is left is the economy itself, which is a deck question.

  • The bot cannot get a crew next to an industry, so Flying Switch never fires. Industries are now stub-only and the bot places 2.23 a game (was 3.84), in districts averaging under two rows deep. flyingSwitch is exempted by name in the reachability sweep in sim.test.ts; deleting that line is the test that this is fixed. Same root cause as the item below.

  • THE RUN-AROUND IS OUT OF REACH OF ANY BOT, AND THE DECK IS WHY — measured, five ways. "Teach the bot to plan across turns" was tried properly and does not work. Every attempt is neutral or negative, and they fail for one reason that the numbers make plain.

    | attempt | result |
    |---|---|
    | hold ALL track for the siding | **−0.70** (t = −3.27) |
    | hold only CURVES, the closing piece | **−0.26** (t = −3.22), district 17.9 → 16.7 cards |
    | finish a run before cutting another way down | 0.00 — 398/400 games identical |
    | treat a second turnout as the closing piece | 0.00 — **400/400 identical** |
    | spend a curve only on a square that CLOSES | −0.11, and only 15 games in 400 differ at all |
    
    **The pieces never meet.** Over 12,000 Local Operations turns: a turnout and a curve are in
    hand together on **0.3%** of them, and a turnout with a MATCHING-hand curve on **0.2%** — about
    once every eight games. A run-around needs five specific pieces of the right hands in a usable
    order; the bot does not get to the two-piece prerequisite.
    
    And it is not hand pressure. The hand is FULL — mean 2.66 cards, at the three-card limit on
    78% of turns. The bot plays 11.4 track cards a game and discards 1.5, so it spends the pieces
    as they arrive because a piece that builds anything outscores holding one that might build
    more later. Holding is the only counter, and holding measures worse every way it is tried.
    
    This is a consequence of moving track into the deck, not a bot weakness: 91 run-arounds per 100
    games when track was a private 26-piece supply the player chose from, 29/100 once it was drawn,
    4/60 now. **If the run-around is meant to be the central switching puzzle — and the rules
    present it that way — the supply has to change, not the player.** Options: give track its own
    hand or yard the way the prototype did, raise the hand limit for track specifically, or print a
    siding as a single card. Nothing else reaches it.
    
  • CLEARING AN INBOUND BOX MINTS A CAR — measured at 1.29 a game against a supply of 80. Answered, and it is duplication after all. The two directions are not symmetrical: - Outbound is paid for. stockToOutbound SPLICES a loaded car out of the Division Yard to become the load, and loadCompleted swaps the emptied car into the Classification Yard as the loaded one takes its place on the industry track. Objects in, objects out. - Inbound is not. unloadBegan (apply.ts) turns one loaded car into an empty car on the track plus a load on MEN|AT|WORK, and passengersDetrained does the same to a coach — one loaded coach becomes an empty coach in the train plus an object in the red box. Then inboundCleared pushes that object into the Classification Yard as a car. The cargo becomes rolling stock, while the car it came out of is already back in service.

    Measured over 200 games: cars at the end minus 80, plus collision losses, equals the
    `inboundCleared` count in 71/200 games exactly and 258 against 279 in total — the rest is
    loads still in flight at the final whistle. So the supply inflates by about 1.3 cars a game.
    That is the number `ROLLING_STOCK_SUPPLY` is supposed to control, so **no supply figure below
    can be tuned until this is settled**. The fix is a decision, not a patch: either a load stops
    being a `RollingStock` and becomes its own type, or `inboundCleared` discards rather than
    banking. Found by a conservation audit, not by a failing test.
    
  • state = fold(events) is not literally true, and the README says it is. Replaying the event log onto a fresh state throws: the phase driver mutates state directly and emits a descriptive event afterwards — newTrainPhase does s.trays.set(...) and then pushes trainMadeUp. Replay works because it re-applies INTENTS (fromSave), not because folding events reconstructs the position. Nothing is broken today, but the claim underwrites reconnection and restart recovery, which are unbuilt — so it should be either made true or restated before anything is built on it.

  • Engines are not a SUPPLY yet, only a position. engineAt now records where the engine sits in the tray and the consist shows it, but an engine is still conjured with the tray rather than drawn from the Division Yard and returned to it. The rules put engines in the Division Yard alongside the cars, with a predefined number of them, so running out of engines should be a second way trains get held — today only the Crew Tray count does that. Needs a number to start from, then playtesting.

  • The yards are shown on the play page but not in either replay viewer. The Frame carries them, so it is a rendering job, not a modelling one.

  • The rolling stock supply is a guess. ROLLING_STOCK_SUPPLY (coach 8+8, boxcar 10+10, hopper 8+8, reefer 5+5, tank 6+6, caboose 6) is marked provisional in content.ts and was scaled alongside the Gap 12 industry increase. Now that the Classification Yard returns stock only when the Division Yard empties, these numbers set the real supply pressure. Adjust from playtesting rather than theory, and watch whether industry density feels light or heavy at the same time.

  • Heavy Grade orientation is rolled, not chosen. The card prints "Player sets orientation", but it is dealt during setup and setup has no decision point at all — createGame is a pure function of the seed, which is also what makes a save portable. Rolled from the seed for now. Revisit when setup gains an interactive phase; the orientation matters, because it decides which direction climbs and therefore what Brakeman and Helpers are worth.

  • Measure with error bars from now on. Built: node src/sim/compare.ts 1600 <tweak>=<n> runs the current bot and one variant over the same deals and reports the paired difference. Pairing drops σ from ~9 on the level to 5.3 on the difference, so 1600 seeds gives ±0.13 in about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is t ≥ 3, and the report prints the better/worse/identical split beside the mean, because a mean carried by a skewed tail is a different claim from broad improvement.

  • Re-run the three "worth ~0" action-mix experiments against the new floor. Capping the draw, pairing the two halves of a load, and restricting Enhancements were each measured "within noise of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would have looked like nothing. They are nearly free to re-run now and at least one may have been discarded wrongly.

  • Confirm the Classification Yard rule against the source. Confirmed, and the guess was wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses straight back to the Division Yard, and the Classification Yard empties ONLY when the Division Yard is bare — then all at once. The Day-boundary version I had invented was far more generous and worth +2.42 revenue a game the game does not actually grant. Corrected; revenue 9.67 → 7.25.

  • Enhancements are placed but mostly do nothing. Measured: forbidding every Enhancement except Interlocking is worth -0.01 ± 0.41 (t = -0.04) over 400 paired seeds. They neither pay nor cost. Left alone. Unlocking the Running Track straight put nine kinds on the board (telegraph 0.73, waterColumn 0.57 …), but only Interlocking has a measured effect. The Telegraph/Telephone/Radio chain adds to the other train's number when dispatching facing trains, which may be worth nothing in solitaire; Water Column removes a Watertower; Facing Point Locks prevents Derail, which is multiplayer-only. Worth measuring what each is actually worth before the bot spends actions on them.

  • The marginal Local Operations action is worth ~0, and that is the real ceiling. Three separate attempts to spend the 60 actions better — capping the draw, pairing the two halves of a load, restricting Enhancements — each measured within noise of zero over 400 paired seeds. 76% of the time an outbound industry has neither a stocked green box nor a spotted car, and only 5% of Stages have a single workable facility anywhere, yet redirecting actions at that does nothing. Something upstream limits how much work exists to do at all; find out what before spending more effort on the option mix.

  • Freight was stuck at ~2.7 loads a game and three fixes have not moved it. Sidings, facility placement, car selection and the discarded-load leak all raised revenue (3.2 → 6.5) without raising loadStarted past 2.7. The chain is not leaking and the cars are arriving correctly (57% of drops land on a facility that wants them, 0% on one that does not). The binding constraint is now upstream of routing: 60 Local Operations actions a game, and a load needs a stocked green box AND a spotted car AND a free Laborer to line up in the same Stage. Measure how many Stages have all three before changing any heuristic — the answer may be that the economy, not the bot, is what caps freight.

  • stats.ts undercounts freight. Fixed: both halves counted, freight share 39% → 49%. Worth revisiting the industry density decision below, which was taken on the old number.


Open questions for Jesse

Blocked on a decision, not on work.

  • Q13 — rear-end collisions on a Mainline card. Answered: collide on catching up. Implemented, and not on cards that print "trains may pass". Invisible to a bot that always denies clearance; a bot that always allows drops from 7.34 revenue to -5.13.
  • Poling. The only card in the deck with no defined behaviour — the sheet records its effect as "TBD in the source". A test asserts it stays TBD so nobody invents one.
  • Heavy Grade orientation at setup. The card says "Player sets orientation", but createGame is synchronous and has no decision point, so it is currently rolled from the seed. Should become a real choice when setup gains an interactive phase.

Balance, provisional

Numbers chosen to fix a measured problem rather than taken from the design. Revisit once the victory target is settled and freight carries its intended share.

  • Office card density (Depot 4→8, Station 2→4, Terminal 1→2). Chosen to remove a 25% chance of an unwinnable opening deal. Blunt: it lifts the whole ladder and dilutes every other category. The better answer may be fewer Terminals, a cheaper first upgrade, or more A/D capacity at the Whistle Post itself.

  • Industry density (9 → 27, Gap 12). Restored roughly the prototype ratio. The "freight is only 13–18% of gross" figure that motivated this was partly a measurement bug (see the stats.ts item) and partly the car-selection bug; freight now runs at 37%. Worth re-deciding whether 27 is still the right number now that the industries are actually served.

  • Train density. Left alone by decision, but noted: 22 train cards in 140 are drawn less often than 22 in 115 were, and trains scheduled fell 2.9 → 2.1 as a side effect of the other density changes.

  • The victory target (20 over 5 Days) is out of reach by a factor of about four, and the Office ladder is why. Measured over 800 games with the tuned bot, which no longer throws revenue away on collisions (0.0 a game, down from 0.4):

    | trains scheduled | games | revenue |     | Office reached | games | trains | revenue |
    |---|---|---|---|---|---|---|---|
    | 0 | 110 | 0.67 |  | Whistle Post | 297 | 0.81 | 0.62 |
    | 1 | 379 | 1.69 |  | Depot        | 272 | 1.49 | 2.92 |
    | 2 | 234 | 3.72 |  | Station      | 176 | 1.89 | 4.36 |
    | 3 |  68 | 5.68 |  | Terminal     |  55 | 1.93 | 5.04 |
    | 4 |   9 | 5.78 |  | | | | |
    
    Revenue is almost exactly linear in trains scheduled — about **1.9 a train** — and trains are
    capped by A/D capacity, which is the Office tier, which is a card you have to draw. So the
    whole economy hangs off one valve: **37% of games never leave the Whistle Post and earn 0.62;
    53% of all games earn nothing at all.**
    
    Extrapolating the line, 20 Revenue needs roughly **11 trains and therefore 11 A/D tracks**. A
    Terminal has four. The target is not merely missed, it is structurally unreachable under this
    deck at this Office ladder — no amount of bot skill closes it, and the best game seen in 800
    was 26 against a median of 0.
    
    The three ways out are all yours to choose between, and they are different games:
    1. **Lower the target** to what a 5-Day game can produce (6–8 looks like the honest number).
    2. **Open the valve** — more Office cards, or a cheaper first upgrade, or more A/D capacity at
       the Whistle Post, so the ladder is climbed rather than drawn.
    3. **Raise revenue per arrival.** It is 0.46 today; each arrival can in principle pay 2 for
       passengers alone. That is the freight/passenger conversion problem, not the traffic problem.
    
    Nothing here is a bot weakness any more, which is what this measurement was waiting on.
    

Not yet built

  • Real audio, as committed assets. Everything the game plays is synthesised from oscillators (src/web/sound.ts), which was the honest choice for a site that fetches nothing — but it is a placeholder, not the finished sound. Sound therefore defaults to OFF. - "All aboard" most of all. It currently goes through the browser's speechSynthesis, so it is whatever system voice the player happens to have — a robot, not a conductor. A real clip is the single biggest improvement available here. - Find and add the rest as assets: steam whistle, grade-crossing bell, couplers clashing, a train pulling away. Needs licences that permit redistribution (CC0 or similar), files small enough to commit, and a check that the "fetches nothing external" test still passes — assets must be served from the site's own folder, never hot-linked. - Keep the synthesised versions as the fallback for anything not sourced, so a missing file is a quieter game rather than a broken one.

  • Regions as the primary model (the other half of §8.2). The Division map now DRAWS regions, deriving position from what the crossing already cost. The engine still models a crossing as a countdown of Stages, so two things printed on the cards remain unimplemented: - entryPoints is declared on every Mainline profile and read nowhere. The Heavy Grade card has five named Start positions, and playing Brakeman is supposed to move your entry point along the card. The engine gets the same ANSWER by taking a Stage off the crossing, which is why the derived drawing looks right — but the mechanism is not the printed one, so a card whose starts do not correspond to its speed would be drawn wrong. - implications.md §6 calls this "the single largest mechanical gap" and asks for typed cards with speeds and named entries, with crossing time DERIVED from the region walk. Doing it properly changes movement, so it invalidates every balance figure — revenue 8.7, the freight numbers, all of it — and needs a full paired re-measure over 400 seeds. Needs the source Start-position art for the ten card types before it can begin.

  • Player settings, saved. The district's auto-focus is the first of these: it is DISPLAY state, so in a multiplayer game two players may reasonably want it set differently and it must never become part of game state. It currently resets on reload. Worth a settings object in localStorage — auto-focus mode to start with, and whatever else earns a preference — kept strictly separate from the save, which is the seed plus the intents and has to stay portable.

  • Let the game join a call and talk to the table. Long-term. If the game could join a Zoom, Teams or Jitsi call and post into its chat, it could carry the whole table's shared state without anyone alt-tabbing: the history of actions as they happen, and a prompt when someone is holding the game up — "Now waiting on player Alice to complete the Cargo phase." - Further out, audio into the same call: a crash when a collision happens, a bell as the Stage clock turns over. - Further out still, a nudge on a timer — if a player has not moved within some interval, the game says so, by beep or by spoken line: "Still waiting on Alice to complete the Cargo phase." That turns the turn chart's "waiting on" chip into something a distracted table actually notices.

  • Multiplayer train make-up is a round, not one player's job. When a new train is built, players take turns adding cars to the consist; in solitaire one player does all of it. The engine currently has no per-player turn within the New Train phase, so this is unbuilt rather than wrong.

  • Action cards (10) and Space-use cards (12). Genuinely multiplayer-only — they are played AT an opponent. Rejected with NOT_IMPLEMENTED.

  • Multiplayer proper. The engine runs 2–5 player games and the bot plays them, but there is no server, no turn submission, and no per-player view.


Smaller things

  • Carry links forward in replay frames. Done, and the premise was wrong in an instructive way: measured, links was 5% of the cells payload. What actually cost was the what prose (32%), the facility object stored a second time inside its own cell (24%) and the rest of the static identity (26%). All three are interned now — 3415 KB → 1877 KB, and a round-trip test runs the page's own unpacking function.
  • Undo is unlimited step-back, and that is a decision to revisit. The save is the seed plus the intents, so undo() replays without the last one and can walk all the way to the deal. The RNG advances with the replay, so the same play re-rolls the same 1D12 — you cannot undo your way to a better die. But you CAN see a train's departure Stage and then spend the turn differently, which is an ordinary solitaire take-back and also a real information leak. Options if it starts to feel like cheating: make the Stage boundary a commit point, or cap the depth at the current Stage. Deliberately left open until it has been played with. Multiplayer gets nothing until there is a proposal/agreement flow — undo there is a table decision, not a button.
  • The 5 MB replay size limit is arbitrary. Invented, not a browser constraint. It has earned its place — it caught a 5.2 MB payload that turned out to be the whole grid re-serialised every frame — but the number itself deserves a reason.
  • Curves are drawn as two straight segments meeting, not true arcs. Fine at this size, angular close up.
  • Wide boards scroll. A 40-card district and a 13-section Division both need horizontal scrolling. Legible, not compact.
  • Every published replay was dead. All three replayed 2 intents of roughly 400 and presented as short games, exactly as the item below predicted. Re-recorded from bot games with node src/sim/save-replay.ts, which verifies each save round-trips before writing it, and harness.test.ts now fails if a published replay stops short. The version-stamp item below is still worth doing — this catches the breakage, it does not explain it to a player.
  • Save/restore is not version-aware. A save from an older ruleset stops replaying rather than failing loudly, which is the safe direction but says little about what changed. This has now bitten once: both published replays were dead — one got 42 intents into 360, the other 4 of 338 — and nothing said so; they simply ended early and looked like short games. A save should carry a ruleset stamp and the page should say "this replay was recorded under an older ruleset and stops at Stage N" rather than presenting a truncated game as a whole one.

Done, kept for the reasoning

  • Put rolling stock back into circulation. The Classification Yard was write-only — seven writers, no readers — so 37% of all rolling stock left the game by Day 5. Returning it at the Day boundary is +2.32 ± 0.52 (t = 8.79), the largest single change measured on this bot, and it was ranked THIRD and predicted not to matter because the Division Yard never runs dry. The aggregate was the wrong measure; having the right commodity at the right moment is what counts.
  • Make Enhancements reachable at all. The bot never laid a straight on the Running Track (0.00 in 100 games) because two-arc run-arounds do not need one — so 13 of the 18 Enhancement cards had nowhere to go, including Interlocking, the only cure for the only penalty in the game (no free A/D track, 27% of gross). One scored straight fixed it: enhancements placed 0.64 → 3.01, collision cost 2.70 → 1.91, worst game −47 → −24. Revenue +0.70 ± 0.74 paired over 400 seeds — real but not significant alone; the variance reduction is the clearer win.
  • Stop the bot discarding its own freight. canStockProductively did not check the Division Yard while the engine's stockOutbound does, so Freight Agent was chosen when nothing could be stocked and the follow-through fell through to an unjam that threw a waiting load out of the green box — 3.10 a game against 2.71 started. Now 0.00. Revenue 6.0 → 6.5, wins 5 → 8 in 100. Also confirmed routing was never the problem: 0% of drops land on a facility that does not want the car.
  • Why switching work did not become Revenue. Answered: it was the freight the crew shuffled, not the shuffling. The chain never leaked — 95% of started loads finished — it was barely entered, because a load needs a matching empty car spotted and half the industries never asked for one. Three fixes later (sidings, facility placement, car selection) revenue is 3.2 → 6.0 and freight 26% → 37% of gross.
  • Fix car selection. Three of six industries were invisible to wantedCars — a hand-written industry→car map naming two industries that do not exist and omitting three that do — so tank cars were dropped 0 times in 100 games. Derived from INDUSTRY_PROFILES now, and the second commodity of the two-commodity industries is reachable. Revenue 5.0 → 6.0, freight share 25% → 37%, wins 1 → 5 in 100.
  • Put the industries on the run-around. Facility placement was unscored — the first legal square — so 0.00 facilities a game sat on a loop; now 1.08. The instructive part was the second bug: scoring facilities onto the siding row dropped run-arounds 91→36, because the anchor test asked a card's KIND rather than its PORTS and an industry in the line read as a dead end. Revenue 4.1 → 5.0. Freight did not follow, which is the item above.
  • Make the bot build sidings that are sidings. 0 run-arounds in 100 games → 91. Three bugs, all scoring on local shape without checking it reached anything; the decisive one was that bestTrackLay never declined a piece, so it spent the track supply on whatever was legal.
  • Teach the bot what a siding is for. Nose coupling (§A.3) implemented, so approach direction decides which car is droppable; the bot runs around rather than setting out, when the drop can follow. Switching activity transformed, revenue unchanged.
  • Curve geometry. Curves were topologically identical duplicates of turnouts, and nothing reached north, so a district could only be a vertical column. Now two-port rotatable arcs.
  • Q10 — when track may be laid. During the "draw a card" option, one piece a turn. Track was a card when §6.2 was written; a 26-piece supply has no hand to bound it.
  • Q11 — which way a Heavy Grade climbs. Answered from the card: it prints "(Up)" and "Player sets orientation", so it is a property of the placed card, not a compass constant.
  • §6.2's reshuffle. Implemented, and the deckReshuffled event it had already declared and narrated — but never emitted or reduced — is now real. Not yet reached in play: solitaire Campaign ends with 168.8 of 243 in the deck and four-player Campaign with 86.6, and no run of any length has emptied it. It is a safety net rather than a live mechanic today, which is worth knowing before tuning draw rates.
  • Q12 — Whistle Post lock-in. Players always start at a Whistle Post; office density doubled instead.