22 KiB
To do
Things worth coming back to. Anything noted here should either get done or get an explicit decision not to — the point is that nothing quietly evaporates.
Ordered within each section by how much it is currently costing us.
Next
-
REBALANCE, once the rules are right — deliberately deferred. Card counts, industry counts and the track mix all need a pass together, and none of them should move until the rules stop moving. Standing distortions to account for when it happens: offices are doubled (Q12) and industries tripled (Gap 12), both tuned when the deck held 139 cards and no track; it now holds 243 of which 104 are track, so every draw is diluted by 43% — precisely the pressure those multipliers exist to relieve. Until then, read no balance conclusion from the revenue numbers; they are a functionality signal only.
-
Review the standalone replay against the site's replay viewer.
node src/sim/replay.ts --seed 1234 --out replay.htmlwrites a self-contained HTML file; the site instead reads JSON saves frompublic/replays/. Nothing links to the standalone one and its output is gitignored, so it is a developer tool that happens to look like a product feature. It carries two panels the site viewer does not — the bot's decision trace ("what it chose, why, and what it passed over") and the timetable — which is debugging material rather than something a player wants. Decide: fold the decision trace into the JSON viewer and delete the standalone, or keep it and accept that it is a tool. No action for now. -
The bot was partly living off an illegal placement. Barring curves from the Running Track (they have no east-west road and dead-end the main) cost it districts 28.0 → 19.7 cards and revenue ~2.0 → 0.8. It has no plan for where a curve should go once the easy square is gone. Same root cause as the two items below; fix them together, after the rebalance.
-
The bot cannot get a crew next to an industry, so Flying Switch never fires. Industries are now stub-only and the bot places 2.23 a game (was 3.84), in districts averaging under two rows deep.
flyingSwitchis exempted by name in the reachability sweep insim.test.ts; deleting that line is the test that this is fixed. Same root cause as the item below. -
The bot does not play for a run-around any more, and revenue halved. With track in the deck a run-around needs a turnout, a matching curve, straights, a second curve and a second turnout, all of the right hand, arriving in a three-card hand in a usable order. The bot holds no plan across turns and discards a piece it cannot use immediately: run-arounds fell 70/100 → 29/100 and revenue 5.80 → 2.87. Two test floors in
sim.test.tsare pinned below the measurement as break-detectors rather than targets, and say so. Do the density re-measurement above first — bot weakness and deck density are currently confounded. -
A "load" is stored as a car, so rolling stock cannot be counted.
outboundBox,inboundBoxandmenAtWorkall holdRollingStock, andfreightAgent.stockOutboundtakes a LOADED CAR out of the Division Yard to fill a green box. So a census of every holder comes to 92 against the 80 dealt at setup — not necessarily duplication, because some of those objects are cargo in transit rather than cars, but there is no way to tell them apart. Until a load is its own type, "is any stock being created or destroyed?" is an unanswerable question, and the supply numbers below cannot be tuned with confidence. -
Engines are not a SUPPLY yet, only a position.
engineAtnow records where the engine sits in the tray and the consist shows it, but an engine is still conjured with the tray rather than drawn from the Division Yard and returned to it. The rules put engines in the Division Yard alongside the cars, with a predefined number of them, so running out of engines should be a second way trains get held — today only the Crew Tray count does that. Needs a number to start from, then playtesting. -
The yards are shown on the play page but not in either replay viewer. The Frame carries them, so it is a rendering job, not a modelling one.
-
The rolling stock supply is a guess.
ROLLING_STOCK_SUPPLY(coach 8+8, boxcar 10+10, hopper 8+8, reefer 5+5, tank 6+6, caboose 6) is marked provisional incontent.tsand was scaled alongside the Gap 12 industry increase. Now that the Classification Yard returns stock only when the Division Yard empties, these numbers set the real supply pressure. Adjust from playtesting rather than theory, and watch whether industry density feels light or heavy at the same time. -
Heavy Grade orientation is rolled, not chosen. The card prints "Player sets orientation", but it is dealt during setup and setup has no decision point at all —
createGameis a pure function of the seed, which is also what makes a save portable. Rolled from the seed for now. Revisit when setup gains an interactive phase; the orientation matters, because it decides which direction climbs and therefore what Brakeman and Helpers are worth. -
Measure with error bars from now on.Built:node src/sim/compare.ts 1600 <tweak>=<n>runs the current bot and one variant over the same deals and reports the paired difference. Pairing drops σ from ~9 on the level to 5.3 on the difference, so 1600 seeds gives ±0.13 in about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is t ≥ 3, and the report prints the better/worse/identical split beside the mean, because a mean carried by a skewed tail is a different claim from broad improvement. -
Re-run the three "worth ~0" action-mix experiments against the new floor. Capping the draw, pairing the two halves of a load, and restricting Enhancements were each measured "within noise of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would have looked like nothing. They are nearly free to re-run now and at least one may have been discarded wrongly.
-
Confirm the Classification Yard rule against the source.Confirmed, and the guess was wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses straight back to the Division Yard, and the Classification Yard empties ONLY when the Division Yard is bare — then all at once. The Day-boundary version I had invented was far more generous and worth +2.42 revenue a game the game does not actually grant. Corrected; revenue 9.67 → 7.25. -
Enhancements are placed but mostly do nothing.Measured: forbidding every Enhancement except Interlocking is worth -0.01 ± 0.41 (t = -0.04) over 400 paired seeds. They neither pay nor cost. Left alone. Unlocking the Running Track straight put nine kinds on the board (telegraph 0.73, waterColumn 0.57 …), but only Interlocking has a measured effect. The Telegraph/Telephone/Radio chain adds to the other train's number when dispatching facing trains, which may be worth nothing in solitaire; Water Column removes a Watertower; Facing Point Locks prevents Derail, which is multiplayer-only. Worth measuring what each is actually worth before the bot spends actions on them. -
The marginal Local Operations action is worth ~0, and that is the real ceiling. Three separate attempts to spend the 60 actions better — capping the draw, pairing the two halves of a load, restricting Enhancements — each measured within noise of zero over 400 paired seeds. 76% of the time an outbound industry has neither a stocked green box nor a spotted car, and only 5% of Stages have a single workable facility anywhere, yet redirecting actions at that does nothing. Something upstream limits how much work exists to do at all; find out what before spending more effort on the option mix.
-
Freight was stuck at ~2.7 loads a game and three fixes have not moved it. Sidings, facility placement, car selection and the discarded-load leak all raised revenue (3.2 → 6.5) without raising
loadStartedpast 2.7. The chain is not leaking and the cars are arriving correctly (57% of drops land on a facility that wants them, 0% on one that does not). The binding constraint is now upstream of routing: 60 Local Operations actions a game, and a load needs a stocked green box AND a spotted car AND a free Laborer to line up in the same Stage. Measure how many Stages have all three before changing any heuristic — the answer may be that the economy, not the bot, is what caps freight. -
Fixed: both halves counted, freight share 39% → 49%. Worth revisiting the industry density decision below, which was taken on the old number.stats.tsundercounts freight.
Open questions for Jesse
Blocked on a decision, not on work.
Q13 — rear-end collisions on a Mainline card.Answered: collide on catching up. Implemented, and not on cards that print "trains may pass". Invisible to a bot that always denies clearance; a bot that always allows drops from 7.34 revenue to -5.13.- Poling. The only card in the deck with no defined behaviour — the sheet records its effect as "TBD in the source". A test asserts it stays TBD so nobody invents one.
- Heavy Grade orientation at setup. The card says "Player sets orientation", but
createGameis synchronous and has no decision point, so it is currently rolled from the seed. Should become a real choice when setup gains an interactive phase.
Balance, provisional
Numbers chosen to fix a measured problem rather than taken from the design. Revisit once the victory target is settled and freight carries its intended share.
- Office card density (Depot 4→8, Station 2→4, Terminal 1→2). Chosen to remove a 25% chance of an unwinnable opening deal. Blunt: it lifts the whole ladder and dilutes every other category. The better answer may be fewer Terminals, a cheaper first upgrade, or more A/D capacity at the Whistle Post itself.
- Industry density (9 → 27, Gap 12). Restored roughly the prototype ratio. The "freight is
only 13–18% of gross" figure that motivated this was partly a measurement bug (see the
stats.tsitem) and partly the car-selection bug; freight now runs at 37%. Worth re-deciding whether 27 is still the right number now that the industries are actually served. - Train density. Left alone by decision, but noted: 22 train cards in 140 are drawn less often than 22 in 115 were, and trains scheduled fell 2.9 → 2.1 as a side effect of the other density changes.
- The victory target itself (20 over 5 Days). 5 wins in 100, up from 1, and the bot now does exploit sidings — so that unclaimed gain has been claimed and the target is still missed by a wide margin (mean 6.0 against 20). This is the next real balance question.
Not yet built
-
Real audio, as committed assets. Everything the game plays is synthesised from oscillators (
src/web/sound.ts), which was the honest choice for a site that fetches nothing — but it is a placeholder, not the finished sound. Sound therefore defaults to OFF. - "All aboard" most of all. It currently goes through the browser'sspeechSynthesis, so it is whatever system voice the player happens to have — a robot, not a conductor. A real clip is the single biggest improvement available here. - Find and add the rest as assets: steam whistle, grade-crossing bell, couplers clashing, a train pulling away. Needs licences that permit redistribution (CC0 or similar), files small enough to commit, and a check that the "fetches nothing external" test still passes — assets must be served from the site's own folder, never hot-linked. - Keep the synthesised versions as the fallback for anything not sourced, so a missing file is a quieter game rather than a broken one. -
Regions as the primary model (the other half of §8.2). The Division map now DRAWS regions, deriving position from what the crossing already cost. The engine still models a crossing as a countdown of Stages, so two things printed on the cards remain unimplemented: -
entryPointsis declared on every Mainline profile and read nowhere. The Heavy Grade card has five named Start positions, and playing Brakeman is supposed to move your entry point along the card. The engine gets the same ANSWER by taking a Stage off the crossing, which is why the derived drawing looks right — but the mechanism is not the printed one, so a card whose starts do not correspond to its speed would be drawn wrong. -implications.md§6 calls this "the single largest mechanical gap" and asks for typed cards with speeds and named entries, with crossing time DERIVED from the region walk. Doing it properly changes movement, so it invalidates every balance figure — revenue 8.7, the freight numbers, all of it — and needs a full paired re-measure over 400 seeds. Needs the source Start-position art for the ten card types before it can begin. -
Player settings, saved. The district's auto-focus is the first of these: it is DISPLAY state, so in a multiplayer game two players may reasonably want it set differently and it must never become part of game state. It currently resets on reload. Worth a settings object in localStorage — auto-focus mode to start with, and whatever else earns a preference — kept strictly separate from the save, which is the seed plus the intents and has to stay portable.
-
Let the game join a call and talk to the table. Long-term. If the game could join a Zoom, Teams or Jitsi call and post into its chat, it could carry the whole table's shared state without anyone alt-tabbing: the history of actions as they happen, and a prompt when someone is holding the game up — "Now waiting on player Alice to complete the Cargo phase." - Further out, audio into the same call: a crash when a collision happens, a bell as the Stage clock turns over. - Further out still, a nudge on a timer — if a player has not moved within some interval, the game says so, by beep or by spoken line: "Still waiting on Alice to complete the Cargo phase." That turns the turn chart's "waiting on" chip into something a distracted table actually notices.
-
Multiplayer train make-up is a round, not one player's job. When a new train is built, players take turns adding cars to the consist; in solitaire one player does all of it. The engine currently has no per-player turn within the New Train phase, so this is unbuilt rather than wrong.
-
Action cards (10) and Space-use cards (12). Genuinely multiplayer-only — they are played AT an opponent. Rejected with
NOT_IMPLEMENTED. -
Multiplayer proper. The engine runs 2–5 player games and the bot plays them, but there is no server, no turn submission, and no per-player view.
Smaller things
- Carry
linksforward in replay frames. Done, and the premise was wrong in an instructive way: measured,linkswas 5% of thecellspayload. What actually cost was thewhatprose (32%), the facility object stored a second time inside its own cell (24%) and the rest of the static identity (26%). All three are interned now — 3415 KB → 1877 KB, and a round-trip test runs the page's own unpacking function. - Undo is unlimited step-back, and that is a decision to revisit. The save is the seed plus
the intents, so
undo()replays without the last one and can walk all the way to the deal. The RNG advances with the replay, so the same play re-rolls the same 1D12 — you cannot undo your way to a better die. But you CAN see a train's departure Stage and then spend the turn differently, which is an ordinary solitaire take-back and also a real information leak. Options if it starts to feel like cheating: make the Stage boundary a commit point, or cap the depth at the current Stage. Deliberately left open until it has been played with. Multiplayer gets nothing until there is a proposal/agreement flow — undo there is a table decision, not a button. - The 5 MB replay size limit is arbitrary. Invented, not a browser constraint. It has earned its place — it caught a 5.2 MB payload that turned out to be the whole grid re-serialised every frame — but the number itself deserves a reason.
- Curves are drawn as two straight segments meeting, not true arcs. Fine at this size, angular close up.
- Wide boards scroll. A 40-card district and a 13-section Division both need horizontal scrolling. Legible, not compact.
- Save/restore is not version-aware. A save from an older ruleset stops replaying rather than failing loudly, which is the safe direction but says little about what changed. This has now bitten once: both published replays were dead — one got 42 intents into 360, the other 4 of 338 — and nothing said so; they simply ended early and looked like short games. A save should carry a ruleset stamp and the page should say "this replay was recorded under an older ruleset and stops at Stage N" rather than presenting a truncated game as a whole one.
Done, kept for the reasoning
- Put rolling stock back into circulation. The Classification Yard was write-only — seven writers, no readers — so 37% of all rolling stock left the game by Day 5. Returning it at the Day boundary is +2.32 ± 0.52 (t = 8.79), the largest single change measured on this bot, and it was ranked THIRD and predicted not to matter because the Division Yard never runs dry. The aggregate was the wrong measure; having the right commodity at the right moment is what counts.
- Make Enhancements reachable at all. The bot never laid a straight on the Running Track
(0.00 in 100 games) because two-arc run-arounds do not need one — so 13 of the 18 Enhancement
cards had nowhere to go, including Interlocking, the only cure for the only penalty in the
game (
no free A/D track, 27% of gross). One scored straight fixed it: enhancements placed 0.64 → 3.01, collision cost 2.70 → 1.91, worst game −47 → −24. Revenue +0.70 ± 0.74 paired over 400 seeds — real but not significant alone; the variance reduction is the clearer win. - Stop the bot discarding its own freight.
canStockProductivelydid not check the Division Yard while the engine'sstockOutbounddoes, so Freight Agent was chosen when nothing could be stocked and the follow-through fell through to an unjam that threw a waiting load out of the green box — 3.10 a game against 2.71 started. Now 0.00. Revenue 6.0 → 6.5, wins 5 → 8 in 100. Also confirmed routing was never the problem: 0% of drops land on a facility that does not want the car. - Why switching work did not become Revenue. Answered: it was the freight the crew shuffled, not the shuffling. The chain never leaked — 95% of started loads finished — it was barely entered, because a load needs a matching empty car spotted and half the industries never asked for one. Three fixes later (sidings, facility placement, car selection) revenue is 3.2 → 6.0 and freight 26% → 37% of gross.
- Fix car selection. Three of six industries were invisible to
wantedCars— a hand-written industry→car map naming two industries that do not exist and omitting three that do — so tank cars were dropped 0 times in 100 games. Derived fromINDUSTRY_PROFILESnow, and the second commodity of the two-commodity industries is reachable. Revenue 5.0 → 6.0, freight share 25% → 37%, wins 1 → 5 in 100. - Put the industries on the run-around. Facility placement was unscored — the first legal square — so 0.00 facilities a game sat on a loop; now 1.08. The instructive part was the second bug: scoring facilities onto the siding row dropped run-arounds 91→36, because the anchor test asked a card's KIND rather than its PORTS and an industry in the line read as a dead end. Revenue 4.1 → 5.0. Freight did not follow, which is the item above.
- Make the bot build sidings that are sidings. 0 run-arounds in 100 games → 91. Three bugs,
all scoring on local shape without checking it reached anything; the decisive one was that
bestTrackLaynever declined a piece, so it spent the track supply on whatever was legal. - Teach the bot what a siding is for. Nose coupling (§A.3) implemented, so approach direction decides which car is droppable; the bot runs around rather than setting out, when the drop can follow. Switching activity transformed, revenue unchanged.
- Curve geometry. Curves were topologically identical duplicates of turnouts, and nothing reached north, so a district could only be a vertical column. Now two-port rotatable arcs.
- Q10 — when track may be laid. During the "draw a card" option, one piece a turn. Track was a card when §6.2 was written; a 26-piece supply has no hand to bound it.
- Q11 — which way a Heavy Grade climbs. Answered from the card: it prints "(Up)" and "Player sets orientation", so it is a property of the placed card, not a compass constant.
- §6.2's reshuffle. Implemented, and the
deckReshuffledevent it had already declared and narrated — but never emitted or reduced — is now real. Not yet reached in play: solitaire Campaign ends with 168.8 of 243 in the deck and four-player Campaign with 86.6, and no run of any length has emptied it. It is a safety net rather than a live mechanic today, which is worth knowing before tuning draw rates. - Q12 — Whistle Post lock-in. Players always start at a Whistle Post; office density doubled instead.