Files
station-master/TODO.md
T

58 KiB
Raw Blame History

To do

Things worth coming back to. Anything noted here should either get done or get an explicit decision not to — the point is that nothing quietly evaporates.

Ordered within each section by how much it is currently costing us.


Next

  • THE BOT CANNOT SPOT A CAR AT A STUB INDUSTRY, and the cut-ordering rules made that visible. Coupling is mandatory on your own square now (v0.4.7), so a crew that sets a car out between itself and the only way out picks it straight back up. At a stub industry that is every set-out the bot makes: its trains run engine-first with all four cars behind, so the tail cut always lands on the exit side. The correct play is §A.5's facing point move — couple the car onto the nose, shove it into the stub, set out off the nose, back away — which is the same cross-turn planning already recorded as out of reach of any bot two items below.

    Measured over 200 paired seeds: **-0.55 revenue** (t = -3.63) and freight revenue 1.11 → 0.56.
    Filtering self-recoupling moves out of the bot's options took recoupling from **625 of 1,029
    set-outs in 60 games to 101 of 677**, and all 101 that remain are this case. Nothing is broken —
    the game models the difficulty correctly and the bot cannot yet play it — but **every revenue
    figure in this file measured before v0.4.7 is now low by roughly half a point** and the rebalance
    pass should not read the drop as a deck problem.
    
  • WHERE THE LOCAL'S COACH STANDS WHILE ITS ENGINE WORKS (§A.4) — now hit in play, still open. Trains 7/8 print "coach must remain on station track if switching", read as "the coach is never set out". A cut comes off an OUTER end, so a coach on one outer end with the engine on the other locks the train completely: it cannot set its freight car out, and cannot uncouple to run around either, because that leaves the coach standing. Measured over 60 games — 1,181 positions where a set-out should have been possible, every one refused; no other train blocked once. Two of the six possible arrangements lock, and ENGINE boxcar coach — the one that locks — is both prototypical and what make-up naturally produces.

    **Worked around, not solved.** The make-up panel now tells the player to add the coach first
    (v0.4.6), which produces `ENGINE coach boxcar` and works. The rules question is untouched: if
    the coach may be set out **at the Office**, which is what the card's wording plainly says and
    what a real mixed train does, then the prototypical make-up works and the advice becomes
    unnecessary. That needs one exception to §A.4's blanket refusal to leave Rolling Stock at the
    Office, for the coach and only on the Local.
    
  • 3/4 EXPRESS PRINTS A RULE IT CAN NEVER USE — Jesse's call. The card says "may drop or pick up one freight car at every location" and also prints Expedite. Expedite means the train departs the Stage it arrives (Q3): it arrives in the Mainline phase, stands through Cargo, and highballs in Supervisor Shift — so it is never on the board during a Local Operations phase, which is the only phase in which freight is coupled or set out. Measured over 40 bot games: 31 Office visits, 31 of them with no Local Operations turn. X14 Fruit Growers Express is in the same position, though its "may pick up one extra loaded reefer" is only a note today.

    The four options put to Jesse, unchanged: drop Expedite from 3/4 only (the other Expedite
    trains all print "no switching" and lose nothing); leave Q3 alone and strike the freight line
    from the card; drop Expedite everywhere (it partly exists to relieve Crew Tray scarcity, so
    this needs re-measuring); or move the Express's freight budget into the Cargo phase, where an
    expedited train does still get a turn. **Nothing is broken** — this is a contradiction between
    two lines on one card, and the timing rule itself is behaving exactly as recorded.
    
  • A DISTRICT CAN NOW ONLY WIDEN AS FAR AS ITS MAIN REACHES (v0.4.8) — worth watching in the rebalance rather than acting on now. Track stays inside the Limits at every row, so extending the Running Track is the only way to buy room for sidings, and a straight laid on the sign is worth more than it was. The bot barely notices — it built outside its own Limits 5 times in 100 games — but the bot also builds close to its Office; a human building deliberately hits this on the first wide district, which is how it was reported. If territory turns out to be the real constraint on freight, this is one of the two places to look (the other is the track supply, below).

  • REBALANCE, once the rules are right — deliberately deferred. Card counts, industry counts and the track mix all need a pass together, and none of them should move until the rules stop moving. Standing distortions to account for when it happens: offices are doubled (Q12) and industries tripled (Gap 12), both tuned when the deck held 139 cards and no track; it now holds 235 of which 96 are track, so every draw is diluted by 41% — precisely the pressure those multipliers exist to relieve. The 8 sharp curves have already been taken out on that argument; offices and industries are the two left. Until then, read no balance conclusion from the revenue numbers; they are a functionality signal only.

  • RE-MEASURE THE BOT AT THE NEW DEFAULTS. Both provisional rules below are now settings on the New Game dialog rather than fixed choices, and the defaults are not what the numbers in this file were measured under: the opening hand defaults to three random cards (the prototype rule) rather than 3+3, and train revenue per transit defaults to 0 rather than 1. That second one is the big move — it was worth ~5.4 of a 7.0 mean, so the bot's revenue should fall to roughly the working freight-and-passenger economy alone, which is the number this game has actually been trying to read all along. Every mean, floor and threshold quoted below and in the tests predates it. The three revenue rates run 0–5, so the useful next step is a sweep rather than a single re-run.

  • REVIEW THE TWO NEW RULES ONCE THEY HAVE BEEN PLAYED — both went in provisional, and both are now selectable rather than fixed. Jesse's call, both implemented and measured, both flagged in rules-v0.2.md. What follows is what was measured when each was the only option.

    **The opening deal (3 track + 3 other, from two separately shuffled piles).** It did what it
    was aimed at, modestly: run-arounds **4/60 → 7/60** and districts **17.9 → 20.3 cards**, with
    revenue unmoved on its own (−0.1, inside noise). Still nowhere near the 91/100 of the
    private-supply era, so the supply question is softened rather than answered. Two things to
    watch at the table: whether opening with six against a limit of three is a real decision or
    just bookkeeping, and whether three is the right number of each.
    
    **~~One Revenue for every train that clears your section.~~ Now: one Revenue to EVERY player
    when a train completes its run.** Jesse's revision in v0.4.2. The first version paid the Office
    a train departed, which on a five-Office railroad paid five separate times for one train and
    paid most to whoever it passed first. It pays once now, when the train runs off the end of the
    Division, and it pays the whole table — getting a train the length of the railroad is the
    shared achievement, and every Office it crossed had to clear it.
    Solitaire is nearly unmoved (7.0 → 7.3 mean over 200 games) because one player's departures and
    completions run at almost the same rate; **in a multi-player game the shape is completely
    different** and needs measuring once multiplayer exists — N players × 1 per completed run
    against the old N payments per train. **The victory-target question stays live**: 20 over 5 Days
    is still reachable largely on traffic, which is either the intent or an argument for raising it
    — and at the new default of 0 per transit it is not reachable on traffic at all, which is the
    first thing a playtest should check.
    
  • BOT DRIFT ACROSS THIS RELEASE — four measurements, all for the rebalance pass. Recorded together so the pattern is visible rather than four relaxed thresholds nobody adds up: - Switching work down ~16%, 1.76 → 1.48 productive acts a game (400 games), because an expedited train now stands at the Office for a Stage instead of passing straight through, and a train on the A/D track and the Office square is in the crew's way. That is the change doing its job rather than a fault — but it is drift. Collisions also went 0.05 → 0.06 and the worst game went −3 → −9, same cause: the Office fills up. Separately, the work > 2 floor that caught this had never actually been met — it read 2.16 at 150 games and 1.76 at 400, so it was passing on which seeds the sample happened to include. Now 400 games and a floor of 1.2. - Track laid badly, 7% → 15% of pieces butting a card that cannot accept them. Forced to shed on turn one, the bot would rather lay a piece than discard it; a player would discard the ones with nowhere good to go. It also means the district-size gain from the new deal is partly padding rather than useful railroad. - Interlocking placed, 15/60 → 7/60 games. Departure Revenue pulls the bot toward other work and it spends its opening on the track it was dealt. - Aimless shuttling in 3 games of 16 — thirteen are clean, so this is a minority behaviour rather than the every-game waste the detector was written for. Each floor was moved to match what is measured, with the reasoning written into the test. None is a crisis on its own; together they say the bot spends its openings worse than it did.

  • Review the standalone replay against the site's replay viewer. node src/sim/replay.ts --seed 1234 --out replay.html writes a self-contained HTML file; the site instead reads JSON saves from public/replays/. Nothing links to the standalone one and its output is gitignored, so it is a developer tool that happens to look like a product feature. It carries two panels the site viewer does not — the bot's decision trace ("what it chose, why, and what it passed over") and the timetable — which is debugging material rather than something a player wants. Decide: fold the decision trace into the JSON viewer and delete the standalone, or keep it and accept that it is a tool. No action for now.

  • The bot was partly living off an illegal placement. Barring curves from the Running Track (they have no east-west road and dead-end the main) cost it districts 28.0 → 19.7 cards and revenue ~2.0 → 0.8. It has no plan for where a curve should go once the easy square is gone. Same root cause as the two items below; fix them together, after the rebalance.

  • THE BOT'S PRIORITIES ARE NOT THE PROBLEM — measured. Ten heuristic variations, each paired over 400+ seeds. Every reordering of what the bot prefers came out inside the noise; the only thing that moved revenue was refusing to schedule a train the Office cannot hold (+1.09 ± 0.16, t = 6.79 at 1600 seeds, revenue 1.19 → 2.28). Notable failures, all instructive: - Refusing to bury the engine costs more than it saves (−0.35, t = −2.98). It works — burial falls from 8.4 decisions a game to 0.03 — and freight halves with it, because coupling is mandatory (§A.4): the moves that bury the engine ARE the moves that pick cars up. Burial is the price of collecting, not a mistake. - Reserving Moves to get home costs 0.55 (t = −2.32), though 62 of 120 trains left on the board at game end were stranded in the district. The switching work is worth more than the departures. - Granting clearance when the train ahead has one Stage left is −0.98 (t = −5.24). Trains move in numeric order, so a follower can enter the region the leader still occupies before the leader moves. "About to leave" is not "gone". - Preferring coaches at make-up, stocking the platform first, playing Interlocking earlier, hunting the Depot in the Departments: all within noise, and three of them were exact no-ops — Interlocking sits in hand alongside a train card 0.04 decisions a game. The funnel says why: only 8% of Cargo phases have a stocked green box, and the bot already takes 42% of the turns where stocking is productive. The opportunities are not there to be prioritised better. What is left is the economy itself, which is a deck question.

  • The bot cannot get a crew next to an industry, so Flying Switch never fires. Industries are now stub-only and the bot places 2.23 a game (was 3.84), in districts averaging under two rows deep. flyingSwitch is exempted by name in the reachability sweep in sim.test.ts; deleting that line is the test that this is fixed. Same root cause as the item below.

  • THE RUN-AROUND IS OUT OF REACH OF ANY BOT, AND THE DECK IS WHY — measured, five ways. "Teach the bot to plan across turns" was tried properly and does not work. Every attempt is neutral or negative, and they fail for one reason that the numbers make plain.

    | attempt | result |
    |---|---|
    | hold ALL track for the siding | **−0.70** (t = −3.27) |
    | hold only CURVES, the closing piece | **−0.26** (t = −3.22), district 17.9 → 16.7 cards |
    | finish a run before cutting another way down | 0.00 — 398/400 games identical |
    | treat a second turnout as the closing piece | 0.00 — **400/400 identical** |
    | spend a curve only on a square that CLOSES | −0.11, and only 15 games in 400 differ at all |
    
    **The pieces never meet.** Over 12,000 Local Operations turns: a turnout and a curve are in
    hand together on **0.3%** of them, and a turnout with a MATCHING-hand curve on **0.2%** — about
    once every eight games. A run-around needs five specific pieces of the right hands in a usable
    order; the bot does not get to the two-piece prerequisite.
    
    And it is not hand pressure. The hand is FULL — mean 2.66 cards, at the three-card limit on
    78% of turns. The bot plays 11.4 track cards a game and discards 1.5, so it spends the pieces
    as they arrive because a piece that builds anything outscores holding one that might build
    more later. Holding is the only counter, and holding measures worse every way it is tried.
    
    This is a consequence of moving track into the deck, not a bot weakness: 91 run-arounds per 100
    games when track was a private 26-piece supply the player chose from, 29/100 once it was drawn,
    4/60 now. **If the run-around is meant to be the central switching puzzle — and the rules
    present it that way — the supply has to change, not the player.** Options: give track its own
    hand or yard the way the prototype did, raise the hand limit for track specifically, or print a
    siding as a single card. Nothing else reaches it.
    
  • CLEARING AN INBOUND BOX MINTS A CAR — measured at 1.29 a game against a supply of 80. Re-audited in v0.4.3: rolling stock is EXACTLY CONSERVED, 100 games out of 100, range 0..0. The old audit's premise was right — the two directions were not symmetrical — but the asymmetry has since been closed from the other end. unloadBegan and passengersDetrained now take their replacement empty OUT of the Division Yard rather than conjuring it, so a load is a car that moved rather than a car that appeared: one leaves the yard, one arrives in Classification. Jesse's description of the tabletop procedure confirms this is the intended model — the token you push along the MEN|AT|WORK sign IS the car, fetched from the yard by the Freight Agent and swapped onto the industry track at the end. Both conjuring fallbacks now throw rather than minting, so the leak cannot silently return; neither fired across the suite or 100 audited games. ROLLING_STOCK_SUPPLY is therefore unblocked — it was waiting on this and can now be tuned in the rebalance pass. (The first audit's own arithmetic was off in the same way mine was on the first attempt: cars set out on a card, in card.standing, are easy to leave out of the count and make a conserved game look like a leaking one.)

  • state = fold(events) was not true, and the docs said it was. Settled: the INTENTS are canonical. Measured before deciding — advance.ts never calls reduce, so 14 of the 46 event types are never reduced: the clock, and the entire Mainline phase, which is every train movement in the game. Folding the log rebuilds a district and not a railroad. Jesse's call, and the cheap one: the plan never needed fold — persistence is { engineVersion, seed, config, history } (multiplayer.md §10) and the wire carries Frames, not events (D2/D3), so reconnection is a fresh Frame rather than an event tail. Making the phase driver reduce would have been a rewrite of the most rule-dense code in the project to buy something nothing uses. Corrected in the README, four architecture documents and six source comments; test/events.test.ts pins the unreduced set so that closing the gap later is a deliberate act, and asserts the property that does hold. If you ever do make the phase driver reduce, that test fails and tells you which docs now understate the engine.

  • Engines are not a SUPPLY yet, only a position. engineAt now records where the engine sits in the tray and the consist shows it, but an engine is still conjured with the tray rather than drawn from the Division Yard and returned to it. The rules put engines in the Division Yard alongside the cars, with a predefined number of them, so running out of engines should be a second way trains get held — today only the Crew Tray count does that. Needs a number to start from, then playtesting.

  • The yards are shown on the play page but not in either replay viewer. The Frame carries them, so it is a rendering job, not a modelling one.

  • The rolling stock supply is a guess. ROLLING_STOCK_SUPPLY (coach 8+8, boxcar 10+10, hopper 8+8, reefer 5+5, tank 6+6, caboose 6) is marked provisional in content.ts and was scaled alongside the Gap 12 industry increase. Now that the Classification Yard returns stock only when the Division Yard empties, these numbers set the real supply pressure. Adjust from playtesting rather than theory, and watch whether industry density feels light or heavy at the same time.

  • Heavy Grade orientation is rolled, not chosen. The card prints "Player sets orientation", but it is dealt during setup and setup has no decision point at all — createGame is a pure function of the seed, which is also what makes a save portable. Rolled from the seed for now. Revisit when setup gains an interactive phase; the orientation matters, because it decides which direction climbs and therefore what Brakeman and Helpers are worth.

  • Measure with error bars from now on. Built: node src/sim/compare.ts 1600 <tweak>=<n> runs the current bot and one variant over the same deals and reports the paired difference. Pairing drops σ from ~9 on the level to 5.3 on the difference, so 1600 seeds gives ±0.13 in about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is t ≥ 3, and the report prints the better/worse/identical split beside the mean, because a mean carried by a skewed tail is a different claim from broad improvement.

  • Re-run the three "worth ~0" action-mix experiments against the new floor. Capping the draw, pairing the two halves of a load, and restricting Enhancements were each measured "within noise of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would have looked like nothing. They are nearly free to re-run now and at least one may have been discarded wrongly.

  • Confirm the Classification Yard rule against the source. Confirmed, and the guess was wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses straight back to the Division Yard, and the Classification Yard empties ONLY when the Division Yard is bare — then all at once. The Day-boundary version I had invented was far more generous and worth +2.42 revenue a game the game does not actually grant. Corrected; revenue 9.67 → 7.25.

  • Enhancements are placed but mostly do nothing. Measured: forbidding every Enhancement except Interlocking is worth -0.01 ± 0.41 (t = -0.04) over 400 paired seeds. They neither pay nor cost. Left alone. Unlocking the Running Track straight put nine kinds on the board (telegraph 0.73, waterColumn 0.57 …), but only Interlocking has a measured effect. The Telegraph/Telephone/Radio chain adds to the other train's number when dispatching facing trains, which may be worth nothing in solitaire; Water Column removes a Watertower; Facing Point Locks prevents Derail, which is multiplayer-only. Worth measuring what each is actually worth before the bot spends actions on them.

  • The marginal Local Operations action is worth ~0, and that is the real ceiling. Three separate attempts to spend the 60 actions better — capping the draw, pairing the two halves of a load, restricting Enhancements — each measured within noise of zero over 400 paired seeds. 76% of the time an outbound industry has neither a stocked green box nor a spotted car, and only 5% of Stages have a single workable facility anywhere, yet redirecting actions at that does nothing. Something upstream limits how much work exists to do at all; find out what before spending more effort on the option mix.

  • Freight was stuck at ~2.7 loads a game and three fixes have not moved it. Sidings, facility placement, car selection and the discarded-load leak all raised revenue (3.2 → 6.5) without raising loadStarted past 2.7. The chain is not leaking and the cars are arriving correctly (57% of drops land on a facility that wants them, 0% on one that does not). The binding constraint is now upstream of routing: 60 Local Operations actions a game, and a load needs a stocked green box AND a spotted car AND a free Laborer to line up in the same Stage. Measure how many Stages have all three before changing any heuristic — the answer may be that the economy, not the bot, is what caps freight.

  • stats.ts undercounts freight. Fixed: both halves counted, freight share 39% → 49%. Worth revisiting the industry density decision below, which was taken on the old number.


From playtesting, 2026-08-12

Jesse played and reported nine things. All of them are now done — the entries are kept because each carries the decision behind it, and two of the nine turned out not to be bugs. The measurements and what went wrong on the way are in CHANGELOG.md.

  • LEFT AND RIGHT ARE ON THE WRONG DIAGONAL — for turnouts and for curves, the same way. The engine's left turnout is {stem:'w', through:'e', diverge:'s'}: a train entering at the points from the west heads east and the diverging route leaves to its right. The engine's left curve is arc sw, which turns an eastbound train right as well. One consistent sign error in the hand↔diagonal mapping, and it mislabels every track card a player ever holds. The artwork is right and does not change — board-svg.ts:462 draws rails from connectionsFor() geometry alone, so only words are wrong. Decision: flip the hand value on the TRACK_CARDS rows AND the two variant tables in the same commit, so the code keeps speaking left/right like the physical supply and now means it. Keep the row order in TRACK_CARDS untouched: setup.ts:95 builds the deck by iterating that array, so flipping only the labels leaves pre-shuffle slot 32 holding a ne_sw curve before and after, and variantsFor(…)[0] still 'sw' — same seed, same board, and every published replay still plays. Also: track.ts:126 arc fallback, track.ts:191-212 doc block, view.ts:1041 and view.ts:1208 diagonal phrases, bot.ts:837-852 arcInHand, seven test files, and the prose plus ~14 data-tip="Turnout · left" strings in docs/design/track-geometry.html — which has no generator and must be hand-edited. Verify by fingerprinting a fixed seed's board before and after: identical geometry, different words.
  • A turnout should be playable as an UPGRADE, on top of a card already down. On a straight, or on a curve whose arc matches the turnout's diverging leg. Nothing like this exists — the only "upgrade" in the game is the Office tier change, which is explicitly not a card swap (apply.ts:1290), and canPlaceAt hard-stops at if (existing && !isMovableSign) return false (track.ts:518). Decision: cars AND enhancements both block it — standing.length > 0 → UPGRADE_OCCUPIED, enhancements.length > 0 → UPGRADE_ENHANCED, so an Interlocked straight stays a straight. Express the curve rule on arcs, not hands, so it survives the flip above. No extra connection requirement: all four turnout orientations are port supersets of a straight and of any same-arc curve, so an upgrade can never sever an existing join, and the new leg is allowed to dangle — that is what it is for. The lifted card leaves play, which is already how board cards behave (apply.ts:1249 salvages only cards that were not placed). Reuse card.play with a placement on an occupied square; legal.ts:252 must offer those squares for turnouts, and the attachments set at legal.ts:137 is already exactly that list. The other half of this report needs no work: a turnout carries the through route (carriesThroughTrack), so it is already legal at a Limit, along the Running Track and on Secondary Track. Confirmed, not re-investigated.
  • A Depot shows three MEN AT WORK boxes it can never work. Offices are Passenger Facilities — no freight — but setup.ts:176 gives every one a three-slot menAtWork array, and the three renderers loop it with no guard while the green and red rows beside them are guarded (board-svg.ts:548, panels.ts:218, replay.ts:522). Decision: all three tiers — Depot, Station and Terminal are all Passenger Facilities and all get laborers: 0, so the boxes are inert on every one. Fix it in the model, not the renderers, so it cannot reappear in a fourth place: make menAtWork nullable and null for passenger facilities, which is what the "Freight only" comment at state.ts:153 has claimed all along. Freight handling must be neither allowed nor displayed there. Two more artifacts of the same "an Office is a Facility" modelling go with it: board-svg.ts:562 calls a Depot "SHIPS + RECEIVES", an industry's flow word, and board-svg.ts:599's Math.max(1, trackCap) draws it a siding slot although its industry track has length 0. FacilityView needs a kind field; it has no freight/passenger flag today.
  • Two Ice Houses can be built in one district. Industries already ban duplicates per Office Area — isLockedOut (apply.ts:1774) covers the Mine Tipple half of the report — but Ice House is a Modifier, not an industry, and checkPlay's modifier branch (apply.ts:662) has no duplicate check at all. Decision: extend the ban to modifier kinds, leave enhancements alone. An Interlocking is a plant at one junction, so a second on another straight is a different installation, and the Telegraph → Telephone → Radio chain is already gated per card. Expect modifiers with copies > 1 to go partly dead in solitaire, where there is one Office Area — correct, since the spare copies exist for other players' districts.
  • A drawn card lands at the far end of the hand with nothing to mark it. apply.ts:1219 pushes, the hand renders in state order (game.ts:355), and with flex-wrap the newest card is exactly where the eye is least likely to be. Decision: reverse in the display layer, not the engine — view.ts:898-899 (reversing hand and handWhat identically, or they desync) covers the replay viewer and the standalone replay together, and game.ts:355 covers the play page. Keeping hand.push means the bot's option-iteration order does not move, so every revenue figure in this file stays comparable; an engine unshift would invalidate the lot. The marker is play-page only — a Frame carries card names, not ids, so a replay cannot say which card arrived that step. Follow the game.scheduled precedent for a justDrawn field, but make the badge persist rather than flash: it says which card is new, not something just happened. A static ::before badge as .handcard.target does it (panels.ts:262), no keyframe — the innerHTML rebuild would restart an animation on every render. Clear it in renderUndo() and after fromSave(), or a fresh page load badges last session's draw.
  • The inbound boxes render green in the side panel. Green is outbound and red is inbound everywhere the colour carries direction — board-svg.ts:528-545, replay.ts:452, rules §9.1 — except panels.ts, whose shared boxes() helper (panels.ts:162) emits class f for every filled box regardless of direction, so the red row at panels.ts:222 comes out green. The board SVG on the same page draws it correctly, which makes the panel actively contradict the board. Give boxes() its class from the caller as replay.ts already does, add .box.r in board-svg.ts:820's palette, and take the siding off green in both panels and replay.ts:525.
  • An Interlocking on the board is a bare label, and it does nothing. The enhancement text is drawn with no tooltip (board-svg.ts:651) while the copy already exists as data in ENHANCEMENT_CARDS (content.ts:589). Done: every card prints its effect, and the one nothing reads says so. CORRECTED — the first pass had five of the ten statuses wrong. It claimed only Telegraph/Telephone/Radio were live, because the survey grepped for four helper function names and read "no match" as "no implementation". In fact seven are live: those three plus Interlocking (advance.ts:770), Yard Office (advance.ts:750), Small Yard (apply.ts:405) and ABS Signals (advance.ts:599, stored on the Mainline node). Facing Point Locks and Water Column are wired but dormant in solitaire; Overpass alone has no code path at all. The shipped tooltip briefly told players four working cards did nothing, which is worse than the bare label it replaced — enhancements.test.ts had passing tests for all four the whole time.
  • THE INDUSTRY TABLE STILL DISAGREES WITH THE CARD REFERENCE — two items left, Jesse's call. v0.4.7 corrected the DIRECTIONS: the Grocer's Warehouse and the Oil Refinery are flow: 'both', as card-reference.md always said, which is what let an Ice House finally give a Grocer's its outbound slot. Two discrepancies remain and both are deliberate for now. (a) Base capacities. The reference prints Grocer's 2/2 with 2 Laborers and the Refinery 2/2 with 3; the engine gives every industry 1 per direction it allows, Mine Tipple included. Raising one alone would be a balance change rather than a correction. (b) The Freight House card. The reference is explicit — "'Freight House' is not a card. It is the collective term for a freight facility that loads and unloads" — and the engine deals 6 copies of one. Removing them is a deck-composition change worth measuring, not a quiet delete.
  • A modifier's grant can land on a direction its host cannot use, and nothing says so. THE "NO BUG" VERDICT BELOW WAS WRONG, and v0.4.7 corrected it. The reasoning was that a Grocer's Warehouse is inbound-only. It is not — the card reference says "Both" — so the grant was being dropped on a direction the facility should have had. Reported again in play as "grocer's warehouse didn't get extra outbound slot for truck dock". The suppression machinery itself was right and is kept: it still fires for a passenger Modifier beside a Whistle Post, which is not a Passenger Facility — and v0.4.7 gives that grant BACK when the Office is upgraded, which it never used to. Original note follows. Reported as "Ice House added the laborer but not the outbound slot" — checked, and there is no bug: Ice House prints +1 outbound, a Grocer's Warehouse is flow: 'inbound' so allows.outbound is false, and the capacity was raised on a direction that can never render or be stocked. The laborer arrived because laborers have no direction gate. Decision: an industry's printed flow is absolute — drop the grant on hosts that cannot use it rather than opening the direction up. So applyModifier (apply.ts:1788) gates each capacity grant on allows and grows industryTrack.length only by what was actually applied. Audit all 17 profiles for the same trap: iceHouse, truckDock and forklifts all print addOut: 1 and list grocersWarehouse; waitingArea, restaurant and hotel print addOut: 1 for hosts: ['office'], which includes a Whistle Post. Then say it in both directions — the hand tooltip naming which printed hosts cannot use which half (computable from the profiles, no host on the board needed), and the panel's "prints N, Modifiers add M" line showing a suppressed grant as suppressed instead of quietly omitting it. That delta display already cites the Ice House as the bug that motivated it.

Open questions for Jesse

Blocked on a decision, not on work.

  • Q13 — rear-end collisions on a Mainline card. Answered: collide on catching up. Implemented, and not on cards that print "trains may pass". Invisible to a bot that always denies clearance; a bot that always allows drops from 7.34 revenue to -5.13.
  • Nine of the twelve special-train rules are declared and read by nothing. Done — all nine enforced, and one of them deleted instead. copiesNextScheduled was never carried by any train card: a Second Section is a Maneuver with its own working intent, so the flag was an unreachable second description of an existing mechanic. Cost 0.8 revenue and half the wins (8.0 → 7.2, 14/200 → 6/200), which is what enforcing restrictions does.
  • THE LOCAL'S COACH HAS NOWHERE TO STAND, AND THAT IS A §A.4 QUESTION. Trains 7/8 print "coach must remain on station track if switching", which should mean the coach is set out at the station while the engine works. It cannot be: §A.4 refuses the Office square to every drop ("Rolling Stock may not be left there"), so "only at the Office" and "nowhere" are the same rule. It is enforced as "the coach is never set out", which keeps the Local from abandoning it at an industry but loses the drop-and-collect pattern a real local works. Decision needed: should the Office square — or a station track beside it — accept a parked coach? That is a change to §A.4, not to the train card, and it would also give the A/D tracks something to do.
  • Poling. The only card in the deck with no defined behaviour — the sheet records its effect as "TBD in the source". A test asserts it stays TBD so nobody invents one.
  • Heavy Grade orientation at setup. The card says "Player sets orientation", but createGame is synchronous and has no decision point, so it is currently rolled from the seed. Should become a real choice when setup gains an interactive phase.

Balance, provisional

Numbers chosen to fix a measured problem rather than taken from the design. Revisit once the victory target is settled and freight carries its intended share.

  • Office card density (Depot 4→8, Station 2→4, Terminal 1→2). Chosen to remove a 25% chance of an unwinnable opening deal. Blunt: it lifts the whole ladder and dilutes every other category. The better answer may be fewer Terminals, a cheaper first upgrade, or more A/D capacity at the Whistle Post itself.

  • Industry density (9 → 27, Gap 12). Restored roughly the prototype ratio. The "freight is only 13–18% of gross" figure that motivated this was partly a measurement bug (see the stats.ts item) and partly the car-selection bug; freight now runs at 37%. Worth re-deciding whether 27 is still the right number now that the industries are actually served.

  • Train density. Left alone by decision, but noted: 22 train cards in 140 are drawn less often than 22 in 115 were, and trains scheduled fell 2.9 → 2.1 as a side effect of the other density changes.

  • The victory target (20 over 5 Days) is out of reach by a factor of about four, and the Office ladder is why. Measured over 800 games with the tuned bot, which no longer throws revenue away on collisions (0.0 a game, down from 0.4):

    | trains scheduled | games | revenue |     | Office reached | games | trains | revenue |
    |---|---|---|---|---|---|---|---|
    | 0 | 110 | 0.67 |  | Whistle Post | 297 | 0.81 | 0.62 |
    | 1 | 379 | 1.69 |  | Depot        | 272 | 1.49 | 2.92 |
    | 2 | 234 | 3.72 |  | Station      | 176 | 1.89 | 4.36 |
    | 3 |  68 | 5.68 |  | Terminal     |  55 | 1.93 | 5.04 |
    | 4 |   9 | 5.78 |  | | | | |
    
    Revenue is almost exactly linear in trains scheduled — about **1.9 a train** — and trains are
    capped by A/D capacity, which is the Office tier, which is a card you have to draw. So the
    whole economy hangs off one valve: **37% of games never leave the Whistle Post and earn 0.62;
    53% of all games earn nothing at all.**
    
    Extrapolating the line, 20 Revenue needs roughly **11 trains and therefore 11 A/D tracks**. A
    Terminal has four. The target is not merely missed, it is structurally unreachable under this
    deck at this Office ladder — no amount of bot skill closes it, and the best game seen in 800
    was 26 against a median of 0.
    
    The three ways out are all yours to choose between, and they are different games:
    1. **Lower the target** to what a 5-Day game can produce (6–8 looks like the honest number).
    2. **Open the valve** — more Office cards, or a cheaper first upgrade, or more A/D capacity at
       the Whistle Post, so the ladder is climbed rather than drawn.
    3. **Raise revenue per arrival.** It is 0.46 today; each arrival can in principle pay 2 for
       passengers alone. That is the freight/passenger conversion problem, not the traffic problem.
    
    Nothing here is a bot weakness any more, which is what this measurement was waiting on.
    

Not yet built

  • INVESTIGATE: how would a player publish a replay so other people can watch it? Today "Save replay" downloads a JSON file to the player's own machine, and the only way it reaches the site is by sending it to Jesse to drop into public/replays/ and redeploy. The question is what a self-service version would look like.

    **The constraint.** The site is fully static — `dist/` is uploaded to File Browser and Start9
    Pages serves the folder — and the replay list is a build-time `manifest.json` because static
    hosting cannot list a directory. So publishing needs something that accepts a write.
    
    **The one measurement that matters:** a full 5-Day game is **451–1017 bytes** compressed
    (brotli), about **600–1150 characters** base64. A save is the seed plus the intents and the
    engine recomputes the board, so a whole game fits in a URL.
    
    Four shapes, roughly costed:
    1. **Share by link, no server (~1–2 hours).** Put the compressed save in the URL fragment
       (`replays.html#s=…`); "Share replay" copies a link and anyone opening it watches the game.
       The viewer already parses saves and already has a file-open path, so this is compression, a
       hash reader and a copy button. The fragment never reaches the host. It is a link rather than
       a gallery: nobody discovers a game they were not sent.
    2. **A write endpoint (a day or two, and it is a service).** Accepts a POST, validates the save
       by replaying it through the engine — `save-replay.ts` already does exactly that check —
       writes the file and regenerates the manifest. The work is the surround: auth or rate
       limiting, abuse handling for a public write, CORS, and a deploy story. It also ends "static
       hosting is all this needs", which has been load-bearing.
    3. **Browser writes to File Browser directly — rejected.** It needs FB credentials in a static
       page, so anyone viewing source gets write access to the whole File Browser, and the manifest
       would need a read-modify-write from the browser that loses a save when two people publish at
       once.
    4. **Curated, manual (zero code).** What happens today, and it composes with (1): players send
       links, Jesse publishes the good ones.
    
    **The question behind the question is whether a gallery of strangers' games is wanted on a
    personal StartOS box at all.** If it is, (1) is the piece (2) would need anyway, so it is the
    right thing to build first either way.
    
  • Real audio, as committed assets. Everything the game plays is synthesised from oscillators (src/web/sound.ts), which was the honest choice for a site that fetches nothing — but it is a placeholder, not the finished sound. Sound therefore defaults to OFF. - "All aboard" most of all. It currently goes through the browser's speechSynthesis, so it is whatever system voice the player happens to have — a robot, not a conductor. A real clip is the single biggest improvement available here. - Find and add the rest as assets: steam whistle, grade-crossing bell, couplers clashing, a train pulling away. Needs licences that permit redistribution (CC0 or similar), files small enough to commit, and a check that the "fetches nothing external" test still passes — assets must be served from the site's own folder, never hot-linked. - Keep the synthesised versions as the fallback for anything not sourced, so a missing file is a quieter game rather than a broken one.

  • Regions as the primary model (the other half of §8.2). The Division map now DRAWS regions, deriving position from what the crossing already cost. The engine still models a crossing as a countdown of Stages, so two things printed on the cards remain unimplemented: - entryPoints is declared on every Mainline profile and read nowhere. The Heavy Grade card has five named Start positions, and playing Brakeman is supposed to move your entry point along the card. The engine gets the same ANSWER by taking a Stage off the crossing, which is why the derived drawing looks right — but the mechanism is not the printed one, so a card whose starts do not correspond to its speed would be drawn wrong. - implications.md §6 calls this "the single largest mechanical gap" and asks for typed cards with speeds and named entries, with crossing time DERIVED from the region walk. Doing it properly changes movement, so it invalidates every balance figure — revenue 8.7, the freight numbers, all of it — and needs a full paired re-measure over 400 seeds. Needs the source Start-position art for the ten card types before it can begin.

  • Player settings, saved. The district's auto-focus is the first of these: it is DISPLAY state, so in a multiplayer game two players may reasonably want it set differently and it must never become part of game state. It currently resets on reload. Worth a settings object in localStorage — auto-focus mode to start with, and whatever else earns a preference — kept strictly separate from the save, which is the seed plus the intents and has to stay portable.

  • Let the game join a call and talk to the table. Long-term. If the game could join a Zoom, Teams or Jitsi call and post into its chat, it could carry the whole table's shared state without anyone alt-tabbing: the history of actions as they happen, and a prompt when someone is holding the game up — "Now waiting on player Alice to complete the Cargo phase." - Further out, audio into the same call: a crash when a collision happens, a bell as the Stage clock turns over. - Further out still, a nudge on a timer — if a player has not moved within some interval, the game says so, by beep or by spoken line: "Still waiting on Alice to complete the Cargo phase." That turns the turn chart's "waiting on" chip into something a distracted table actually notices.

  • Multiplayer train make-up is a round, not one player's job. When a new train is built, players take turns adding cars to the consist; in solitaire one player does all of it. The engine currently has no per-player turn within the New Train phase, so this is unbuilt rather than wrong.

  • MULTIPLAYER — three things deliberately deferred while planning the server. Decisions and reasoning are in docs/architecture/multiplayer.md §11; these are the ones left open. - Bots should take minimally damaging, defensive actions when a player steps away, so a game is not permanently halted. Deliberately NOT automatic today: a turn timer forfeiting is different from a bot competing, and the clearance ruling is the one decision that changes another player's score. Bots fill empty seats at lobby time only (D8). - Let a player resign and hand their railroad to a bot to finish. Same care needed as above, but it is consented rather than imposed. - A forcing turn timer — explicitly NOT in the design. lobby-and-sessions.md §5 used to specify one: on expiry the server took "the safest legal action", including denying a clearance. Cut in the review, because it is the same objection as a bot playing for an absent player — the clearance decision changes somebody else's score, so anything that answers it automatically changes the game. Explore later if halted games turn out to be a real problem at a real table; the reasoning worth keeping is that deny is the safe default, since a held train costs a Stage and a wrecked one costs 5 Revenue and feeds the collision floor. - The opening D12 for the Eastern Division Point (§4.4) decides nothing. Done in v0.4.1 — it orders the whole chain now, west to east by ascending roll. The lobby still owes it a display: state.openingRolls is kept so clients can show the rolls forming the chain rather than only the result (lobby-and-sessions.md §4). - Revisit the join secret (D14). One server-wide secret, passed out of band, gates create and join. Enough for a private box, probably not enough if stationmaster.<domain> is pointed at the open internet for long. Note that one-game-at-a-time per person is expected usage and deliberately NOT enforced — enforcing it needs cross-game state whose only job is deciding when to release someone, and getting that wrong locks a player out.

  • WHY DOES A 4-PLAYER COMPETITIVE GAME END AFTER ~16 STAGES OF A POSSIBLE 60? Measured while sizing multiplayer: 8 games, all reaching Day 5, but only ~16 distinct (day, stage) pairs each and ~355 intents. Most likely the collision or revenue floor (§3.4) firing early, which would make a competitive game about an hour rather than four. Worth knowing whether that is the design working or a balance bug — it decides what a lobby should tell players about length.

  • THE 22 OPPONENT-DIRECTED CARDS — 10 Action, 12 Space-use — ARE OUT OF EVERY DECK UNTIL THEY ARE BUILT. Jesse's call. They were already cut from solitaire (Q6, no legal target with one player); they are now cut from the competitive deck too, because checkPlay answers both categories NOT_IMPLEMENTED and dealing them would make ~9% of draws reject outright. Flip opponentCardsInDeck in setup.ts when they land. They are played AT another player — Watertower, Derail, Railroad Crossing and so on — so they are genuinely multiplayer work, and three Enhancements are waiting on them: Facing Point Locks, Water Column and Overpass are wired and read, and fire only against these cards. Until then those three are dormant by design rather than broken.

  • Multiplayer proper — Phases 0 and 1 done (v0.4.0), Phases 2–6 to go. The full plan is docs/architecture/multiplayer.md §12. The engine now has seat/player separation and per-player turn state, the page renders from Frame + Menu alone and talks to a Session rather than to the engine — so a RemoteSession can be dropped in without the page changing. Still no server, no turn submission and no per-player push: that is Phase 2, and it is deliberately held until the two provisional rules have been playtested, because a rule change after the wire format is live is much more expensive than one before it.


Smaller things

  • Carry links forward in replay frames. Done, and the premise was wrong in an instructive way: measured, links was 5% of the cells payload. What actually cost was the what prose (32%), the facility object stored a second time inside its own cell (24%) and the rest of the static identity (26%). All three are interned now — 3415 KB → 1877 KB, and a round-trip test runs the page's own unpacking function.
  • Undo is unlimited step-back, and that is a decision to revisit. The save is the seed plus the intents, so undo() replays without the last one and can walk all the way to the deal. The RNG advances with the replay, so the same play re-rolls the same 1D12 — you cannot undo your way to a better die. But you CAN see a train's departure Stage and then spend the turn differently, which is an ordinary solitaire take-back and also a real information leak. Options if it starts to feel like cheating: make the Stage boundary a commit point, or cap the depth at the current Stage. Deliberately left open until it has been played with. Multiplayer gets nothing until there is a proposal/agreement flow — undo there is a table decision, not a button.
  • The 5 MB replay size limit is arbitrary. Invented, not a browser constraint. It has earned its place — it caught a 5.2 MB payload that turned out to be the whole grid re-serialised every frame — but the number itself deserves a reason.
  • The test suite fails at random under npm test, and it is the runner rather than the code. node --test test/**/*.test.ts runs the files in parallel and three suites write and read the same dist/ — the static build, the published-replay check and "the three places a game is drawn stay in step". Back-to-back full runs measured 9 failures then 0; run one file at a time and every suite passes. That is worse than a slow suite: it trains us to shrug at a red run, which is exactly how a real regression gets waved through. Give the build test its own output directory, or mark the trio to run serially.
  • Curves are drawn as two straight segments meeting, not true arcs. Fine at this size, angular close up.
  • Wide boards scroll. A 40-card district and a 13-section Division both need horizontal scrolling. Legible, not compact.
  • Every published replay was dead. All three replayed 2 intents of roughly 400 and presented as short games, exactly as the item below predicted. Re-recorded from bot games with node src/sim/save-replay.ts, which verifies each save round-trips before writing it, and harness.test.ts now fails if a published replay stops short. The version-stamp item below is still worth doing — this catches the breakage, it does not explain it to a player.
  • Save/restore is not version-aware. A save from an older ruleset stops replaying rather than failing loudly, which is the safe direction but says little about what changed. This has now bitten once: both published replays were dead — one got 42 intents into 360, the other 4 of 338 — and nothing said so; they simply ended early and looked like short games. A save should carry a ruleset stamp and the page should say "this replay was recorded under an older ruleset and stops at Stage N" rather than presenting a truncated game as a whole one.

Done, kept for the reasoning

  • Put rolling stock back into circulation. The Classification Yard was write-only — seven writers, no readers — so 37% of all rolling stock left the game by Day 5. Returning it at the Day boundary is +2.32 ± 0.52 (t = 8.79), the largest single change measured on this bot, and it was ranked THIRD and predicted not to matter because the Division Yard never runs dry. The aggregate was the wrong measure; having the right commodity at the right moment is what counts.
  • Make Enhancements reachable at all. The bot never laid a straight on the Running Track (0.00 in 100 games) because two-arc run-arounds do not need one — so 13 of the 18 Enhancement cards had nowhere to go, including Interlocking, the only cure for the only penalty in the game (no free A/D track, 27% of gross). One scored straight fixed it: enhancements placed 0.64 → 3.01, collision cost 2.70 → 1.91, worst game −47 → −24. Revenue +0.70 ± 0.74 paired over 400 seeds — real but not significant alone; the variance reduction is the clearer win.
  • Stop the bot discarding its own freight. canStockProductively did not check the Division Yard while the engine's stockOutbound does, so Freight Agent was chosen when nothing could be stocked and the follow-through fell through to an unjam that threw a waiting load out of the green box — 3.10 a game against 2.71 started. Now 0.00. Revenue 6.0 → 6.5, wins 5 → 8 in 100. Also confirmed routing was never the problem: 0% of drops land on a facility that does not want the car.
  • Why switching work did not become Revenue. Answered: it was the freight the crew shuffled, not the shuffling. The chain never leaked — 95% of started loads finished — it was barely entered, because a load needs a matching empty car spotted and half the industries never asked for one. Three fixes later (sidings, facility placement, car selection) revenue is 3.2 → 6.0 and freight 26% → 37% of gross.
  • Fix car selection. Three of six industries were invisible to wantedCars — a hand-written industry→car map naming two industries that do not exist and omitting three that do — so tank cars were dropped 0 times in 100 games. Derived from INDUSTRY_PROFILES now, and the second commodity of the two-commodity industries is reachable. Revenue 5.0 → 6.0, freight share 25% → 37%, wins 1 → 5 in 100.
  • Put the industries on the run-around. Facility placement was unscored — the first legal square — so 0.00 facilities a game sat on a loop; now 1.08. The instructive part was the second bug: scoring facilities onto the siding row dropped run-arounds 91→36, because the anchor test asked a card's KIND rather than its PORTS and an industry in the line read as a dead end. Revenue 4.1 → 5.0. Freight did not follow, which is the item above.
  • Make the bot build sidings that are sidings. 0 run-arounds in 100 games → 91. Three bugs, all scoring on local shape without checking it reached anything; the decisive one was that bestTrackLay never declined a piece, so it spent the track supply on whatever was legal.
  • Teach the bot what a siding is for. Nose coupling (§A.3) implemented, so approach direction decides which car is droppable; the bot runs around rather than setting out, when the drop can follow. Switching activity transformed, revenue unchanged.
  • Curve geometry. Curves were topologically identical duplicates of turnouts, and nothing reached north, so a district could only be a vertical column. Now two-port rotatable arcs.
  • Q10 — when track may be laid. During the "draw a card" option, one piece a turn. Track was a card when §6.2 was written; a 26-piece supply has no hand to bound it.
  • Q11 — which way a Heavy Grade climbs. Answered from the card: it prints "(Up)" and "Player sets orientation", so it is a property of the placed card, not a compass constant.
  • §6.2's reshuffle. Implemented, and the deckReshuffled event it had already declared and narrated — but never emitted or reduced — is now real. Not yet reached in play: solitaire Campaign ends with 168.8 of 243 in the deck and four-player Campaign with 86.6, and no run of any length has emptied it. It is a safety net rather than a live mechanic today, which is worth knowing before tuning draw rates.
  • Q12 — Whistle Post lock-in. Players always start at a Whistle Post; office density doubled instead.