Files
station-master/CHANGELOG.md
T

26 KiB
Raw Blame History

Changelog

Detail behind each commit. Commit messages stay high level; the reasoning, the measurements, and the things that turned out to be wrong live here.

Measured figures are 100 solitaire Standard games with the developer bot unless stated otherwise. The target is 20 Revenue over 5 Days.


Unreleased

Rolling stock was leaving the game

Four things were tried against bot revenue. One of them was worth more than everything else in this changelog combined, and it is the one that had been ranked third and predicted not to matter.

The Classification Yard was write-only. Seven places pushed cars into it — retired trains, collisions, unjams, set-outs — and nothing in the engine ever read it. Rolling stock drained one way out of the game: 30 cars dead by the end of a game, 37% of the 80 dealt at setup. Gap 2c says engines and cabooses return to the Division Yard and everything else to Classification; the recovered rules never say how Classification empties.

ASSUMPTION, flagged rather than derived: sorting cars for redistribution is what a classification yard is for, and a Day is its natural cycle, so they now return at the Day boundary. This is a rules decision that wants confirming against the source.

Paired over 400 seeds: +2.32 ± 0.52, t = 8.79, 194 seeds better against 44. Revenue 6.70 → 9.02.

Why it was mis-ranked is worth recording. The Division Yard does not run dry in five Days — 16.6 loaded freight cars remain, empty in 2 games of 100 — so it was reasoned that supply could not be binding. The aggregate was never the point: what starves freight is not having the RIGHT commodity at the moment a green box needs stocking, and returning classified cars keeps the mix alive.

Three things that did not work, kept because the measurement is the result

Measured paired over the same 400 seeds, which is the only way to see effects this size.

Capping the draw option: worth nothing. 62% of all Local Operations actions went to drawing and 12.6 of 29 cards drawn were discarded, so a draw into a full hand converts straight into a discard. Refusing it: -0.10 ± 0.13, and 379 of 400 seeds byte-identical. The branch almost never fires.

Pairing the two halves of a load: worth nothing. Sampling every outbound industry at every loadUnload phase, 76% of the time it had NEITHER a stocked green box NOR a spotted car, and only 5% of Stages had a single workable facility. Letting the bot switch for a stocked box without waiting for a train at the Office changed nothing measurable — folded into the -0.10 above. The diagnosis was right and the prescription did not address it: there is usually nothing to switch.

Waking a dead branch made things worse. chooseLocalOption tested options.some(i => i.type === 'mainline.modify'), but that intent requires s.turn.option === 'draw' and the test runs while the option is still null — legalActions had already filtered it out, so the branch could never fire and never had. Rewriting it to check the hand cost -0.64 ± 0.32 (t = -3.90), 104 seeds worse against 35. Its comment claimed the value compounds like a train card's; it does not. The branch is now deleted, with the measurement in its place.

The Enhancements other than Interlocking are worth exactly nothing — and cost nothing either: -0.01 ± 0.41 (t = -0.04) when the bot is forbidden to place any of them. Left as they are; there is nothing to gain by restricting them. Note the first attempt at this experiment showed 0/400 seeds changed, which was the experiment failing rather than the answer: a card.play fallback with no placement filter was still placing them.

before after
revenue (200 games) 6.7 8.7
wins 6.5% 12%
freight loads 2.8 3.2
cars dead in the Classification Yard 28.4 ~0

A Running Track with nowhere to put an Enhancement

Every penalty in the game has one cause. Across 100 games, all 48 were collision: no free A/D track — 2.70 revenue a game, 27% of gross, concentrated in about a fifth of games and responsible for the −47 tail.

The rules already answer it. Interlocking prints "may stop an inbound train on the Limit Track", and advance.ts:621 holds the train at the Limits instead of colliding. It had never once been placed. Nor had any train ever been held at the Limits.

The reason was structural and nothing to do with Interlocking. The bot builds minimal two-arc run-arounds — an ne arc meets an nw arc directly — so it never needed a straight and laid none: 0.00 straights on the Running Track across 100 games. Interlocking, Water Column and Telegraph all require one; Yard Office and Small Yard want a Secondary Track Straight; Telephone and Radio chain off Telegraph. Thirteen of the eighteen Enhancement cards that go on the board were unplayable. They were drawn 3.78 a game and placed 0.64.

bestTrackLay now scores one straight onto the Running Track. Enhancements placed went 0.64 → 3.01 across nine types where only Overpass had ever appeared, and Interlocking now reaches the board in about a quarter of games.

On whether it pays — measured properly, and the answer is qualified. Revenue per game has a standard deviation of ~9, so a 100-game run carries roughly ±1.0 of noise. Run against the same 400 seeds, paired:

without with
revenue 5.99 6.70
collisions cost 2.70 1.91
worst single game −47 −24
most collisions in a game 10 6
wins 5.3% 6.5%

Paired per-seed the change is +0.70 ± 0.74 (95% CI), t = 1.87 — not significant on its own. And 154 seeds improved against 165 that got worse: the mean gain comes from removing catastrophes, not from making a typical game better. What justifies keeping it is that the mechanism is measured directly and accounts for the whole effect — collision cost falls 0.79, and revenue rises 0.70.

A correction to the numbers already in this file. The per-100-game revenue figures reported for the earlier changes carry the same ±1.0 noise, so the individual steps (5.0 → 6.0 → 6.5) were stated more precisely than the sample supports. The cumulative move from 3.2 to ~6.7 is far larger than the noise and stands; the individual increments should be read as indicative only.

The bot was throwing away its own freight

Routing turned out not to be the problem, which is why it was worth measuring first. Of 1650 drops across 100 games, 940 (57%) landed on a facility that wanted the car and zero landed on one that did not. The crew makes 9.4 correctly-targeted drops a game; the one-move-only destination test costs nothing measurable. "70% of waiting loads have nothing spotted" was a misleading signal — the cars arrive.

Following the freight from the other end found the leak. Loads were being destroyed:

stockToOutbound   9.45/game
loadStarted       2.71/game
facilityUnjammed  3.10/game  from outbound  <- healthy waiting loads, discarded

facilityUnjammed from outbound splices the load out of the green box and pushes it to the classification yard. That load cost a Local Operations action to stock, so discarding it is strictly negative — and the bot did it more often than it started a load.

Two fallbacks meeting, and the root cause is a familiar one: two rules for one act. canStockProductively decided the Freight Agent option was worth taking if a matching empty car was spotted. The engine's stockOutbound additionally requires a loaded car of that commodity in the Division Yard — which the predicate never checked. So the option was chosen believing a box could be stocked when none could; the follow-through then found nothing stuck, nothing to clear and nothing stockable, and fell through to "clear whatever is stuck" with nothing stuck.

Fixed by making the predicate ask the same question the engine does, by no longer using Freight Agent as the idle default (switching at worst moves the crew toward the Office, which a train must reach to depart at all, §8.1), and by ordering the last-resort unjam by what it costs to lose — MEN|AT|WORK first, then the red box whose Revenue is already banked, and the green box last.

cars fixed freight kept
loads discarded from a green box 3.10 0.00
revenue 6.0 6.5
wins 5/100 8/100
trains scheduled 3.0 3.4
cards played 16.6 19.3

The gain is development, not freight. Loads started held at 2.70 and freight revenue at 2.6 — the recovered Local Operations actions went into drawing and switching rather than into the freight chain. Green-box stocking fell 9.45 → 6.34 because the bot no longer stocks boxes it cannot serve. The regression test asserts outbound unjams stay at zero and that genuine MEN|AT|WORK jams are still cleared, so gutting the fallback would not pass it.

Half the industries never asked for a car

Freight had not moved through two rounds of fixing the district, and this is why: three of the six industries were invisible when the bot chose what to put on a train.

wantedCars consulted a hand-written industry -> car type switch that had drifted from the sheet. It named produceShed and oilRefinery — neither is an industry — and omitted freightHouse, refinery and packingSheds, which are. An industry it could not name returned null and was skipped entirely, so it never requested a car. The Refinery is the only source of tank traffic, so tank cars boarded a train 0.07 times a game and were dropped by a crew zero times in 100 games, while 23 of 79 waiting loads sat at an industry that wanted one.

It now derives the commodities from INDUSTRY_PROFILES, which is the sheet. The switch is deleted rather than corrected — a second copy of the mapping is the bug, not the values in it.

A second, narrower collapse. facilityCarType returned carTypes[0], so the second commodity of a two-commodity industry was unreachable: a Power Plant burns coal OR oil, a Grocer's Warehouse receives dry goods OR perishables. facilityCarTypes (plural) now returns the full set, and the bot's spotting and switching checks accept any of them.

Worth stating precisely: the engine's WRONG_CAR_TYPE gate was corrected to use the full set too, but that changes nothing today — both two-commodity industries are inbound-only, so freightAgent.stockOutbound rejects them before the commodity is examined. That fix is latent. The measured gain is entirely the bot side.

facilities fixed cars fixed
revenue 5.0 6.0
wins 1/100 5/100
freight revenue 1.3 2.6
freight share of gross 25% 37%
loads completed 1.27 2.60
tank cars dropped 0.00 0.62
reefers dropped 0.04 0.67

Both regression tests were checked against the bug they guard: restoring the stale map fails the tank assertion, and collapsing carTypes to its first entry fails the profile assertion.

Still 70% of waiting loads have nothing spotted at all. The commodity mix is right now; the volume reaching the industry tracks is not. That is a routing question — which facility the crew takes a car to — rather than a car-choice one.

Putting the industries on the siding

The run-arounds were being built and the industries were somewhere else — facilities sitting on one stayed at 0.00 a game even at 91 districts in 100 with a closed loop. The cause was blunt: facility placement was options.find(i => i.placement !== undefined), the first legal square the generator happened to list, unscored, while track laying had sixty lines of scoring beside it.

Facilities are now scored, on the thing that decides whether a crew can serve them at all: a placement in line with a siding still being built becomes part of the loop itself (its own through track is a segment), so the crew reaches it from either end and can pass its standing cars (§A.5). Merely touching reachable track is worth less; the Running Track is a penalty, because a car left standing there is hit by the next arrival (§11.2).

A second bug surfaced immediately, and it is the interesting one. Scoring facilities onto the siding row sent run-arounds down, 91 games in 100 to 36 — the industry took the square and the loop stopped closing around it. runsAcross tested the card's KIND ("a track straight running east-west"), so an industry standing in the line read as a dead end and the run refused to extend through it. It now asks the card's PORTS instead. A Facility carries its own rails (§11.2), and so do the Office and a Limits sign; what matters is whether a port faces this way.

The reachability walk is now shared between track laying and facility placement rather than written twice — two copies would eventually disagree about whether a district connects, which is the one thing both decisions rest on.

before sidings sidings fixed facilities fixed
facilities on a run-around 0.00 0.00 1.08
games with a run-around 0/100 91/100 71/100
revenue 3.2 4.1 5.0
collisions — — 0.4 (was 0.6)
track pieces spent 19.5 16.1 13.6

Run-arounds fall from 91 to 71 because facilities now compete for the siding squares — which is the trade being made deliberately: a loop with an industry on it is worth more than an empty one.

Freight did not follow. Loads completed 1.42 → 1.27 and freight's share of gross 31% → 25%; the revenue gain is passengers and fewer collisions. Of facilities holding a load, 77% still have nothing spotted at all. The industries are now reachable and the right cars still are not arriving — which is the car-selection problem in TODO, untouched by any of this.

The bot was building stubs, not sidings

Measured first: 0 run-arounds in 100 games. A run-around is the engine's own definition of a useful siding (track.ts) — double-ended, both ends reaching the main, and §A.5's facing-point move is impossible without one. Every district the bot built was dead-end stubs, 3.86 of them a game, plus 2.89 cards below the main that reached nothing at all. Before trusting a zero the detector was handed a run-around built on purpose and found it from both ends.

Three bugs, all the same shape: scoring on local form without checking it reaches anything.

  • The +12 rule was commented "close the loop back up to the main: the run-around is complete" and only tested that a neighbour ran east-west — never that the card above had a south port to join. The siding terminated in an arc pointing north into empty space. It now requires a way up above, asked of the engine's hasPort so the Office counts too; divergesSouth had looked for a turnout and missed the one way down that is on every board. An arc's facing also decides what it meets — nw joins west, ne joins east — so closing from the wrong side connected nothing.
  • The east-west extension had no stopping condition, so the run went on past the last column it could rejoin at, in 96 of 100 games. The loop then missed by one card.
  • bestTrackLay never declined. This was the one that actually mattered, and the first two fixes barely moved the overshoot without it: the function returned its best-scoring option unconditionally, so once the useful squares were taken it kept laying track because track was legal. Bonuses are now tracked apart from the distance score, and a piece that earns none is not laid — the Stage falls through to playing a card instead.

Anchors also have to be reachable from the main now. A stranded east-west straight made both its neighbours look like legal extensions, so a fragment joined to nothing grew in both directions.

before after
games with a run-around 0/100 91/100
track laid east of the last way up 96/100 0/100
track pieces spent 19.5 16.1
revenue 3.2 4.1
freight revenue 2.6 3.8
cars dropped 11.6 20.9
loads completed 0.99 1.42

Two regression tests, both walking the district with the engine's own exitsFrom so they cannot credit a connection §A.1 forbids: one asserts run-arounds get closed, the other that no track is spent east of the last column with a way up.

Still open. Facilities sitting on a run-around: 0.00. The loops get built and the industries are not on them, so the run-around is not yet paying for itself in freight — which is the next thread, not a finished one.

Freight, drawn where the work happens

The load pipeline moved onto the card. A load crosses green → MEN | AT | WORK → a spotted car, and that journey is freight. It was drawn only in the side panel, so a Laborer action — the whole of the freight game — changed nothing on the card the player was looking at. The squares now sit under the track, which is where the printed cards put them and why they are printed at all. The tooltip names the same thing in words: which square the load is on, and what the next Laborer action does with it.

A collision found while placing them. overpass and facingPointLocks are placed onCard, so they can land on a facility, and the enhancement label's baseline ran straight through the new squares. The label moved into the gap between the crew tray and the pipeline; a test now asserts the two do not overlap rather than trusting the two constants to stay apart.

Teaching the bot to use a siding

Nose coupling, which the rules had and the engine did not. §A.3: "engines also have couplers on the front end, so a train can pick cars up onto its nose". Coupling always appended to the back, so which end cars landed on did not exist — and that is precisely what a run-around is for. Cars met running FORWARD now couple in front, cars met BACKING couple behind, so the approach decides which car is next off the tail:

running forward -> reefer, boxcar, hopper   next off: hopper
backing up      -> boxcar, hopper, reefer   next off: reefer

The bot prefers a run-around to setting a car down, since it keeps the car — but only when the drop can actually follow. Without that guard it ran the loop for its own sake (8 a game, 63 Moves), which the shuttling regression correctly failed.

A serious bug found on the way. The carsCoupled reducer cleared every card in the Office Area, not the ones the crew ran over — its own comment said "every card along the path" while the code iterated the whole grid. One coupling anywhere deleted every standing car and every industry track in the district, so loads worked over several Stages vanished when a crew picked up an unrelated boxcar somewhere else. The event now names the cards it lifted from.

A test replaced rather than relaxed. "Under 60 Moves a game" started failing. That threshold was calibrated when the crew coupled ~0.4 cars a game; it now couples ~8 and drops ~11, and a run-around is supposed to cost several Moves. Counting Moves was always a proxy for aimlessness, so the test now measures the thing itself — Moves per drop-or-coupling — which still fails when the bot is made to wander.

before sidings now
revenue 3.9 3.2
freight 0.9 1.0
drops per game 3.3 11.1
couplings per game ~0.4 7.9

Switching is transformed; revenue is not. The crew now does real work — roughly 4 Moves per productive act, which is what running a loop costs — but that work is not yet converting into Revenue. Where it goes next is an open question rather than a known fix.


Board rendering, three-page site, curve geometry

Drawing the board as track, splitting the site, tooltips, and the curve fix.

Board rendering

Two renderers in src/sim/board-svg.ts, chosen because they fail in opposite places:

  • Office Area — the district stays a map. Cards on a grid with the rails drawn edge to edge, so a join is rail meeting rail rather than two descriptions that happen to agree. The through rail sits at a constant height on every card, which is the alignment the printed cards use.
  • The Division — not a map but a queue of sections with hard capacities, so it is a dispatcher's diagram: one line per track, 1 of 2 free under each section, no limit — trains queue at the Division Points.

Both are self-contained — no imports, no module-level helpers — because the playable app imports them normally while the replay is a single HTML file with an inline script that cannot import anything, and embeds them via Function.toString(). One implementation either way; a second copy would eventually draw a different board from the same state.

Rails come from the engine's own connectionsFor, now exported. A drawn rail cannot claim a connection the rules do not have — visible in that Modifier cards draw no rails at all.

Capacity, confirmed from the rules and now shown: Division Points unlimited, most Mainline cards 1, Double Track and Uncontrolled Siding 2, Offices 1/2/3/4 by tier.

Three pages

index.html is now a splash with two doors. The game moved to play.html; replays.html is a directory.

A replay is a save, and a save is {seed, history}. The engine is deterministic and runs in the browser, so re-submitting the same moves rebuilds the position exactly — 18 KB instead of 3.4 MB, about 190× smaller, small enough to email. A save that would describe an impossible position cannot be replayed at all, because every step goes back through applyIntent; the viewer stops and says so rather than showing a board the rules could not produce.

Static hosting cannot list a directory, so replays/manifest.json is generated at build time from whatever is in public/replays/, with each file validated first.

Tooltips

Reference detail is read once and printing it costs the space the board and action list need. A single delegated tooltip (src/web/tooltip.ts) now carries card effects, action reasoning, and facility state. Not the native title: that waits a second, cannot be styled, and never appears for keyboard users.

Curve geometry — the significant fix

A curve was modelled as a through track plus a diverging leg, which is a turnout. Consequences:

  • Curves were topologically identical duplicates of turnouts — same connections, no distinction beyond the sharp curve's 2-Move cost.
  • Every diverging leg went south, so no piece anywhere reached north except an n-s straight, which connects n↔s and nothing else.
  • Therefore a district could only ever be a vertical column: no siding, no parallel track, no run-around, no second turnout back to the main.

The printed cards (docs/tracks.png, rows 3–4) show a curve as a single arc from edge to edge with no through track. A curve is now a two-port arc, rotatable to ne/nw/se/sw. A sharp curve is geometrically identical and costs 2 Moves. Turnouts keep §A.1 unchanged — a two-port arc has no third port, so the absent-edge rule cannot apply to it.

Verified through the engine, not the types: a crew leaves the main at a turnout, runs a siding parallel, and rejoins at the far end.

Measurements

revenue freight sidings built
before 4.1 0.5 impossible
geometry only 3.9 0.9 possible, not sought
bot builds sidings 3.3 1.1 59/60 games

Freight more than doubled and the harness recorded its first win, but overall revenue is down from geometry-alone and the bad tail worsened (−29 → −59).

Capping turnouts at one or two measured worse (revenue 2.5, freight 0.6) — more ways off the main means more industries the crew can reach, and that beats a tidy Running Track. The uncapped version stands, with that measurement recorded at the code.

The bot builds sidings but does not exploit them. It pays about 20 of its 26 track pieces for them while its switching logic still sets out dead weight on a spur rather than planning a run-around. That is the next piece of work and where the revenue should appear.


Earlier commits

Recorded from memory of the work rather than written at the time; detail thins going back.

f00255c — more fixes and tweaks

Card descriptions on every card in hand and every face-up slot, the three Local Operations options explained, the objective and pace in the header, and rotations named in words rather than "rotation 2". Build stamp added (version, git SHA, -dirty for an uncommitted tree) because nothing tracked what was deployed. Extra X22 was not a bug — a per-diem train whose card calls for a caboose and nothing else — but "loaded caboose" was.

d1f689d — name the caboose, the train and the Running Track hazard

maneuver.redFlags was labelled "Red Flags on tray3": the earlier tray-id fix checked the history log, and the action list was a surface it missed. An industry on the Running Track does not block traffic (§11.2 gives it rails) but a car left standing there is hit by the next arrival (§10).

dd300ac — fix caching, Limits placement, and explain the board better

The first deploy served a fresh index.html against a cached main.js, which threw missing element: target and never started. Every module import now carries the build tag. The Limits fix was half-done: the sign could move, but nothing forbade building past it, so one straight at (0,2) produced limits · Whistle Post · limits · straight · limits.

2eca9de — playable browser build

The solitaire game as a static site, proven to need no server: full games run with every Node global replaced by a throwing stub. Train cards stopped accepting a board placement — seed 555 had offered Extra X15 at six squares with six rotations, all identical.

160190d — the bot's reasoning in the replay

Decision panel showing what the bot chose, why, and every option it passed over, with the reasons reported by the bot itself rather than re-derived by the viewer.

d261ad7 — parking trains, freight deadlock, Whistle Post lock-in

Three faults found by measurement rather than failing tests. Trains parked because destinationsFor computed reverse as facing === 'e' ? 'w' : 'e', so a crew facing south reversed to east — a port a north-south card does not have. Freight ran at a 3% load completion rate because a load started with no spotted car parks on WORK and locks the industry track, blocking the very car that would clear it. Office density doubled after 25 of 100 games never drew a Depot and never escaped a one-track Whistle Post.