Files
station-master/TODO.md
T

294 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# To do
Things worth coming back to. Anything noted here should either get done or get an explicit decision
not to — the point is that nothing quietly evaporates.
Ordered within each section by how much it is currently costing us.
---
## Next
- [ ] **REBALANCE, once the rules are right — deliberately deferred.** Card counts, industry counts
and the track mix all need a pass together, and none of them should move until the rules stop
moving. Standing distortions to account for when it happens: offices are doubled (Q12) and
industries tripled (Gap 12), both tuned when the deck held 139 cards and **no track**; it now
holds 243 of which 104 are track, so every draw is diluted by 43% — precisely the pressure
those multipliers exist to relieve. Until then, read no balance conclusion from the revenue
numbers; they are a functionality signal only.
- [ ] **Review the standalone replay against the site's replay viewer.** `node src/sim/replay.ts
--seed 1234 --out replay.html` writes a self-contained HTML file; the site instead reads JSON
saves from `public/replays/`. Nothing links to the standalone one and its output is gitignored,
so it is a developer tool that happens to look like a product feature. It carries two panels
the site viewer does not — the bot's decision trace ("what it chose, why, and what it passed
over") and the timetable — which is debugging material rather than something a player wants.
Decide: fold the decision trace into the JSON viewer and delete the standalone, or keep it and
accept that it is a tool. No action for now.
- [ ] **The bot was partly living off an illegal placement.** Barring curves from the Running Track
(they have no east-west road and dead-end the main) cost it districts 28.0 → 19.7 cards and
revenue ~2.0 → 0.8. It has no plan for where a curve should go once the easy square is gone.
Same root cause as the two items below; fix them together, after the rebalance.
- [ ] **The bot cannot get a crew next to an industry, so Flying Switch never fires.** Industries are
now stub-only and the bot places 2.23 a game (was 3.84), in districts averaging under two rows
deep. `flyingSwitch` is exempted by name in the reachability sweep in `sim.test.ts`; deleting
that line is the test that this is fixed. Same root cause as the item below.
- [ ] **The bot does not play for a run-around any more, and revenue halved.** With track in the deck
a run-around needs a turnout, a matching curve, straights, a second curve and a second turnout,
all of the right hand, arriving in a three-card hand in a usable order. The bot holds no plan
across turns and discards a piece it cannot use immediately: run-arounds fell 70/100 → 29/100
and revenue 5.80 → 2.87. Two test floors in `sim.test.ts` are pinned below the measurement as
break-detectors rather than targets, and say so. Do the density re-measurement above first —
bot weakness and deck density are currently confounded.
- [ ] **A "load" is stored as a car, so rolling stock cannot be counted.** `outboundBox`,
`inboundBox` and `menAtWork` all hold `RollingStock`, and `freightAgent.stockOutbound` takes a
LOADED CAR out of the Division Yard to fill a green box. So a census of every holder comes to
92 against the 80 dealt at setup — not necessarily duplication, because some of those objects
are cargo in transit rather than cars, but there is no way to tell them apart. Until a load is
its own type, "is any stock being created or destroyed?" is an unanswerable question, and the
supply numbers below cannot be tuned with confidence.
- [ ] **Engines are not a SUPPLY yet, only a position.** `engineAt` now records where the engine
sits in the tray and the consist shows it, but an engine is still conjured with the tray
rather than drawn from the Division Yard and returned to it. The rules put engines in the
Division Yard alongside the cars, with a predefined number of them, so running out of engines
should be a second way trains get held — today only the Crew Tray count does that. Needs a
number to start from, then playtesting.
- [ ] **The yards are shown on the play page but not in either replay viewer.** The Frame carries
them, so it is a rendering job, not a modelling one.
- [ ] **The rolling stock supply is a guess.** `ROLLING_STOCK_SUPPLY` (coach 8+8, boxcar 10+10,
hopper 8+8, reefer 5+5, tank 6+6, caboose 6) is marked provisional in `content.ts` and was
scaled alongside the Gap 12 industry increase. Now that the Classification Yard returns stock
only when the Division Yard empties, these numbers set the real supply pressure. Adjust from
playtesting rather than theory, and watch whether industry density feels light or heavy at the
same time.
- [ ] **Heavy Grade orientation is rolled, not chosen.** The card prints "Player sets orientation",
but it is dealt during setup and setup has no decision point at all — `createGame` is a pure
function of the seed, which is also what makes a save portable. Rolled from the seed for now.
Revisit when setup gains an interactive phase; the orientation matters, because it decides
which direction climbs and therefore what Brakeman and Helpers are worth.
- [x] ~~**Measure with error bars from now on.**~~ Built: `node src/sim/compare.ts 1600 <tweak>=<n>`
runs the current bot and one variant over the same deals and reports the paired difference.
Pairing drops σ from ~9 on the level to **5.3 on the difference**, so 1600 seeds gives ±0.13 in
about 1m45s — the noise floor is now ~±0.15 rather than ±1.0. Threshold to keep a heuristic is
**t ≥ 3**, and the report prints the better/worse/identical split beside the mean, because a
mean carried by a skewed tail is a different claim from broad improvement.
- [ ] **Re-run the three "worth ~0" action-mix experiments against the new floor.** Capping the draw,
pairing the two halves of a load, and restricting Enhancements were each measured "within noise
of zero" over 400 games — but at 400 games the standard error is ±0.33, so a real +0.5 would
have looked like nothing. They are nearly free to re-run now and at least one may have been
discarded wrongly.
- [x] ~~**Confirm the Classification Yard rule against the source.**~~ Confirmed, and the guess was
wrong. The rule is: used Rolling Stock to the Classification Yard, used engines and cabooses
straight back to the Division Yard, and the Classification Yard empties ONLY when the Division
Yard is bare — then all at once. The Day-boundary version I had invented was far more generous
and worth **+2.42 revenue a game the game does not actually grant**. Corrected; revenue 9.67
→ 7.25.
- [x] ~~**Enhancements are placed but mostly do nothing.**~~ Measured: forbidding every Enhancement
except Interlocking is worth **-0.01 ± 0.41 (t = -0.04)** over 400 paired seeds. They neither
pay nor cost. Left alone. Unlocking the Running Track straight put
nine kinds on the board (telegraph 0.73, waterColumn 0.57 …), but only Interlocking has a
measured effect. The Telegraph/Telephone/Radio chain adds to the other train's number when
dispatching facing trains, which may be worth nothing in solitaire; Water Column removes a
Watertower; Facing Point Locks prevents Derail, which is multiplayer-only. Worth measuring
what each is actually worth before the bot spends actions on them.
- [ ] **The marginal Local Operations action is worth ~0, and that is the real ceiling.** Three
separate attempts to spend the 60 actions better — capping the draw, pairing the two halves of
a load, restricting Enhancements — each measured within noise of zero over 400 paired seeds.
76% of the time an outbound industry has neither a stocked green box nor a spotted car, and
only 5% of Stages have a single workable facility anywhere, yet redirecting actions at that
does nothing. Something upstream limits how much work exists to do at all; find out what
before spending more effort on the option mix.
- [ ] **Freight was stuck at ~2.7 loads a game and three fixes have not moved it.** Sidings,
facility placement, car selection and the discarded-load leak all raised revenue (3.2 → 6.5)
without raising `loadStarted` past 2.7. The chain is not leaking and the cars are arriving
correctly (57% of drops land on a facility that wants them, 0% on one that does not). The
binding constraint is now upstream of routing: 60 Local Operations actions a game, and a load
needs a stocked green box AND a spotted car AND a free Laborer to line up in the same Stage.
Measure how many Stages have all three before changing any heuristic — the answer may be that
the economy, not the bot, is what caps freight.
- [x] ~~**`stats.ts` undercounts freight.**~~ Fixed: both halves counted, freight share 39% → 49%.
Worth revisiting the **industry density** decision below, which was taken on the old number.
---
## Open questions for Jesse
Blocked on a decision, not on work.
- [x] ~~**Q13 — rear-end collisions on a Mainline card.**~~ Answered: collide on catching up.
Implemented, and not on cards that print "trains may pass". Invisible to a bot that always
denies clearance; a bot that always allows drops from 7.34 revenue to **-5.13**.
- [ ] **Poling.** The only card in the deck with no defined behaviour — the sheet records its effect
as "TBD in the source". A test asserts it stays TBD so nobody invents one.
- [ ] **Heavy Grade orientation at setup.** The card says "Player sets orientation", but `createGame`
is synchronous and has no decision point, so it is currently rolled from the seed. Should become
a real choice when setup gains an interactive phase.
---
## Balance, provisional
Numbers chosen to fix a measured problem rather than taken from the design. Revisit once the victory
target is settled and freight carries its intended share.
- [ ] **Office card density** (Depot 4→8, Station 2→4, Terminal 1→2). Chosen to remove a 25% chance
of an unwinnable opening deal. Blunt: it lifts the whole ladder and dilutes every other
category. The better answer may be fewer Terminals, a cheaper first upgrade, or more A/D
capacity at the Whistle Post itself.
- [ ] **Industry density** (9 → 27, Gap 12). Restored roughly the prototype ratio. The "freight is
only 13–18% of gross" figure that motivated this was partly a measurement bug (see the
`stats.ts` item) and partly the car-selection bug; freight now runs at 37%. Worth re-deciding
whether 27 is still the right number now that the industries are actually served.
- [ ] **Train density.** Left alone by decision, but noted: 22 train cards in 140 are drawn less often
than 22 in 115 were, and trains scheduled fell 2.9 → 2.1 as a side effect of the other density
changes.
- [ ] **The victory target itself** (20 over 5 Days). 5 wins in 100, up from 1, and the bot now does
exploit sidings — so that unclaimed gain has been claimed and the target is still missed by a
wide margin (mean 6.0 against 20). This is the next real balance question.
---
## Not yet built
- [ ] **Real audio, as committed assets.** Everything the game plays is synthesised from oscillators
(`src/web/sound.ts`), which was the honest choice for a site that fetches nothing — but it is a
placeholder, not the finished sound. Sound therefore defaults to OFF.
- **"All aboard" most of all.** It currently goes through the browser's `speechSynthesis`, so
it is whatever system voice the player happens to have — a robot, not a conductor. A real
clip is the single biggest improvement available here.
- **Find and add the rest as assets**: steam whistle, grade-crossing bell, couplers clashing,
a train pulling away. Needs licences that permit redistribution (CC0 or similar), files small
enough to commit, and a check that the "fetches nothing external" test still passes — assets
must be served from the site's own folder, never hot-linked.
- Keep the synthesised versions as the fallback for anything not sourced, so a missing file is
a quieter game rather than a broken one.
- [ ] **Regions as the primary model (the other half of §8.2).** The Division map now DRAWS regions,
deriving position from what the crossing already cost. The engine still models a crossing as a
countdown of Stages, so two things printed on the cards remain unimplemented:
- `entryPoints` is declared on every Mainline profile and read nowhere. The Heavy Grade card
has five named Start positions, and playing Brakeman is supposed to move your entry point
along the card. The engine gets the same ANSWER by taking a Stage off the crossing, which is
why the derived drawing looks right — but the mechanism is not the printed one, so a card
whose starts do not correspond to its speed would be drawn wrong.
- `implications.md` §6 calls this "the single largest mechanical gap" and asks for typed cards
with speeds and named entries, with crossing time DERIVED from the region walk.
Doing it properly changes movement, so it invalidates every balance figure — revenue 8.7, the
freight numbers, all of it — and needs a full paired re-measure over 400 seeds. Needs the
source Start-position art for the ten card types before it can begin.
- [ ] **Player settings, saved.** The district's auto-focus is the first of these: it is DISPLAY
state, so in a multiplayer game two players may reasonably want it set differently and it must
never become part of game state. It currently resets on reload. Worth a settings object in
localStorage — auto-focus mode to start with, and whatever else earns a preference — kept
strictly separate from the save, which is the seed plus the intents and has to stay portable.
- [ ] **Let the game join a call and talk to the table.** Long-term. If the game could join a Zoom,
Teams or Jitsi call and post into its chat, it could carry the whole table's shared state
without anyone alt-tabbing: the history of actions as they happen, and a prompt when someone
is holding the game up — "Now waiting on player Alice to complete the Cargo phase."
- Further out, audio into the same call: a crash when a collision happens, a bell as the Stage
clock turns over.
- Further out still, a nudge on a timer — if a player has not moved within some interval, the
game says so, by beep or by spoken line: "Still waiting on Alice to complete the Cargo
phase." That turns the turn chart's "waiting on" chip into something a distracted table
actually notices.
- [ ] **Multiplayer train make-up is a round, not one player's job.** When a new train is built,
players take turns adding cars to the consist; in solitaire one player does all of it. The
engine currently has no per-player turn within the New Train phase, so this is unbuilt rather
than wrong.
- [ ] **Action cards (10) and Space-use cards (12).** Genuinely multiplayer-only — they are played AT
an opponent. Rejected with `NOT_IMPLEMENTED`.
- [ ] **Multiplayer proper.** The engine runs 2–5 player games and the bot plays them, but there is no
server, no turn submission, and no per-player view.
---
## Smaller things
- [x] **Carry `links` forward in replay frames.** Done, and the premise was wrong in an instructive
way: measured, `links` was 5% of the `cells` payload. What actually cost was the `what` prose
(32%), the facility object stored a second time inside its own cell (24%) and the rest of the
static identity (26%). All three are interned now — 3415 KB → 1877 KB, and a round-trip test
runs the page's own unpacking function.
- [ ] **Undo is unlimited step-back, and that is a decision to revisit.** The save is the seed plus
the intents, so `undo()` replays without the last one and can walk all the way to the deal. The
RNG advances with the replay, so the same play re-rolls the same 1D12 — you cannot undo your
way to a better die. But you CAN see a train's departure Stage and then spend the turn
differently, which is an ordinary solitaire take-back and also a real information leak. Options
if it starts to feel like cheating: make the Stage boundary a commit point, or cap the depth at
the current Stage. Deliberately left open until it has been played with. Multiplayer gets
nothing until there is a proposal/agreement flow — undo there is a table decision, not a
button.
- [ ] **The 5 MB replay size limit is arbitrary.** Invented, not a browser constraint. It has earned
its place — it caught a 5.2 MB payload that turned out to be the whole grid re-serialised every
frame — but the number itself deserves a reason.
- [ ] **Curves are drawn as two straight segments meeting**, not true arcs. Fine at this size, angular
close up.
- [ ] **Wide boards scroll.** A 40-card district and a 13-section Division both need horizontal
scrolling. Legible, not compact.
- [ ] **Save/restore is not version-aware.** A save from an older ruleset stops replaying rather than
failing loudly, which is the safe direction but says little about what changed. **This has now
bitten once**: both published replays were dead — one got 42 intents into 360, the other 4 of
338 — and nothing said so; they simply ended early and looked like short games. A save should
carry a ruleset stamp and the page should say "this replay was recorded under an older
ruleset and stops at Stage N" rather than presenting a truncated game as a whole one.
---
## Done, kept for the reasoning
- [x] **Put rolling stock back into circulation.** The Classification Yard was write-only — seven
writers, no readers — so 37% of all rolling stock left the game by Day 5. Returning it at the
Day boundary is **+2.32 ± 0.52 (t = 8.79)**, the largest single change measured on this bot,
and it was ranked THIRD and predicted not to matter because the Division Yard never runs dry.
The aggregate was the wrong measure; having the right commodity at the right moment is what
counts.
- [x] **Make Enhancements reachable at all.** The bot never laid a straight on the Running Track
(0.00 in 100 games) because two-arc run-arounds do not need one — so 13 of the 18 Enhancement
cards had nowhere to go, including Interlocking, the only cure for the only penalty in the
game (`no free A/D track`, 27% of gross). One scored straight fixed it: enhancements placed
0.64 → 3.01, collision cost 2.70 → 1.91, worst game −47 → −24. Revenue +0.70 ± 0.74 paired
over 400 seeds — real but not significant alone; the variance reduction is the clearer win.
- [x] **Stop the bot discarding its own freight.** `canStockProductively` did not check the Division
Yard while the engine's `stockOutbound` does, so Freight Agent was chosen when nothing could be
stocked and the follow-through fell through to an unjam that threw a waiting load out of the
green box — 3.10 a game against 2.71 started. Now 0.00. Revenue 6.0 → 6.5, wins 5 → 8 in 100.
Also confirmed **routing was never the problem**: 0% of drops land on a facility that does not
want the car.
- [x] **Why switching work did not become Revenue.** Answered: it was the freight the crew shuffled,
not the shuffling. The chain never leaked — 95% of started loads finished — it was barely
entered, because a load needs a matching empty car spotted and half the industries never asked
for one. Three fixes later (sidings, facility placement, car selection) revenue is 3.2 → 6.0
and freight 26% → 37% of gross.
- [x] **Fix car selection.** Three of six industries were invisible to `wantedCars` — a hand-written
industry→car map naming two industries that do not exist and omitting three that do — so tank
cars were dropped **0 times in 100 games**. Derived from `INDUSTRY_PROFILES` now, and the
second commodity of the two-commodity industries is reachable. Revenue 5.0 → 6.0, freight
share 25% → 37%, wins 1 → 5 in 100.
- [x] **Put the industries on the run-around.** Facility placement was unscored — the first legal
square — so 0.00 facilities a game sat on a loop; now 1.08. The instructive part was the
second bug: scoring facilities onto the siding row dropped run-arounds 91→36, because the
anchor test asked a card's KIND rather than its PORTS and an industry in the line read as a
dead end. Revenue 4.1 → 5.0. Freight did **not** follow, which is the item above.
- [x] **Make the bot build sidings that are sidings.** 0 run-arounds in 100 games → 91. Three bugs,
all scoring on local shape without checking it reached anything; the decisive one was that
`bestTrackLay` never declined a piece, so it spent the track supply on whatever was legal.
- [x] **Teach the bot what a siding is for.** Nose coupling (§A.3) implemented, so approach direction
decides which car is droppable; the bot runs around rather than setting out, when the drop can
follow. Switching activity transformed, revenue unchanged.
- [x] **Curve geometry.** Curves were topologically identical duplicates of turnouts, and nothing
reached north, so a district could only be a vertical column. Now two-port rotatable arcs.
- [x] **Q10 — when track may be laid.** During the "draw a card" option, one piece a turn. Track was a
card when §6.2 was written; a 26-piece supply has no hand to bound it.
- [x] **Q11 — which way a Heavy Grade climbs.** Answered from the card: it prints "(Up)" and "Player
sets orientation", so it is a property of the placed card, not a compass constant.
- [x] **§6.2's reshuffle.** Implemented, and the `deckReshuffled` event it had already declared and
narrated — but never emitted or reduced — is now real. Not yet reached in play: solitaire
Campaign ends with 168.8 of 243 in the deck and four-player Campaign with 86.6, and no run of
any length has emptied it. It is a safety net rather than a live mechanic today, which is worth
knowing before tuning draw rates.
- [x] **Q12 — Whistle Post lock-in.** Players always start at a Whistle Post; office density doubled
instead.