Fix an inescapable debt trap the playtest logs exposed
A 182-turn session spent 30% of its turns below zero, went negative on turn 18 and never recovered, and chose the same hard-times option — "take every shift going" — thirty-seven times. The cause was a tuning change made without re-deriving what it meant. Rent went from 385 to 455 while chasing a different problem, which took the mailroom baseline from the designed -$3/day to -$13/day. Worse, at the debt that run accumulated the weekly service charge came to $180 against an escape option worth $120. The hole was inescapable by arithmetic, whatever the player did. Nothing caught it. Every archetype either optimised money or got promoted out of the problem before it bit, and the dominance test passed because rent_bounced was only 17% of turns. Fixes: mailroom wages 980 -> 1120, restoring the -$3/day baseline; the escape option 120 -> 220, so it is worth more than a week's rent; rent_bounced from weight 5000 with no cooldown to 900 on a cooldown of 3, because being broke should colour a run rather than replace it; debt_collector from weight 25 to 45, since at 25 it appeared six times in 262 turns and debt became a ratchet. Three guards so this class of bug cannot recur quietly. The baseline is now asserted directly: a test sums the pack's own upkeep over twenty fortnights and requires the mailroom to net between -90 and +10, and Dispatch to be better but not so much better that money stops mattering. That would have failed the moment rent changed. Outcome tests over simulated play are a slow and noisy way to detect a number that is simply wrong. A `lifer` archetype refuses any option that would change stage — generically, by looking for an effect on `stage` — and otherwise plays for people. It reproduces the session that found this; no other archetype can. And no archetype may spend more than a quarter of a run below zero. Hard times is a state a player passes through; living in it is the failure mode. Also fixed: the analyser reported "every ~-5 turns". One export can hold several playthroughs and turn numbers restart with each, so spans are now accumulated per run and pooled rather than measured across the seam between two. 121 tests. Reasoning in docs/DECISIONS.md §29-30. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VMSFHyVPitUoosW5wyEADj
This commit is contained in:
co-authored by
Claude Opus 5
parent
76072edf4e
commit
5c1784d990
@@ -365,3 +365,52 @@ there when given no argument. The logs are gitignored: they record what a real
|
||||
person did turn by turn, including anything they typed into the feedback box.
|
||||
That is not repository content, and making it so once would put it in the
|
||||
history for good.
|
||||
|
||||
## 29. The economy needs a test on its arithmetic, not just on its outcomes
|
||||
|
||||
A real 182-turn session spent 30% of its turns below zero, went negative on turn
|
||||
18 and never recovered, and chose the same hard-times option — "take every shift
|
||||
going" — thirty-seven times.
|
||||
|
||||
The cause was a tuning change made without re-deriving what it meant. Rent went
|
||||
from 385 to 455 while chasing a different problem, which took the mailroom
|
||||
baseline from the designed −$3/day to −$13/day. Worse, at the debt that run
|
||||
accumulated, the weekly service charge came to $180 against an escape option
|
||||
worth $120: the hole was inescapable *by arithmetic*, whatever the player did.
|
||||
|
||||
Nothing caught it. Every archetype either optimised money or was promoted out of
|
||||
the problem before it bit, and the 25%-dominance test passed because
|
||||
`rent_bounced` was only 17% of turns.
|
||||
|
||||
Three things came out of it.
|
||||
|
||||
The baseline is now asserted directly — a test sums the pack's own upkeep over
|
||||
twenty fortnights and requires the mailroom to net between −90 and +10, and
|
||||
Dispatch to be better but under +250. That test would have failed the moment
|
||||
rent changed. Outcome tests over simulated play are worth having, but they are a
|
||||
slow and noisy way to detect a number that is simply wrong.
|
||||
|
||||
A `lifer` archetype was added: it refuses any option that would change stage —
|
||||
generically, by looking for an effect on `stage` — and otherwise plays for
|
||||
people. It reproduces the session that found this, and no other archetype can.
|
||||
|
||||
And a test now holds that no archetype spends more than a quarter of a run below
|
||||
zero. Hard times is a state a player passes through. Living in it is the
|
||||
failure mode.
|
||||
|
||||
Fixes: mailroom wages 980 → 1120, the escape option 120 → 220 so it is worth
|
||||
more than a week's rent, `rent_bounced` from weight 5000 with no cooldown to 900
|
||||
on a cooldown of 3 — being broke should colour a run, not replace it — and
|
||||
`debt_collector` from weight 25 to 45, since at 25 it appeared six times in 262
|
||||
turns and debt became a one-way ratchet.
|
||||
|
||||
## 30. Gaps are measured within a run, never across two
|
||||
|
||||
`tools/analyze-log.js` reported "every ~-5 turns". One export can hold several
|
||||
playthroughs and turn numbers restart with each, so a first-seen in run two
|
||||
compared against a last-seen in run one produces nonsense.
|
||||
|
||||
Spans and pair-counts are now accumulated per run and pooled, so the figure is a
|
||||
proper weighted average of within-run gaps. It is worth stating the general
|
||||
form: any statistic over a log has to respect the run boundary, because a log is
|
||||
not one sequence.
|
||||
|
||||
Reference in New Issue
Block a user