added harness to support bots being able to run / test strategies.

This commit is contained in:
Jesse
2026-08-08 15:07:42 -04:00
parent 5d825b97d2
commit 35575147cd
9 changed files with 866 additions and 23 deletions
+19
View File
@@ -64,6 +64,25 @@ npm run typecheck # tsc --noEmit
Because Node strips types rather than compiling them, the codebase is restricted to **erasable
syntax**: no `enum`, no parameter properties, no namespaces. `tsconfig.json` enforces this.
### Measuring the bot
```sh
node src/sim/harness.ts 200 # how the bot does, with the funnel
node src/sim/compare.ts 1600 trainCapSlack=1 # one change, paired against the current bot
```
**Never judge a heuristic on an unpaired run.** Revenue has σ ≈ 9 across games, so two runs of the
*identical* bot differ by about a point through nothing but the deal. `compare.ts` gives both
policies the same seed and reports the per-seed difference, where σ is 5.3 — 1600 seeds puts the
standard error at ±0.13, in under two minutes. Keep a change at **t ≥ 3**, and read the
better/worse/identical split beside the mean: a gain carried by a few rescued games is a different
claim from one spread across the field.
Variants come from `makeDeveloperBot(tweaks)`. A tweak is **temporary** — when it measures well it
becomes the default and the flag is deleted in the same commit; when it measures badly it is deleted
with the finding recorded in `CHANGELOG.md`. A bot that accumulates switches nobody can account for
is the thing this machinery exists to prevent.
## Design notes worth knowing
- **The rules engine is pure.** No I/O, no clock, no sockets, and all randomness derives from one