Files
interactive-story/backend/tools/measure_bundle.py
T
parththakkar106andClaude Opus 5 a7bf47a35e Let a backup carry a story that went two ways
A bundle had one list and a forked adventure has two stories, so export was
emitting every branch's turns interleaved by index — a mangled story rather
than lost data, and unreachable only because forking has no UI yet.
`ai-dnd-adventure-v2` carries the branches, the depth each one left its
parent at, which attempt at every turn is the story, and what each node
left behind. That last one is not decoration: the after-snapshots are what
a branch switch puts back, and a bundle without them imports a tree nobody
can switch inside.

`app/bundle.py` owns both formats and nothing else knows either. The v1
reader stays — those files are already on people's disks — and it is now
the only place a `variants` array exists anywhere.

The rule the module is built on is that a bundle carries what was chosen
and never what is derived. The head branch, the fork points, the live flags
and the anchors are decisions somebody made. The lineage, the head depth,
the legacy `index` and the variant ordinals are computed from those and are
rebuilt on the way in, because a bundle is a text file anybody can edit and
a derived field shipped beside its source is a chance for the file to
disagree with itself where no read would report it.

`index` is the one that stops being academic here. It agreed with `depth`
until SP5, and this is the first writer that has to fill it for a forked
story, where two branches both hold a node at depth 4. It is allocated one
per turn instead: siblings share it, no two coordinates do.

Everything a hand-edited file can get wrong about the shape of a tree is a
400 raised before the adventure row exists, because a half-applied import
is exactly the failure this phase exists to end — a story that goes quiet.
A file wrong about which attempt is live is corrected rather than refused;
that is an invariant of the database, not of the format.

Measured on the 600-action fixture: 587 kB to 911 kB, and all of the
increase is the outcomes at 489 B a node — the coordinates themselves save
57.5 B a node against the old turn-and-variants shape. Twenty forks add
660 B. 4.3% of the import body cap.

381 tests green, 16 of them new in test_bundle_v2.py. No migration, no
vacuum owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30

150 lines
5.7 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""What a v2 bundle costs, measured on a production-sized adventure.
SP6 gave the bundle coordinates, and coordinates cost bytes: a branch number
and a depth on every node, plus the two after-snapshots a switch needs to put
back. This puts a number on that, against the v1 shape it replaces, on the same
fixture the egress work is measured on.
python -m tools.measure_bundle --actions 600 --rich
python -m tools.measure_bundle --actions 600 --rich --forks 20
Reuses `tools.stress_session`'s fixture, so read that module's warnings first:
a `--rich` run is a correctness fixture, and its byte figures are not
comparable to a plain one.
"""
import json
import random
import sys
# `stress_session` FIRST, and it is not a style preference. Importing it is what
# points `AIDND_DB_PATH` at a throwaway file, and `app.database` reads that at
# module scope — so an `app` import above this line silently runs the whole
# fixture against `backend/data.db` instead. It fails by *working*: the first
# run seeds a synthetic user and adventure into the local database and reports
# perfectly good numbers, and only the second run trips over the unique email.
from tools import stress_session as stress # noqa: I001 (see above)
from app import bundle, models, tree
from app.database import SessionLocal
def _v1_shape(v2: dict) -> dict:
"""The same story as v1 would have written it, for a like-for-like count.
One entry per turn, the siblings folded back into a `variants` array, no
coordinates and no outcomes — which is exactly what v1 could carry.
"""
turns: dict[tuple[int, int], list[dict]] = {}
order: list[tuple[int, int]] = []
for node in v2["actions"]:
key = (node["branch"], node["depth"])
if key not in turns:
turns[key] = []
order.append(key)
turns[key].append(node)
actions = []
for i, key in enumerate(order):
group = turns[key]
live = next((n for n in group if n["live"]), group[0])
actions.append({
"index": i, "type": live["type"], "text": live["text"],
"reasoning": live.get("reasoning"),
"variants": [
{"text": n["text"], "reasoning": n.get("reasoning"),
"createdAt": n["createdAt"]}
for n in group
] if len(group) > 1 else None,
"variantIndex": group.index(live),
"createdAt": live["createdAt"],
})
old = {k: v for k, v in v2.items() if k not in ("branches", "headBranch")}
old["format"] = bundle.LEGACY_FORMAT
old["actions"] = actions
old["memories"] = [
{k: v for k, v in m.items() if k not in ("branch", "depth")}
for m in v2["memories"]
]
old["memoryCursor"] = 0
old["summaryCursor"] = 0
return old
def _bytes(obj) -> int:
return len(json.dumps(obj).encode("utf-8"))
def main(argv=None) -> int:
argv = list(sys.argv[1:] if argv is None else argv)
forks = 0
if "--forks" in argv:
i = argv.index("--forks")
forks = int(argv[i + 1])
del argv[i:i + 2]
args = stress.parse_args(argv)
rng = random.Random(args.seed)
stress._build_text(args, random.Random(args.seed ^ 0x5F5F))
adv_id, _ = stress.build_fixture(args, rng)
db = SessionLocal()
try:
adventure = db.get(models.Adventure, adv_id)
if forks:
# Twenty divergences off one line, each a little deeper — the shape
# SP5 measured the fork cost on. The nodes are chosen up front and
# the session is flushed after every fork: `fork` moves a row onto
# its new branch, and with `autoflush=False` a query issued before
# that move is written still finds the node where it used to be.
root = adventure.head_branch_id
candidates = [
node.id for node in
db.query(models.Action)
.filter(
models.Action.adventure_id == adventure.id,
models.Action.branch_id == root,
models.Action.depth > 1,
)
.order_by(models.Action.depth)
.all()
]
step = max(len(candidates) // (forks + 1), 1)
made = 0
for node_id in candidates[::step][:forks]:
tree.fork(db, adventure, db.get(models.Action, node_id))
db.flush()
made += 1
db.commit()
print(f"forked {made} times")
v2 = bundle.export(db, adventure)
v1 = _v1_shape(v2)
nodes = len(v2["actions"])
turns = len({(n["branch"], n["depth"]) for n in v2["actions"]})
branches = len(v2["branches"])
stripped = json.loads(json.dumps(v2))
for node in stripped["actions"]:
node.pop("stateAfter", None)
node.pop("worldStateAfter", None)
node.pop("worldDelta", None)
v2_bytes, v1_bytes, bare = _bytes(v2), _bytes(v1), _bytes(stripped)
print(f"{turns} turns · {nodes} nodes · {branches} branches")
print(f"v1 shape {v1_bytes:>12,} B")
print(f"v2 {v2_bytes:>12,} B "
f"{v2_bytes / v1_bytes:.3f}× v1")
print(f"v2 w/o outcomes {bare:>12,} B "
f"{bare / v1_bytes:.3f}× v1")
print(f"the outcomes {v2_bytes - bare:>12,} B "
f"{(v2_bytes - bare) / nodes:.1f} B/node")
print(f"coordinates {bare - v1_bytes:>12,} B "
f"{(bare - v1_bytes) / nodes:+.1f} B/node")
print(f"import cap {bundle.__name__}: "
f"{v2_bytes / (20 * 1024 * 1024):.1%} of MAX_IMPORT_BODY_BYTES")
finally:
db.close()
return 0
if __name__ == "__main__":
raise SystemExit(main())