Commit Graph
45 Commits
Author SHA1 Message Date
JesseMarkowitzandClaude Opus 5 db7b309e3d v1.1 closeout: accept integrated release validation
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-16 07:12:23 -04:00
JesseMarkowitz 87a40326a2 v1.1: harden recovery and control boundaries
WP-D and WP-E complete the planned v1.1 implementation packages.

WP-D — recovery honesty:
- backups verify the completed copy with PRAGMA integrity_check
- corruption missed by quick_check is detected by the full check
- existing good backups remain protected
- oversized exports are still delivered but declare whether this version can
  import them, while the 20 MB import limit remains unchanged
- backup was exercised through the real browser UI on both the normal campaign
  database and a campaign-shaped database over 100 MB

WP-E — control-boundary contrast:
- interactive control boundaries meet the WCAG 1.4.11 3:1 target
- the contrast audit is now a failing gate rather than an advisory
- rendered browser measurements pass for the composer, controls, tabs and nav
- text contrast and focus visibility remain intact
- owner reviewed and approved the before/after screenshots

Reports:
- planning/reports/v1.1/V1.1-WP-D-REPORT.md
- planning/reports/v1.1/V1.1-WP-E-REPORT.md

All planned v1.1 work packages A-E are now complete. Release validation has not
yet begun.
2026-09-16 05:37:13 -04:00
JesseMarkowitzandClaude Opus 5 59b5ebc2d8 v1.1 WP-C: browser release coverage
Drives in a real browser the reader workflows v1 proved only through the API
or the component suite, including an export that leaves the browser as a
file. Final run: 91 checks (the 38 existing M11 checks plus 53 new), 0 failed,
0 skipped, on the production build over trusted-LAN HTTPS.

- tools/m11_browser.py: scenarios for Retry and takes, Save Point create /
  restore / Redo, state correction (accepted, and a refused correction with
  its reason), narration length reaching each turn's prompt, failed
  generation (an unserved model blocked up front; a listed model that cannot
  narrate failing in the open) and recovery, and export download from the
  library and from campaign settings, imported into a fresh application.
  Rows are tagged M11 / WP-C and counted separately; --only for development.
  The M11 checks now wait on conditions instead of sleeping.
- tools/m11_webdriver.py: Firefox download preferences, a $HOME-only
  download folder, a download wait that ignores partial, empty, pre-existing
  and still-growing files, centred real clicks, tabs, and condition waits.
- tests/test_v11_c_browser_helpers.py: the download wait, prefs and $HOME
  guard, without a browser.
- frontend: a correction the story refused was presented as "Generation
  failed" with a Retry offer and a typed-input claim. It is now "That
  correction was not applied", not retryable, with the reason kept
  (errors.js, FailureNotice.jsx; 3 regression tests).
- DEVELOPMENT.md: the harness command, download profile and $HOME rule,
  what counts as a finished download, and the no-sleep rule.
- docs: V1.1-PLAN, VERSION v4.4, planning README,
  reports/v1.1/V1.1-WP-C-REPORT.md.

Open for the owner: "Correct" on an Important Facts row is always refused
(K1), and a partly refused correction is not reachable from the reader UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 18:29:45 -04:00
JesseMarkowitzandClaude Opus 5 0c1ba836ba v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 11:21:53 -04:00
JesseMarkowitzandClaude Opus 5 beb17ada10 v1.1 WP-B.1: diagnose independent long-term memory retention
Diagnostic only; no memory behaviour changes.

- tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage
  diagnosis (created / retained / ranked / injected) with a verdict, a
  production-ranking replica, deterministic summariser/embedder/narrator
  stubs and seven scenarios (default, past capacity, pinned, low top_k,
  long-block early/late, lineage control)
- tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of
  a finished real campaign
- tools/m11_long_run.py: opt-in --independent-fact mode with per-turn
  isolation tracking and the recovered_through_memory_independent verdict;
  M04 verdicts unchanged
- tests: diagnostic stages, eviction, creation window, ranking, lineage and
  authority controls; two strict xfails record the diagnosed retention and
  creation defects for WP-B.2 to flip
- planning/reports/v1.1/V1.1-WP-B1-REPORT.md

First failing stage: ranking (real model); retention past capacity and
creation for early facts in long blocks (deterministic, same on v1.0.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 20:50:05 -04:00
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00
JesseMarkowitzandClaude Opus 5 432f04100b M11 closeout: accept v1 release validation
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
The browser, offline and identity runs had last been taken on ef25b0a. The
closeout repeated them on the exact release-candidate tree, 3652dc6, whose
product code is identical to 96c1bf5, where the 100-turn evidence was run. No
product code changed, so the long-run evidence stands.

On 3652dc6:
- backend suite: 1421 passed, 17 skipped, 0 failed
- frontend suite: 161 of 161; lint clean
- production build clean
- docker build --no-cache: image SPA byte-identical to the local build
- browser regression: 38 of 38
- no-network container: 23 of 23
- identity diagnostic: 0 signals; the scripted self-test's detectors fire
- release-shaped smoke test from the image: 14 of 14

- M11 report: new S (exact-tree verification, including the identity
  results the report never carried) and T (acceptance record). Corrections:
  the REQUIRED FOR V1 count is 82, not 85, and L's browser narrator was
  qwen2.5:3b-instruct. P gains risks 16 and 17; risk 6 is widened.
- BUILD-MILESTONES.md: M11 COMPLETE / ACCEPTED, and a post-v1 backlog.
- V1-ACCEPTANCE-TESTS.md: the P release gate's result, and the P3
  disposition's run.
- planning/README.md, VERSION.md (v4.0), README.md: status, map, stop rule.
- tools/m11_browser.py: the G01 import wait could not fail, because the
  scenario's campaign is titled "Hidden Knowledge". It now waits for the
  imported source's row.

Found and carried, not fixed. The identity run stored protocol shapes the
extractor leaves, on 4 of 10 turns at a 4,096 window: event-call syntax and a
parroted length hint. The owner chose residual risk. The state rule's example
is fantasy, and the state lagged the narration. None occurs in the 100-turn
evidence.

No requirement changes. No release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CBTT3qbkGoemWD7BRvvxpT
2026-09-14 06:07:16 -04:00
JesseMarkowitzandClaude Opus 5 96c1bf5ded Measure M04 by where the planted turn is, and catch a section of the model's own
The M04 re-run on 0c7316f ran 101 turns with no failures, and the verdict
still came out `precondition_not_met`. That verdict was wrong. The planted
player turn (depth 1) was 65 actions outside the history window, whose floor
was 66. The sentinel's text was in recent history for two other reasons.

- On 5 turns the narrator wrote a section of its own, `## Established:` over
  indented facts, with the planted clue copied into it from the state section.
  The extractor passed it: it was a single heading, and the `## ` meant it did
  not match. The harness leak count passed it the same way, and read 1 where
  5 turns leaked.
- On 4 turns the narrator used the sentinel as a name inside ordinary
  sentences ("the SILVER-KEY-CRYPT-OLD-ABBEY, flickers with latent power"). That
  is story text and cannot be stripped.

So the sentinel's text in history can never be the precondition. M04's
written pass is "Fact/event remains recoverable without entire transcript in
prompt" (V1-ACCEPTANCE-TESTS.md). The owner agreed on 2026-09-13 that the
precondition is positional, and that recovery through authoritative state
counts; memory is not required.

Extractor:
- A state-section heading is recognised with any markdown the model wrapped
  it in (`## Established:`, `**Held:**`, `> Held:`).
- One heading with an indented entry under it now qualifies as protocol.
  Before, a block needed two headings, or one heading and the scene line. A
  heading followed by unindented prose is still story. The earlier guard test
  "Held:" with an indented line is now protocol, and the case was rewritten
  unindented.

Harness:
- It records `planted_depth` when the clue is planted, and carries it across
  --resume. No endpoint reports an action's depth, but a fresh campaign's path
  is the opening, the planted turn and its reply, so the depth is
  `total_actions - 2`. That was checked against the database in three runs.
- `_recall` reports `planted_depth`, `history_floor_depth` and
  `planted_turn_in_history_window`. The verdict is `precondition_not_met` only
  when the planted turn is still in the window, and `precondition_unknown`
  when its depth was never recorded. `clue_in_recent_history_window` stays as a
  fact about the prompt.
- The protocol-leak count matches a state heading with an indented entry under
  it, in any markdown.

Every AI turn in five real runs was replayed through the new extractor: draco
M01, the two 26-turn GPU trials, the f8d4010 M01 run and the M04 re-run. 443
turns in all. No turn the old extractor had left clean changed, and the M04
re-run lost 7 more leaks. The sentinel-as-a-name turns are story and remain.
Trial 1's model-invented headings remain, as before.

The M04 re-run's own evidence, reclassified under the new precondition from
its database and its recall-turn snapshot, reads
`recovered_through_state_only`. The original recall.json is kept unchanged
beside `recall-reclassified.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
2026-09-13 22:10:33 -04:00
JesseMarkowitzandClaude Opus 5 0c7316f951 Keep the state section and its proposal out of the story
The first complete M01 run with the memory bank on (f8d4010, 101 turns on
a GPU host) reported "complete". It still did not prove M04. The planted
clue was found at turn 100 only because the narrator had pasted the
narrative-state section into its own prose, and the paste was still in
recent history. No memory and no summary carried the clue.

The narrator is a small local model. It wrote protocol into its stored
narration on 42 of 104 turns, starting at depth 2, in four shapes:

- a copy of the state section: `Scene:`, `Who and what exists:`, `Held:`,
  `Established:`, `Still open:`
- that copy above a correct ```state block, which was stripped while the
  copy stayed
- the copy, then a bare `State` heading, then a `> {"events": ...}`
  proposal quoted like a player turn, sometimes with story after it
- the same block cut off by the output-token limit, on 10 turns

Stored text is replayed verbatim as history, so each leak also put a second,
older account of the state into the next prompt. That is what M5 review
Finding 4 removed from history replay, and every leak gave the model another
example to copy.

The extractor now removes:

- a pasted state section, recognised by at least two of the renderer's own
  headings as whole lines. The headings are constants in `render.py`, so the
  renderer and the extractor cannot drift apart. One heading alone, or a
  `Scene:` line of prose, is left.
- an unfenced proposal that starts a line, quoted or not, when it parses and
  is a proposal. With no fence it becomes the turn's proposal. A `State`
  heading directly above goes with it. Candidates are taken outermost first,
  so a finished event line inside an unfinished block is never taken as a
  proposal by itself.
- an unfinished unfenced proposal at the end that reads as protocol.
- whatever is left at the end: a `State` heading, a bare `>`, a parroted
  reminder or continue hint (closed or not), and a ```json fence cut off
  before it names its events. These are cut repeatedly until nothing more
  comes off.

This also fixes an older bug. `_STATE_FENCE_RE` read "a ```state block"
inside a parroted reminder as a fence opening and cut out the middle of the
reminder. The label must now end its line or run straight into the payload.

A reply whose only removal is a pasted state section records no raw block,
so the turn is not marked unparseable for a block it never started.

Every AI turn in four real runs was replayed through the new extractor:
draco M01, the two 26-turn GPU trials, and this M01 run. 339 turns in all.
No turn the old extractor had left clean changed. Every leak of our own
protocol is gone: 42 of 42 in this M01 run, 5 in trial 2, 3 on draco.
Trial 1 still has model-invented headings ("Identifiers established:",
"Set of events made true:") on 10 turns. They paraphrase the instruction and
are not our renderer's text, so they are left, not guessed at.

The long-run harness now records an explicit M04 verdict, which is never a
recovery while the clue is still in recent history. It also counts the AI
turns in the export that still carry protocol.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
2026-09-13 21:18:18 -04:00
JesseMarkowitzandClaude Opus 5 f8d401029f Stop a turn locking out its own memory bank, and let the long run notice
The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.

The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.

- `retrieve_memories` now only reads. `record_use` writes the counters in the
  turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.

The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.

The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".

Both new application tests fail on fec46f6: the lock probe sees
`database is locked`, and memory status stays `idle`. The full backend suite
passes (1392 passed, 17 skipped). A 26-turn re-run against the same host had
0 lock errors, wrote 7 memories and 2 summaries, and used them in the prompt
from turn 8.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0136VBTMUKWYeU6G9HgbDbND
2026-09-13 20:21:21 -04:00
JesseMarkowitzandClaude Opus 5 fec46f66bb Turn the memory bank on for the long run, and refuse one that cannot use it
The first complete hundred-turn campaign did not exercise M01's
"summary/memory activation" step. Memory bank and auto-summarize are
per-campaign switches that default to off, and m11_long_run never
turned them on: summary_tokens and memory_tokens were 0 on every turn,
memories_used was empty, and M04's clue was recalled through narrative
state alone. The retrieval path M6 built was never asked, and nothing
in the evidence said so except a row of zeros.

setup now PATCHes both switches on, reads the campaign back, and stops
before the first turn if either did not take. memories_in_bank is
recorded on every turn, in the final summary and in the recall, and the
recall also says whether a summary exists, so which of the two recall
paths succeeded is stated rather than implied.

The embedding model is now required. Without one the summary pass still
writes memories, but memorybank.retrieve answers "No embedding model
configured" and returns none -- the same unexercised path in a fuller
bank. The harness refuses before it starts a server or claims --out.

tests/test_m11_long_run_memory.py drives setup against the real
application in-process: the switches are on afterwards, a server that
ignores the PATCH is refused before any state is written, the bank
count comes from the application and reads -1 rather than raising when
it cannot, and a run with no embedding model is refused. The four that
exercise setup and the bank count were run against the previous
harness and fail there; the premise test (a fresh campaign has both
switches off) passes on both, as it should.

The 2026-09-10 run in ~/m11-evidence/m01 therefore does not count as
M01. It has to be run again on this harness.

Backend 1,382 passed, 18 skipped, 0 failed. The eighteenth skip is
test_built_spa_fetches_no_fonts_remotely, which wants a built
frontend/dist this worktree does not have; it is an environment
condition, not a change here. The frontend is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XKWHt2DXuvqP83cAk6Zq88
2026-09-12 21:56:14 -04:00
JesseMarkowitzandClaude Opus 5 ef25b0a876 Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
2026-09-10 06:13:55 -04:00
JesseMarkowitzandClaude Opus 5 144406cd48 M11: what the server will actually read
The release-validation milestone, and the thing it had to settle first was
whether any of the earlier evidence meant what it said. M8 measured a deployment
enforcing a 4,096-token input window while the application budgeted 16,384.
Every request returned 200. What Ollama does with the excess is drop the oldest
tokens, and the oldest tokens here are the system block — the narrator's rules
and the campaign canon. A hundred-turn certification against that server would
have looked perfect and proved nothing, which is why this milestone could not
begin with a hundred turns.

So the application asks now. Ollama's window is a property of how a model was
loaded rather than of the request — sending num_ctx is accepted, ignored, and
worse, reloads the model at the server's own default — so the only honest move
is to find out and then tell the truth about it. /api/ps reports what a resident
model is being served with, /api/show what an unloaded one will load with, both
on the same host inference already uses, through the same endpoint policy and
the same TLS trust store. A verified window is a ceiling on the budget; an
unverified one leaves the budget alone and is recorded as unverified in the
turn's own provenance, so an old turn can be asked afterwards whether it was
built against a checked window. There is no third behaviour, and in particular
no hard-coded 4,096: a number the server did not say would be right on one
machine and wrong on the next.

The proof that this is doing something is a campaign whose canon sits at the
front of the prompt, 120 turns of history, and a 4,096-token window. The canon
is still there afterwards and the oldest history is gone. The same campaign
built the old way produces a prompt more than twice the window — the defect,
reproduced, so the fix is measured against it rather than asserted.

Two defects the validation found on its own, and they are the same defect twice:
something was true and nobody was told. A manual state correction of four
changes with one bad reference applied three, returned 201, and said nothing —
while recording the refusal on the audit row nobody reads. It came to light
because the identity diagnostic's own fixture was refused that way and the whole
run proceeded on a campaign with no scene, which would have read as a model
failure. And the narration-length setting moved no number: brief, medium and
long each became one English sentence, while the numeric hint the model actually
reads was derived from the global reply cap and said the same thing for all
three. Both now say what they did.

The other two post-M8 findings are closed as well. The tab said AI D&D, which no
document had ever claimed it did not; it says Interactive Story now, with the
open campaign first, and the name is the owner's decision rather than a
find-and-replace to something narrower than the engine. After an Undo the reader
could not tell where they had landed; the control row now ends with
"Moment 11 · later story ahead", from the server's own answer, in the word the
transcript already uses, with none of head, branch or depth anywhere near it.

The identity diagnostic exists and the root cause does not. That campaign was
destroyed, so no cause can be established — what M11 owes the finding is
something that can classify the next occurrence, and a diagnostic that makes only
the judgements a program can honestly make: duplicate keys, shared names,
protagonist drift, state and context disagreeing. Whether prose misattributed a
line is left to a person reading it beside its prompt, because a regex cannot
read dialogue and one that pretended to would produce exactly the confident wrong
answer this finding is about. Its detectors are proved to fire against a planted
second Alice.

Two entities may still share a display name. That was checked first, as the
finding asked, and left permitted: a mother and a daughter, or a stranger giving
a false name, are ordinary fiction, and refusing them to guard against a model
mistake would refuse the wrong thing. What was missing was that it happened
silently. It is reported now.

Evidence, not inference: a hundred accepted turns against a real narrator with
genuine process restarts; a real browser against the built SPA; a container with
no network at all; a campaign moved into a data directory that never existed.
Each was discarded and re-run whenever the product changed under it, and the runs
that were thrown away are listed in the report with the reason, along with ten
defects in the harnesses themselves — because a harness that has only ever
agreed with itself is not evidence, and two of M8's five harness defects were
masking real ones.

No dependency was added, removed or upgraded. No acceptance test was retired,
relaxed or reclassified. M11 is implemented and verified; it is not accepted, and
there is no release tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 14:01:20 -04:00
JesseMarkowitzandClaude Opus 5 1013c94eb1 M10: the seam for media, and no media
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
The media extension contract asks for a scene snapshot a future image or video
provider could be handed: location, who is present, what they hold, what must
stay true, and where in the story it sits. Building one was the milestone's
obvious first task, and it was the wrong one. That snapshot has existed since
M5. `narrative_state["scene"]` holds the summary, the location, the cast and the
coordinate it was written at; a validated `set_scene` event writes it, every
position snapshots it, and every head move restores it. It survives Undo, Redo,
Retry, divergence, Save Point restore and a process restart because it is the
authoritative state rather than a copy of it.

So there is no scenes table here. A second scene store would have been a second
answer to "where is the story now", with its own lineage rules to get wrong —
and the lineage rules are the expensive part, which is the argument for reusing
the ones that already work rather than against it. The Scene Packet is derived
on read, and its identity is computed from the campaign and the position rather
than allocated: the same position yields the same id in another process, after a
restart, and after the packet is thrown away and rebuilt, with no row to keep in
step. That is the part of a future media_assets table that would be expensive to
retrofit, so it is fixed now even though the table is not built.

One table, then: visual_profiles, the only thing the contract's scene list asks
for that nothing already stored. Campaign-scoped and not per-position, because a
character does not change appearance when the story forks — a reader who
diverged would otherwise lose their cast, and the same descriptors would land in
every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn
campaign to say something that never varies. Keyed by the M5 entity key rather
than a new identity namespace, and one table for characters, locations and items
alike, because a location is an entity with a type and splitting them would
reintroduce the genre shape M5 spent a milestone removing.

What the packet leaves out is the more interesting half. Not the transcript, and
not imported knowledge — none of it, not merely the sources marked hidden. The
rule is what the story established at this position, not everything the narrator
was told, and drawing it by class is what makes it hold for a secret nobody
thought to mark. A hidden Canon source proves it, with a positive control
showing the narrator did receive the sentinel the packet does not carry. Once a
validated event puts the observer in the room, the observer is in the packet:
that is no longer narrator-only knowledge, and a packet that hid it would be
hiding the story from itself.

The providers are contracts and nothing else. Protocols for image, video, audio,
speech and transcription, an empty registry, no adapter, no dependency, no
socket, and no media setting to point anywhere — a setting that exists can be
pointed at a cloud by mistake. A future provider endpoint must be loopback,
stricter than narration's trusted-LAN allowance, because a picture of a scene
carries the scene with it. Transcription returns an editable draft with no
commit method, so STT structurally cannot bypass the authoritative path.

Nothing here can write the story. Not by convention: no module under media/
imports the code that writes state, no media event type exists in the state
vocabulary, and every test in the authority suite compares the authoritative
document byte for byte either side of a media operation — including one where a
provider insists Alice is in a red coat in a corridor, and the campaign goes on
disagreeing.

One defect, found by the milestone's own tests. M10 first added a migration
creating an index that create_all already builds from the column, so an upgraded
database ended up with two indexes and a fresh install with one. Comparing the
two schemas is what caught it; neither database examined alone would have. The
migration is gone rather than renamed, and the right number of migrations for a
new table whose indexes are declared on its columns is zero.

Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145
passed. Lint, production build and Docker build clean. No frontend file changed:
M10 adds no reader-facing surface, and ordinary play — turns, state, memory,
knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media
configuration, no warning, no connection attempt and no media row written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 03:41:04 -04:00
JesseMarkowitzandClaude Opus 5 44edece67e M9: a campaign you can actually get back
A campaign could already be exported and imported. What could not survive the
trip was everything that explains it: the state events behind the authoritative
document, the prompt each turn was actually given, the passages it was shown,
the summaries that carry long-story continuity, and which take belonged to which
turn. An imported campaign could be read and could no longer say why it was what
it was — and a manual correction, the one state change no narration explains,
was indistinguishable from something the story had established.

The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather
than a side effect. Everything added here could have been another optional key,
the way persona, Save Points, narrative state and imported knowledge each were.
That mechanism stops working at exactly this addition: a v2 file with no prompt
provenance is ambiguous between "written before M9" and "written by M9 from a
campaign that has none", and those are different facts about a campaign. A
version number is how a recovery file states what it was capable of recording.
v1 and v2 still import, and every seam from pre-active-head onward is tested for
the rule that an older file is never reinterpreted under a newer assumption.

Two categories became three. "Chosen travels, derived is recomputed" was enough
until stored prompts had to be decided: they are derived, and they must travel
anyway. The test that separates evidence from cache is not "could this be
recomputed" but "would a recomputation answer the same question" — a rebuilt
search index answers the same question, a rebuilt prompt says what the turn
would be told *now*, which is the opposite of what the inspector is for.

Also here: a real SQLite backup, through the online backup API rather than a
file copy, taken while the application is running and verified before it is
kept; story cards settled as compatibility-only legacy data and taken out of the
narrator's prompt, because they were the untracked path around knowledge
authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema
change at all, proved against a database M8's own code wrote.

Three defects, found by running the milestone's own tests rather than by reading
them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the
freed ids to the next source imported into any campaign, which failed with an
integrity error that Reindex could not repair — both ends are closed, and a
database already carrying the damage now repairs itself. An imported node with
no state snapshot was being stamped with the campaign's head state, so an Undo
to turn 2 showed what the story knew at turn 20. And the snapshot relink did not
persist at all, because it mutated a dict in place on a column SQLAlchemy tracks
by assignment: it looked correct in memory and wrote the wrong ids to disk.

Carrying per-turn prompts looked like it would halve the length of campaign that
can be restored. Measured — and after compressing them inside the file —
everything M9 added costs 12% of it: the import ceiling moves from about 318
turns to about 279, against a 100-turn certification target. The dominant cost
is not M9's at all. The per-position narrative state document is 74% of a
bundle, and v2 already carried it.

Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint,
production build and Docker build clean. Verified across two server processes
with two data directories, and in a real browser against a real narrator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 01:55:45 -04:00
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00
JesseMarkowitz 8c65ae99de M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The
milestone is subtraction, and what is left is the single-user local
storyteller the specification describes.

Removed in full: campaign scripting and its QuickJS sandbox; multi-user
accounts, guest sessions, login, registration and the shared demo key;
the visitor-analytics tables, dashboard and page beacon; the access log
of sign-ins, addresses and devices; per-IP and per-user rate limiting
and quotas; Render deployment config; Postgres and psycopg; cloud
inference providers, the API-key field and the key encryption that
existed to store it; session-cookie signing. None of it was hidden
behind a flag — the routes are gone and answer 404.

Two things were kept that the brief allowed keeping. The `users` table
and its foreign keys stay as an internal ownership detail, because
rewriting them out means a migration across most of the schema to
delete a column that costs nothing; nothing creates a second user and
no request carries an identity. Five inert tables and four inert
columns stay for the same reason, so an M1 campaign database opens
unchanged.

The one addition is app/endpoints.py, which decides where a story may
be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an
explicit allowlist of networks, not a guess at what `ipaddress` means
by "private", which calls the documentation ranges private and IPv6
loopback reserved. Every address a hostname resolves to must be in it,
so a split answer does not squeak through, and the rule runs both when
the endpoint is saved and before every outbound request, because a name
that resolved to the LAN this morning can resolve elsewhere this
afternoon. Known cloud hosts are named in the refusal so the error says
why rather than looking like broken DNS. TLS is never traded against
it: M1's shared trust context is intact on all four clients and there
is no way to skip verification.

The hardcoded 120-second model timeout is now a setting. That was not
theoretical — on this GPU-less four-core host a cold load of
qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while
turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short
at 10s so a wrong address still fails fast; the read timeout defaults
to 300s and is bounded at 3600, because "wait longer" must stay a
number.

Two defects found while testing and fixed here. An unknown /api path
fell through the SPA catch-all and came back as HTML with status 200,
so a client asking for JSON parsed a web page instead of learning the
route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an
unauthenticated loopback API would hand every page on the Internet a
write handle on the campaign database; it now refuses to start.

Verified rather than assumed. Offline, on a network with no route out
and no DNS: five turns, retry with both takes retained, restart with an
identical transcript digest, a failed model call leaving the accepted
AI-turn count untouched, and a capture with zero non-loopback unicast
packets. Against a real second machine on the LAN over HTTPS with a
private CA: four turns, restart, and a capture showing 289 packets to
the approved host, 344 loopback, zero anywhere else, zero DNS queries.
Cloud and public endpoints refused with their reasons; no API key
settable; every removed route 404.

604 backend tests pass, down from 648 by the fifteen retired with the
subsystems they tested and up by the twenty-nine added for the endpoint
policy and the removed surface. The scripting tests were not deleted:
eight files used a JavaScript counter as instrumentation for the state
snapshot and rollback machinery, which M2 does not touch, so the
counter moved to the world-state engine and those tests still assert
what they always did. Frontend lint and build are clean; the image
builds, and its wheel-building stage is gone with quickjs.

No M3 work. Undo is still destructive and there is still no Redo.
2026-09-02 11:27:14 -04:00
Claude 745a4ea9f3 Stop printing the database password
The report opened with the connection string it was about to work on, taken
straight from AIDND_DATABASE_URL. On the hosted deploy that string carries the
Neon password, so the first line of every run put a live credential into the
console — and from there into scrollback, a screenshot, or a pasted bug report.
Nothing in the output said it was there to notice.

It now prints scheme, user, host and database name. The query string goes whole:
sslmode is the only part worth reading, and some drivers accept a password
there as well.

631 green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
2026-08-31 15:13:47 +00:00
Claude eeab8ef5bb Survive the console this will actually be run from
The report prints an arrow between the old and new word counts, and it prints
the memories themselves, which are model-written prose. A Windows console
defaults to cp1252 and cannot encode either, so the run dies on
UnicodeEncodeError partway through — after some memories have been rewritten
and committed, which is the worst place to stop. STATUS already carries the
same warning for tools/dbmeter.py.

stdout is reconfigured to UTF-8 with errors="replace" instead, so the report
degrades a character at a time rather than failing.

The usage notes were also written for a POSIX shell, and this project is
developed in PowerShell, where `VAR=x command` is not a thing. Both docs now
give the PowerShell form first, via Read-Host so the Neon URL and the secret
stay out of the history file, and name the venv's Python.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
2026-08-31 14:44:16 +00:00
Claude 4c94dd0fc0 Keep a production run to the adventures you meant
The hosted database is not the local one. It holds other people's stories, and
`summary_provider` builds from the adventure owner's Settings, so an unfiltered
`--write` against it would spend other people's money rewriting memories they
never asked about. `--adventure` could already hold a run down, but only if you
knew the ids.

`--email` names accounts instead, and every line of the report now says who owns
the adventure, so a dry run answers "whose keys would this spend" before
anything is written. Guests have no email and stay reachable only by id, which
is the right amount of friction for touching a stranger's bank.

The other half is written down rather than built: stored API keys are encrypted
with AIDND_SECRET_KEY, so a run from a checkout against the Neon database needs
the same value the web service holds. With a different one `decrypt_secret`
returns "" instead of failing, and every adventure is reported as having no key
— a run that looks like it worked and did nothing. The Dockerfile copies
backend/app alone, so tools/ is not on the box either way; the recipe is a
checkout pointed at AIDND_DATABASE_URL.

629 green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
2026-08-31 10:18:41 +00:00
Claude 0633cb624e Run the new memory prompt back over an old bank
The prompt change only reaches memories written after it. plan/18 decided to
leave the existing ones alone and let eviction age them out at
memory_bank_capacity, on the grounds that re-summarizing would duplicate
whatever was still in the bank because nothing deletes the old rows.

That was wrong about the only option. A memory can be rewritten in place. The
row carries more than its text — whether it is pinned, how often it has been
retrieved, and the node it hangs off, which is what makes a fork inherit the
right memories — and rewriting `text` keeps all of it. Deleting the bank and
rewinding the cursor would lose that, and would trickle memories back at
MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never
recover.

tools/rewrite_memories.py does it. Without --write it makes no model calls and
only reports the scope; --write rewrites, --embed re-embeds in the run rather
than leaving it to the app's post-turn pass. It reads whichever database the
app reads, so it works against the hosted Postgres as well as a local file.

Two things it needed from the app. `summarize_block` is now the one place a
memory prompt is assembled, and the post-turn pass calls it too — a backfill
that built its own prompt would be writing memories with a prompt that never
shipped, and nothing would report the drift. `source_block` reads a memory's
block back out of the story, which nothing has ever had to do: it reads on the
lineage of the branch the memory was written on, not the branch being played,
because after a fork the same depths hold different actions on each side and a
read through the adventure's path would summarize the wrong story silently.
It also excludes the discarded attempts at a retried turn, and tolerates a
block an action has since been deleted from.

Left alone: a memory with no source range, which is hand-written or migrated by
62 and may be the player's own words; a memory whose actions are gone; and an
adventure whose owner has no API key, because summarization spends the user's
own key by construction and never the shared demo key. --api-key/--model/
--endpoint override that, the last of them aiming a run at claude_shim.py.

The vector is cleared for every rewrite, because the stored one describes
wording that no longer exists. Re-embedding always uses the owner's own
embedding model, never --endpoint: a vector only means anything against the
vectors it is ranked beside.

17 tests, 627 green. The fork case is the one that would fail quietly, so the
test builds a fork whose depths hold different actions on each side.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW
2026-08-31 10:11:26 +00:00
Claude 712ef44f57 Keep the A/B, and the harness that produced it
The run that justified MEMORY_MAX_WORDS lived in a scratch directory and
would have been gone with the container. The numbers in plan/18 were
therefore assertions nobody could check.

plan/18-appendix-memory-ab-run.md now carries the whole transcript: both
memories, both summaries, and the thirteen-action story they were
written from.

backend/tools/memory_ab.py reproduces it. It replaces the throwaway
script the first run used, and differs in two ways that matter. It goes
through OpenAICompatibleProvider rather than calling a model directly,
so a run exercises the provider, the streaming path and complete()
instead of a stub. And it reads the control prompt out of git at the
commit given to --before, so the thing being compared against cannot
drift from what actually shipped.

There was already a claude_shim.py serving an OpenAI-compatible endpoint
backed by the CLI, which is exactly what the throwaway script had
reinvented. memory_ab.py points at it by default, so a run spends a
Claude subscription rather than API credit, and --endpoint aims it at
the provider the deployed app really uses — which is the one question
this whole exercise could not answer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
2026-08-31 15:03:20 +05:30
parththakkar106andClaude Opus 5 6a5d8e98f9 Give the stress fixture back its sibling attempts
`place_action` used to read `depth` from `index`, so the fixture's retry
attempts all landed on the turn's own depth and formed one sibling group.
Now that depth comes from the head, each attempt moved the head one step
and every retry became a turn of its own. The fixture's own check caught
it: a coordinate holding a single superseded attempt has no live row, and
that turn disappears from the story.

The attempts at one turn share that turn's coordinate, so only the first
one goes through `place_action` and the rest copy its placement. This is
what `attempts.add_attempt` does. The fixture cannot call it directly,
because it makes the newest attempt live and the fixture needs a live
attempt that is often not the newest.

Also drop the `variant_index` and `variant_count` arguments, which raise
`TypeError` now, and read the two mark positions from a helper instead of
from the adventure columns migrations 72 and 73 removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YFQY6WgaE3JaX3dXkxLynV
2026-08-30 18:43:50 +05:30
parththakkar106andClaude Opus 5 f1bebe18d0 Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`,
`variant_count`, `state_before`, and `world_state_before`, plus
`adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in
SQLite, so migration 71 quotes it.

Nothing outside the migrations read these. `models.py`, `schemas.py`, and
`ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders
by `id`, and `attempts.renumber`, `context.history.max_action_index`, and
`nodes.next_index` are deleted.

Two changes keep the migration replayable on a `create_all` database:

- `_split_variants_into_siblings` wrote through the live ORM table, so it
  stopped compiling once migration 66 removed five of its columns. It now
  writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`.
- Five data passes read columns these migrations drop. Each now calls
  `_has_columns` and returns early when the columns are absent.

`bootstrap` takes a `through` version so a migration test can stop at the
schema it asserts on.

555 tests pass, up from 549. Eight of the new cases assert each column is
gone after a real schema-45 database migrates all the way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK
2026-08-29 17:15:49 +05:30
parththakkar106andClaude Opus 5 0623dd8b78 Split Play.jsx into a directory of twelve files
`frontend/src/pages/Play.jsx` was 2280 lines holding 27 components. It is now
`frontend/src/pages/Play/`: `index.jsx` with the page component, five panels,
two drawers, the reports, the take pager, the refresh dialog, and the shared
formatting helpers.

Every line moved verbatim. A coverage check confirms every non-blank line of
the original appears exactly once, in order, across the twelve files, and a
name-resolution check confirms every identifier each file references is defined
or imported there, with no unused imports.

The line split stranded a comment at six of the boundaries. A leading comment
sits above the section it describes, so each cut left one at the end of the file
before it. All six moved to the section they describe, rewritten in the house
style.

`usePlaySession.js` is not here. The page component still owns all of the
session state. Moving eighteen `useState` calls and seven `useEffect` calls is a
rewrite rather than a move, and no frontend test would catch a mistake in it
today, so it waits for Stage 5.

Four stand-in providers under `backend/tools/` now define `last_usage = None`.
The turn engine reads that attribute after every call, and the fixtures never
defined it, so `tools.tree_fixture` crashed. That break predates this branch.

Verified by 549 passing tests, a clean `npm run lint` and `npm run build`, and
by driving the Play screen: the story view, all five panels, both drawers
including the world-state edit form, the branch map, the refresh dialog, and the
take pager stepping onto a take that lives on another branch. No console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 01:33:14 +05:30
parththakkar106andClaude Opus 5 2fa812c056 Split the adventures router into a package
`backend/app/routers/adventures.py` held 2353 lines and 35 endpoints. It is now
a package of 14 modules, the largest 443 lines.

The split moves text rather than rewriting it. An AST comparison against the
old file confirms all 86 definitions are identical, and the OpenAPI schema
still lists the same 35 operations.

Names a test replaces now live in `turns.py` only, and other modules reach them
as `turns.<name>`. Rebinding a re-exported alias changes the alias and leaves
every caller reading the original, so the package root does not re-export them.
A patch aimed at the old target raises `AttributeError` instead of passing while
doing nothing. Tests and the fixtures in `backend/tools/` say
`adventures.turns.<name>`.

The same rule keeps the turn lock working. One module owns `_active_turns`, so
one lock guards one set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 01:05:50 +05:30
parththakkar106andClaude Opus 5 ef4da7fda2 Keep the local Claude test rig, and record what it found
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by
the local `claude` command line tool, so a demo can be played against a real
model with no API key. Each request spawns one `claude --print` process, which
suits the turn engine: the app assembles the whole prompt every turn and
expects a stateless endpoint.

Playing the Pokemon demo through it made all five of the world-state changes in
`plan/16` work, and the refusal loop ran end to end for the first time: turn 3
clamped to nothing, turn 4's assembled prompt carried the correction verbatim,
and the model's next delta was right. Two failures previously blamed on the
code were the demo model. One is a schema fault and is still open: Milo's three
Pokemon share one `active_hp` stat, so a switch leaves the newcomer at 0 HP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 18:42:24 +05:30
parththakkar106andClaude Opus 5 ae39514d1e Name the Pokemon demo, and land a seed rename on its own row
`seed.py` matches a seed file to its scenario by title, so renaming one
inserted a second public scenario and stranded the first. The stranded row
stays public forever and has to be deleted by hand on every deployment, which
is what happened to "Road to the Champion". A seed file now lists its old
titles under `previous_titles`, and the rename updates the existing row.

The cover art is a PNG data URI. `app/images.py` accepts raster formats only,
because SVG can carry script and the bytes are served from the app's own
origin, so an SVG stores but yields an empty `image_url`.
`tools/make_pokeball.py` draws the ball with `zlib` alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 18:42:06 +05:30
Parth e7d75c3b05 Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style

Rewrite the comments and docstrings across the backend core modules so they
read plainly. The previous prose was accurate but dense and figurative, which
made it slow to skim.

Applies the Google developer documentation style guide: short sentences, active
voice, present tense, American spelling, and no metaphors, idioms, or
rhetorical asides. Replaces em-dash chains with separate sentences.
2026-08-26 15:37:25 +05:30
parththakkar106andClaude Opus 5 40d2555f84 Say what the app is now, everywhere it is published
The README, the project page and the engineering guide all describe a linear
story. The tree shipped two days ago. Every published surface is a phase
behind, and the guide is not merely behind — it is wrong in a way that costs a
reader time.

Its 2.2 was "Two coordinate systems, and the bug class they create", and it
explained the codebase through position_of_index, note_action_removed and
settled_story_actions. All three were deleted in SP3. 2.3 explained retry
through Action.variants and state_before. Somebody reading either would go
looking for machinery that is not there, which is worse than a gap.

So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are
written at: the seven bugs that turned out to be one bug, the lineage clause
and the two properties that make fork count free, why takes group by parent_id
rather than by coordinate, cursors becoming anchors, and a closing list of what
the design is honest about. 2.3 is rewritten around state_after and takes, and
1.1 and 1.5 follow, because the pipeline no longer snapshots before the call
and the memory bank no longer holds an action back.

The numbers were simply old: 151 tests where there are 440, 37 migrations where
there are 64, twelve phases where there are fourteen. They appear in four
places across the README, the project page's stat tiles and the guide's results
table. The measured branch cost — 103 B, and 1.007x the page load of the same
story flat — is added beside the egress and turn-cost figures it belongs with,
since it is the number that answers "what does branching cost me".

Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven
through eight written turns with written deltas, three discarded takes forked
onto branches of their own, one off a branch so the map has to nest. Same
reason tree_fixture.py is committed — the shots have to be reproducible and the
frontend still has no test runner. play-world-state.jpg is reshot because it
predates the entire tree UI; the map and the branches panel are new.

Note for next time: docs/guide.html is hand-written, not generated from the
Markdown, so every guide edit is two edits in two vocabularies. Both files were
checked for tag balance and both pages rendered locally before this landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-20 04:19:53 +05:30
parththakkar106andClaude Opus 5 c8e081e9d6 Draw the tree the branch list can only spell out
The Branches panel says which lines exist. It cannot say where they parted
or how much story each one is, because those are the two numbers `GET
/branches` already answers and a list has nowhere to put them. A ⌗ See the
tree button opens a map: one horizontal lane per branch, running from the
moment it left its parent to the moment it ends, joined to the parent by an
elbow at the fork. The horizontal axis is the story's own clock, so two
lanes at the same x are at the same moment and a short branch reads as
short.

This is a branch map, not the per-node map SP7 refused, and that is the
whole reason it was cheap. Lanes are bounded by branch count, not node
count, so the 600-node windowing problem never arrives. It reads the single
request the rail already made and nothing else.

`branches.js` holds the tree maths and `BranchMap.jsx` the drawing. Play.jsx
gives up its private copies of branchLabel and orderBranches: the panel and
the map now label and order a branch through the same functions, so a branch
cannot be called two things by the two views. The three operations stay in
BranchPanel and are passed down, and `run` answers whether it worked so
neither view clears a half-typed name on a refusal.

Three things came out of driving it, none of which a test could have seen.
`clientWidth` counts the canvas padding the ResizeObserver leaves out, so
the first paint drew an svg 24px wider than its box — and the observer's
initial observation never arrived here, so dropping the seed left the map
never drawing at all. Both are needed and the seed subtracts the padding.
Only the name was being clipped, not the meta line under it, so a late fork
ran its text off the right edge; both are clipped now, and a lane starting
in the right third hangs its labels back over the fork, where its own band
guarantees nothing to collide with. And the delete rule lived in the server
and in the map but not in the list, which offered Delete on a branch the
head was forked from and answered with a toast from the server's 400.
`headLineage` is the client's copy of that rule and both views use it; the
server stays the authority.

`tools/tree_fixture.py` is the counterpart to `tools/branch_fixture.py` —
four branches at three fork depths, one forked off a fork. With no frontend
test runner it is the whole of the map's coverage, and it exists to be
looked at. Nothing in the backend changed; the 440 tests pass unmoved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-20 03:38:17 +05:30
parththakkar106andClaude Opus 5 811368d048 Refresh what a switch changes, not what a turn changes
A branch switch does not change the length of the story. It changes which
story it is. Four panels keyed on actions.length and so could not tell the
difference: the Branches panel drew one branch while the reader was already
on a second, Insights showed the prompt built for the path just left, the
script-state drawer kept the other line's numbers, and the Memory Bank did
not notice a deleted branch taking its memories with it.

Only the world-state drawer was right, and only because it happened to
carry stateKey already. They all key on the pair now, and deleting a branch
bumps it too — that is the one operation that changes what is stored
without a turn being played and without the story on the current path
moving by a single action.

tools/branch_fixture.py is the thing that could show it. The stress
fixture's world state is empty, so it cannot answer whether a switch puts
the scoreboard back, and its story is one branch. This builds a small
bootable adventure with a stat schema, a gold script, two takes on one turn
that differ by 35 hit points, a fork, and a memory on each side — with both
branches the same length on purpose, because equal length is precisely the
case a length-based key cannot see.

Verified in a browser with the drawer open: hp 60 to 95 and back, the bar
redrawn, the story swapped to the other take, Insights carrying the scratch
and not the beating. The Memory Bank deliberately does not change on a
switch: the drawer is adventure-wide so a memory is always findable to
delete, and retrieval is the path-scoped half.

396 tests, build clean, no new lint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 a7bf47a35e Let a backup carry a story that went two ways
A bundle had one list and a forked adventure has two stories, so export was
emitting every branch's turns interleaved by index — a mangled story rather
than lost data, and unreachable only because forking has no UI yet.
`ai-dnd-adventure-v2` carries the branches, the depth each one left its
parent at, which attempt at every turn is the story, and what each node
left behind. That last one is not decoration: the after-snapshots are what
a branch switch puts back, and a bundle without them imports a tree nobody
can switch inside.

`app/bundle.py` owns both formats and nothing else knows either. The v1
reader stays — those files are already on people's disks — and it is now
the only place a `variants` array exists anywhere.

The rule the module is built on is that a bundle carries what was chosen
and never what is derived. The head branch, the fork points, the live flags
and the anchors are decisions somebody made. The lineage, the head depth,
the legacy `index` and the variant ordinals are computed from those and are
rebuilt on the way in, because a bundle is a text file anybody can edit and
a derived field shipped beside its source is a chance for the file to
disagree with itself where no read would report it.

`index` is the one that stops being academic here. It agreed with `depth`
until SP5, and this is the first writer that has to fill it for a forked
story, where two branches both hold a node at depth 4. It is allocated one
per turn instead: siblings share it, no two coordinates do.

Everything a hand-edited file can get wrong about the shape of a tree is a
400 raised before the adventure row exists, because a half-applied import
is exactly the failure this phase exists to end — a story that goes quiet.
A file wrong about which attempt is live is corrected rather than refused;
that is an invariant of the database, not of the format.

Measured on the 600-action fixture: 587 kB to 911 kB, and all of the
increase is the outcomes at 489 B a node — the coordinates themselves save
57.5 B a node against the old turn-and-variants shape. Twenty forks add
660 B. 4.3% of the import body cap.

381 tests green, 16 of them new in test_bundle_v2.py. No migration, no
vacuum owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 0a12d9cd47 Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth),
with `live` naming the one the story tells. The JSON repeating group on
`actions.variants` is read one last time, by a migration that writes it
out as the sibling rows it always described, and then goes unread.

The snapshots turn around with it: an action carries the state it left
behind rather than the state it started from, because attempts at one
turn share a starting position and differ exactly in their outcome.
Rolling back is "what the node in front left behind", one lookup on the
path, and it is what undo and retry now both read.

And the memory holdback goes. It existed because retry rewrote a row
under a mark that had already moved past it; a retry writes a sibling
now, and replacing what a coordinate says withdraws what was derived
from it — the same repair undo and delete already made.

The assembled prompt is still stored once per turn: it moves with the
live flag, so a superseded attempt keeps only the few hundred bytes that
were its own. Measured on the 600-action fixture: 700 rows for the same
600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB
and page load 62.7 kB unmoved.

347 tests green. `tests/test_story_tree_baseline.py` and
`tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move
the baseline for the variant-count semantics and did not need to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 c51531709d Mark the story with a node, not with a count
The memory bank and the story summary each kept a cursor: how many story
actions they had already covered. A count is a position in a list, and this
list moves — delete an action in front of the mark and every later one slides
down a slot, so the mark now covers one it has never read. All the cursor
bookkeeping existed to patch that up.

Both marks are now (branch_id, depth): the node up to and including which the
work is done. A depth is a coordinate along a path, not an offset into a list,
so nothing in front of it can move it. That deletes rather than rewrites
`position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`,
`prune_dangling_memories` and the every-pass clamp in `run_post_turn`.

A memory hangs off the node its block ends on, so a fork inherits its
ancestors' memories without copying any, and retrieval selects through the
branch clause over the *whole* lineage — recall is long-range by definition and
cannot be windowed. Measured: 1,807 B on a story forked twenty times against
1,823 B on a flat one of the same length.

Migrations 53-56 translate the old counts into nodes. They rewrite `adventures`
and not `actions`, so this one needs no VACUUM FULL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 d3756abdaa Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a
`branches` table, `branch_id`/`depth` on actions and memories, a head pointer
on adventures, migrations 46-52, and a server-side backfill that re-reads every
existing adventure as a tree with one branch. `depth` holds the number `index`
already held, gaps included, so no story changes — a linear story *is* a tree
with one branch, which is what makes the SP0 baseline passing unmodified the
pass condition rather than a hope.

The writer had to come with it. No migration will ever visit a row written
after it ran, so columns backfilled today and populated next subphase would
leave a hole exactly the width of one deploy, and from SP2 on a row without a
branch is a row no read can see. `app/tree.py` owns that: one module, because a
node written without a branch fails by disappearing rather than by raising.

Three things the schema itself insisted on:

- `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing
  both ways makes the two tables a cycle create_all cannot order, and its
  escape hatch needs an ALTER SQLite does not have. It is a cache, and a head
  naming a branch that is gone recovers onto the root.
- `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a
  second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`,
  because Postgres `json` has no equality operator.
- SQLite will not drop a column a foreign key names, which is how two existing
  tests broke: they simulated an old database by rewinding the stamp while
  leaving the new columns in place. Every ADD COLUMN migration is now
  idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by
  rebuilding three tables from frozen DDL so the real ALTERs run.

297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn;
page load and index are byte-identical to the recorded figures.

The deploy that ships this needs one `VACUUM FULL actions;` on the direct
endpoint afterwards — it rewrites every row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
parththakkar106andClaude Opus 5 5c1bcf7305 Build a fixture that can witness what a tree changes
The stress fixture is sized from production and exists to weigh bytes, so it
leaves every column it does not weigh at its default. Checked against a freshly
built one, that is exactly the set phase 14 has to migrate: state_before and
world_state_before NULL on all 600 rows, no scenario and so no RPG layer, no
adventure scripts, both cursors 0, and a hundred retry histories whose two
attempts carry byte-identical text with variant_index pinned at 0.

That last one matters most. "Which attempt is live?" is the question SP4's
migration answers when it decides which sibling node becomes the head of a
turn, and against that fixture the question had no observable answer -- a
migration that picked wrong would have looked correct.

--rich fills in those columns and no others. A real RPG scenario read from the
seed data rather than invented, with a world state played forward so hp
declines and flags flip; monotonic per-action state snapshots, so a rollback
that does not happen reads as a wrong number instead of as nothing; a gold
script, story cards, non-zero cursors, pinned and forgotten memories; retry
attempts with distinct texts, counts of two and three, and a live attempt that
is often not the last written. Plus a second adventure, because a branch clause
that forgot its adventure still looks right on a database holding one.

The invariant SP4 reads -- text mirrors variants[variant_index] -- is asserted
at build time rather than assumed.

The plain fixture is untouched and re-measured unchanged at 1.8 kB, so the
egress ceilings stay comparable. All six shapes run clean on --rich, including
the turn path with the RPG layer and script pipeline now live. 283 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-18 19:14:07 +05:30
ParthandClaude Opus 5 0f1e05c808 Keep the fixture, so there is something to scroll (#5)
The harness already built a production-shaped 600-action adventure and then
threw it away with the temp file. The one open gap in plan/13 is that nothing
has ever driven the scroll in a browser, and part of why is that there was
never a long adventure to drive it with.

--keep PATH writes the fixture somewhere durable and makes the app able to
serve it. Two edits are needed for that, both of which cost an hour to
rediscover:

  - create_all() builds the current schema but leaves the version stamp at its
    default, and bootstrap() reads a populated-but-unstamped database as
    ancient — it replays every migration against a schema that already has the
    columns, and fails on the first.
  - the fixture's user is a registered one, but local mode looks for the row
    with email IS NULL and is_guest false, so without clearing the email the
    app opens on an empty library.

--keep is read before argparse exists, because where the database lives has to
be settled before app.database is imported. That is the same constraint the
AIDND_STRESS_DATABASE_URL block already lives under. SQLite only; combining it
with a Postgres target is rejected rather than half-honoured.

Verified end to end: the fixture boots with no manual step, action_count 600,
a 60-action first payload, and before_id walks back nine more pages to the
start. Nothing about the default path changed; 259 tests pass.


Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 16:50:41 +05:30
parththakkar106andClaude Opus 5 ae6e5af6c7 Store context_snapshot compressed
One column is 89% of the database and the free tier allows 512 MB. Reads were
already solved -- the column is deferred, so a page load never touches it and
one screen fetches one row at a time -- but nothing had costed storage, and
storage is the constraint with a cliff: 99.6 MB used, ~94 kB of disk per
action, so the ceiling arrives around 5,400 actions and 944 are stored.

Postgres already compresses it and only gets 1.7x. pglz is tuned for fast
decompression of data a query might filter on, and nothing has ever filtered
on an assembled prompt -- it is written once and read whole, rarely, by the
Insights viewer. zlib gets 3.5x on the same text for a decompress on a request
that already made an LLM call.

Done as a TypeDecorator rather than a second column, so every call site still
writes a dict and reads a dict back, and deferred/undefer/load_only keep
naming the same attribute. Only the storage format moves.

Migrations 43-45: add the bytea, convert into it, drop the original, rename.
The backfill is the one destructive step in the file -- 44 removes the only
other copy -- so it decompresses every row and compares it against what went
in, and a row that fails aborts the run. The whole loop is one transaction, so
an abort rolls the DROP back and the prompts are still there.

Verified on real Postgres, replaying 43-45 from a pre-43 schema on a throwaway
Neon database: 720,864 B of JSON became 204,293 B of bytea, 3.53x, the column
came out named context_snapshot, every snapshot compared equal and the one
NULL stayed NULL.

Postgres does not return the disk by itself: DROP COLUMN only marks the column
gone and the backfill leaves a dead tuple per row, so the table peaks near
twice its size before settling. The deploy needs one VACUUM FULL to collect
it; the migration comment says so.

The egress fixture's snapshots are prose now rather than "x" * 20_000, and the
prose generator moved to tools/fakeprose.py so the harness and the tests share
one definition. A repeated character compresses a thousandfold: against the
old fixture a compressed column looked free and the byte ceilings would have
been guarding nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 14:08:23 +05:30
parththakkar106andClaude Opus 5 12d57afdac Put a number on what an endpoint may fetch, not just a column list
test_egress.py asserted which columns a statement names, which is the shape
both of this project's egress blowouts took. It would all still pass if a
response grew tenfold within the columns it is allowed to read -- and a story
that keeps getting longer does exactly that. Production's longest adventure is
607 actions where the plan assumed 200.

So dbmeter, which was built to be importable from tests and was not yet used
by any, now backs four byte ceilings: the page load, the action list, and one
action's snapshot fetched on demand. Budgets are per action rather than
absolute, so they mean the same thing whatever size the fixture is set to, and
generous -- 3 kB against a real 994 B. They are there to catch an order of
magnitude, not to freeze a byte count.

The fourth test is the one that keeps the other three honest. A ceiling proves
nothing unless the thing it excludes would breach it, so it undefers the
snapshot on purpose and asserts the same twelve rows cost more than ten times
the budget. If the fixture ever shrinks below the point where that holds, that
test fails rather than the ceilings quietly passing on nothing.

Meter grows detach() and a context manager. A script exits and takes the
wrapping with it; a test does not, and one test leaving the shared engine
metered would charge bytes to a scope nobody opened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:51:10 +05:30
parththakkar106andClaude Opus 5 be66780a26 Size the stress fixture from what production actually holds
The old fixture was wrong in both directions at once and happened to land near
the right total. Actions were modelled at ~2.1 KB against a real 886 B of
text, and adventures at 200 actions against a real 607. Width flattered,
length did not, and length is what a page load pays for.

Re-sized from the 2026-08-17 measurements: 600 actions, 1700 B of narration
alternating with a one-line player input, 232 KB of context_snapshot a row.
The page-load shape now reports 606.0 kB against the 589.5 kB measured on
production's longest adventure -- 2.8% out, where the old defaults were 28%
out on a story a third of the length.

Filler text is now generated word by word instead of one sentence repeated.
That matters for what comes next: the repeated string compresses 313x and the
generated prose 3.7x, so any compression ratio measured against the old
fixture would have been fiction, and shrinking context_snapshot is the open
question it exists to answer.

context_snapshot also gains a flag of its own rather than being hardcoded, and
the 74 KB figure in the comment -- inherited from models.py -- is corrected:
the real column averages 163 KB a row across the table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 13:48:23 +05:30
parththakkar106andClaude Opus 5 70d024da62 Check the egress work against the database it actually runs on
Everything measured so far ran on SQLite against a synthetic fixture, so two
claims were still on trust: that migration 38 spells BYTEA correctly for a real
server, and that the byte figures survive psycopg's encodings.

Both hold. Production reads schema_version 41 with embedding_blob bytea and
embedded boolean present, the backfill is complete at 134/134, and the packed
vectors are 5.04x smaller than the JSON on real data -- 30,971 to 6,144 bytes a
memory, as predicted. stress_session now takes AIDND_STRESS_DATABASE_URL, and
against a throwaway Neon database every shape lands within 0.5% of the SQLite
run: the warm turn is 121.1 kB against 122.3, with memories down to 1.7 kB of
it.

The harness writes, so it refuses any target whose name does not say stress or
scratch -- pointed at the production database it stops rather than seeding it
with a fake user and 200 fake turns. It also empties a Postgres target before
building, which a fresh SQLite temp file never needed.

Two corrections fall out, both recorded in plan/13. The page-load model has the
wrong shape: real actions are half the fixture's weight but real stories run to
607 actions, not 200, so the worst real page load is 589.5 kB. And the decision
to leave context_snapshot in the database costed egress but never storage --
it is 88.9 MB of a 99.6 MB database against a 512 MB free tier, which is the
ceiling this deploy will hit first.

Measured with counts and octet_length sums only. No user content was read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
2026-08-17 12:37:23 +05:30
parththakkar106andClaude Opus 5 b7e53ae581 Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.

Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.

    one turn      3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
    run_post_turn 3,139.1 kB -> 0.7 kB
    Insights      3,223.7 kB -> 117.9 kB
    Memories drawer  ~3.1 MB -> 23.6 kB

A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.

The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.

memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.

Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:28:00 +05:30
parththakkar106andClaude Opus 5 c56864877a Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.

It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.

Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.

The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.

Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:17:44 +05:30
parththakkar106andClaude Opus 5 7ee5ceea6c Measure database egress in bytes, from inside the repo
Both egress blowouts this project has had were one query fetching a column
nobody read, and a statement count would have shown nothing wrong in either.
So the meter counts bytes, at the DBAPI cursor -- everything that crosses
that line crossed the wire.

tools/stress_session.py drives a production-shaped adventure through the real
routes with only the network faked. It reproduces both figures measured
directly on production: 426.7 kB for a 200-action page load against 423 KB,
and 3,258.7 kB for one turn against 3,153 kB.

The memory bank is on by default, which is the whole point -- the previous
harness ran without an embedding model, so retrieval returned early and the
heaviest read in a turn never happened. --no-embeddings reproduces that
deliberately, and the gap is 29x.

It also turned up two callers the production SQL could not see: run_post_turn
walks the whole bank again every turn, and Insights pays for it a third time.
A played turn costs ~6.4 MB, not 3.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-16 21:11:56 +05:30