main
26
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0c1ba836ba |
v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each verified before the next. Accepted by the owner with a documented reference-model limitation. No schema, bundle format, setting default, lineage, authority or protocol-cleanup change. - B2.1 ranking: the retrieval query is the player's input plus a bounded scene context (state scene + end of the newest narration), embedded in one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical, where lexical is a rarity-weighted share of the input's words, computed per turn over the candidates with no index. Scores and the query are recorded per used memory; pins and redundancy suppression unchanged. - B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest and newest memories are kept, the smallest coverage hole goes first, least-recently-used breaks ties and remains the fallback. Bounded; pins never evicted; frozen-bank protection kept; reads no text or vectors. - B2.3 bounded memory creation: a block longer than 2,000 tokens is shown to the summariser as head + tail with an omission marker, inside the same budget; shorter blocks unchanged; the marker is never stored. - The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt experiment was measured on the reference model, showed no reliable improvement for the target failure (0/5 under both prompts, with new "Memory:"-prefix, second-person and length regressions), and was reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral helper. - tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus the failed block, a deterministic fidelity checker, and a real-model shipped-vs-experiment measurement. - tools/memory_diagnostic.py: ranking replica uses production scoring; ranking_crowded, ranking_context_dependent and independent_full fixtures; per-turn isolation and provenance. - tests: B.1's two strict xfails are now ordinary passes; ranking, eviction and excerpt tests; summariser acceptance tests kept apart from diagnostic-measurement tests. - DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`, which matches nothing; now the OR form. - docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and release criteria 12-13), planning README, VERSION v4.3, reports/v1.1/V1.1-WP-B2-REPORT.md. Deterministic independent-memory recovery: PASS (independent_full fails on v1.0.0 at creation and returns recovered_through_memory_independent here). Reference-model independent recovery: FAILED on the precondition-valid attempt, at memory creation: the summariser omitted a player-established fact from a block it received whole. Accepted as a documented v1.1 residual and carried into the release gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY |
||
|
|
f8d401029f |
Stop a turn locking out its own memory bank, and let the long run notice
The first M01 trial with the memory bank on was 26 turns on a GPU host. It
accepted every turn and reported "complete". It also wrote two memories and
no summary, and logged 180 `database is locked` errors, while derived status
still read `idle`.
The cause was a single uncommitted UPDATE. Retrieval bumped each used
memory's counter before the model call, and the turn commits only after the
reply has streamed. SQLite has one writer, so the turn held the write lock for
the whole reply. Every post-turn memory, summary and status write in that
window waited out the five-second timeout and failed. Recording the failure
needed a write as well, and without a rollback first it raised
PendingRollbackError. The loss therefore reached the log and never reached
the status the Insights panel reads, which F08 forbids. The draco run never
hit this because the bank was off there.
- `retrieve_memories` now only reads. `record_use` writes the counters in the
turn's single commit, so a turn that never lands counts nothing.
- The post-turn task's outer handler rolls back before it records a failure.
The harness could not have caught any of this. It read three prompt sections
under names the builder does not use: `memories` (really `used_memories`),
`story_history` (really `history`/`recent_history`), and a `knowledge` prefix
that matched the fixed instruction section instead of the imported passages.
Memory tokens read 0 whatever the prompt held, and the in-history and
in-memories recall checks could never come out true. The labels are now
constants, pinned by a test against a prompt the real builder assembled.
The harness also stops at the first sign of failed post-turn work. It checks
/derived and new server.log lines after every turn, keeps its log position
across --resume, and waits for background work to settle before its final
checks. A run with no memories or no summaries now ends "failed", not
"complete".
Both new application tests fail on
|
||
|
|
480414efe0 |
M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or Inspiration, and the class is load-bearing rather than a label: it decides the words a passage is framed with in the prompt, the weight it carries when passages are ranked, and which budget it competes in when the context is tight. This is a separate subsystem, which is the Phase 0B decision (IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification, provenance, content identity, chunking, an index or a lifecycle, and they were not promoted into something that does. Nothing here reads or writes one. The subsystem, in backend/app/knowledge/: classes the three classes, their weights, and the prompt framing chunking deterministic, heading-aware, 60-800 tokens, no overlap fts SQLite FTS5 with porter stemming; scoped and bounded in SQL importer validate, hash, store, chunk, index — in one transaction embeddings local Ollama vectors through the shared provider retrieval query construction, hybrid merge, rerank inject the budgeted cut and the rendered prompt sections Relevance admission is a separate stage from ranking, and that separation is the milestone's most expensive lesson. An independent review found the first implementation deciding relevance with a floor expressed as a share of the best candidate — which the best clears by construction — so a passage was admitted on every turn regardless of the scene. A query about tide tables and container tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden Canon among them. So the pipeline is now: candidate generation -> admission -> ranking -> class weighting -> budget Admission reads raw, candidate-set-independent signals: the cosine the model returned, and how many distinct meaningful query terms a passage contains. Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero is not zero. Normalization decides order among things that matched; it can never decide whether anything matched. Authority is applied after admission, so a class orders what matched and never rescues what did not. Retrieval may therefore return nothing, and on a scene unrelated to the library it does. The other decisions that each replaced an obvious wrong one: - The class multiplies relevance rather than adding to it. An additive bonus satisfies "Canon outranks Reference" and makes "do not include irrelevant Canon" impossible, because a large enough constant wins on its own. - The semantic floor is measured, not guessed: 113 production-path pairs against nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at 0.36-0.56, and 0.58 sits between them. Because it is a property of that model and not of cosine similarity, it is keyed to the model rather than applied to whatever is configured: an embedding model with no measured calibration in this build does not borrow the number. Semantic admission is skipped, the campaign retrieves lexically, and the reason is stated in the knowledge status and in the turn's provenance. Degrading to lexical keeps the library usable; lending the threshold to an unmeasured model is how the admitted-everything defect would return. - One lexical term is not evidence. Two distinct meaningful terms, or one that is neither a standing campaign entity nor a negligible share of the query. The stop list grew from 42 words to 261, all function words — no subject matter, because a stop list that removes subject matter stops finding "The Silver Key". - Lexical retrieval is a production path, not a fallback. It finds the proper nouns and invented terms a setting bible is made of, and the library is fully usable with no embedding model configured. Safety is structural rather than filtered. Imported text reaches the prompt whole, inside a section that says what it is, under a rule stating the authority order in words and refusing every instruction inside it. No endpoint accepts a filesystem path, so H08 has no mechanism to escape from. Nothing renders imported content as HTML, so a script tag is five visible characters and a remote image is never fetched. Import, chunking, indexing, retrieval and a turn open no socket at all; only embeddings do, through the endpoint allowlist the memory bank already uses. Provenance is the rendered text, not a foreign key: deleting a source cannot turn a historical turn's evidence into dangling ids. Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5 virtual table attached to knowledge_chunks as a DDL hook so it is created and dropped with the table it indexes. Migration 92. A pre-M7 database opens unchanged and needs no sources to play. Bundle: the source content and the reader's judgements about it travel; the passages, index rows and vectors are rebuilt on import, so a restored campaign is searchable immediately without a reindex step. One runtime dependency: python-multipart, Starlette's multipart parser. It is what makes the upload surface possible, and the upload surface is why no pathname is ever accepted. The test doubles were the reason the defect shipped, so they were corrected too. The retrieval stub scored unrelated text at 0.06-0.20 where the real model scores it at 0.43-0.44, and its docstring said it had deliberately removed the constant component that "would put a similarity floor under every pair" — which is exactly the property real models have. The stub now has that floor, one test fails if it is ever removed, and another reproduces the superseded rule and asserts it is still fooled by the same fixture. Run against the pre-corrective implementation, the new suite fails 13 of 18. Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of which mocks nothing between itself and Ollama and re-measures the similarity separation on every run. 43/43 checks in a real Firefox, reproduced. Docker build clean. Four other defects found by review or by the browser run were fixed here rather than carried: an unreachable relevance constant that appeared to enforce something and did not; acceptance tests using the wrong fixture files, so G07's trap was never exercised; a bidirectional override surviving into displayed filenames; and, from the implementation pass, the Insights panel showing M5's two state sections as raw keys and the source inspector refetching on every keystroke. M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK REQUIRED. Both blocking findings are closed, and closeout resolved the embedding-model calibration boundary the corrective pass had left as debt. planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective closeout and the closeout verification in sequence, none overwriting another. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b |
||
|
|
a6e9c7a32b |
M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against
|
||
|
|
8652fe7cd8 |
M2 review: two regressions the green suite hid, and the reports
The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit
|
||
|
|
d72f7c1bda |
Stop paying twice for a block a retry can still throw away
A memory whose block ends on the newest action is the one memory a player can reach: retry and take-switching both refuse anything else. Each retry of that turn withdrew the memory and wrote it again, and a block closes every six actions while a normal turn writes two, so that was one turn in three. SETTLE_SLACK asks for one action past a block before the block is summarized. The block is still MEMORY_INTERVAL actions; only the moment moves. This is not the pre-SP4 holdback returning: that one was about a retry rewriting text in place, which sibling attempts and forget_node settled, and correctness still rests on the withdrawal rather than on the slack. Undo and delete can carry a summarized node back to the tip, so the withdrawal path stays reachable, just rarely. The slack buys nothing back. The block that just closed is still in the history window in full, so a memory of it says what the model can already read. Tests: the settling suite asserts the new rule and that a retry at the tip finds nothing to withdraw; the rewrite suite builds thirteen actions so both of its blocks settle; the one path test that ended on a block boundary sets the slack to zero, because it is about which actions a block is read from rather than about when a block forms. 632 green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NcrxCJjqgDvAkamKeLWdn |
||
|
|
0633cb624e |
Run the new memory prompt back over an old bank
The prompt change only reaches memories written after it. plan/18 decided to leave the existing ones alone and let eviction age them out at memory_bank_capacity, on the grounds that re-summarizing would duplicate whatever was still in the bank because nothing deletes the old rows. That was wrong about the only option. A memory can be rewritten in place. The row carries more than its text — whether it is pinned, how often it has been retrieved, and the node it hangs off, which is what makes a fork inherit the right memories — and rewriting `text` keeps all of it. Deleting the bank and rewinding the cursor would lose that, and would trickle memories back at MAX_MEMORIES_PER_RUN per turn, so an adventure nobody is playing would never recover. tools/rewrite_memories.py does it. Without --write it makes no model calls and only reports the scope; --write rewrites, --embed re-embeds in the run rather than leaving it to the app's post-turn pass. It reads whichever database the app reads, so it works against the hosted Postgres as well as a local file. Two things it needed from the app. `summarize_block` is now the one place a memory prompt is assembled, and the post-turn pass calls it too — a backfill that built its own prompt would be writing memories with a prompt that never shipped, and nothing would report the drift. `source_block` reads a memory's block back out of the story, which nothing has ever had to do: it reads on the lineage of the branch the memory was written on, not the branch being played, because after a fork the same depths hold different actions on each side and a read through the adventure's path would summarize the wrong story silently. It also excludes the discarded attempts at a retried turn, and tolerates a block an action has since been deleted from. Left alone: a memory with no source range, which is hand-written or migrated by 62 and may be the player's own words; a memory whose actions are gone; and an adventure whose owner has no API key, because summarization spends the user's own key by construction and never the shared demo key. --api-key/--model/ --endpoint override that, the last of them aiming a run at claude_shim.py. The vector is cleared for every rewrite, because the stored one describes wording that no longer exists. Re-embedding always uses the owner's own embedding model, never --endpoint: a vector only means anything against the vectors it is ranked beside. 17 tests, 627 green. The fork case is the one that would fail quietly, so the test builds a fork whose depths hold different actions on each side. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Tqgupw5CZGjSZrUTNUd4fW |
||
|
|
07192767d8 |
Give a memory a word ceiling, after measuring one
Ran the memory prompt end to end against a real model as a controlled A/B: one story generated through the app's own build_context a turn at a time, then both prompts run over the same blocks, so the story is held constant and the prompt is the only variable. Fresh process per call, so neither arm sees the other and the model is never told what is being tested. The control is the exact prompt from 9cdcb55. The reported fault reproduced. Two consecutive memories written from one story minutes apart came back in two different persons — "You crept low through the mist" and "The player asked Gwen to". With the cast brief both named Kaelen. The control also inverted who acted on a move whose player text was "grab her wrist and pull her down", which is the failure the brief predicts: with no cast there is nothing to say whose wrist "her wrist" was. It also found something the prompt review had not. "1-2 plain sentences" is not a length, and the same model wrote 34 words for one block and 105 for the next. A 105-word memory is a paragraph, and `memory_top_k` injects five every turn, so the bank's running cost was set by a number nobody had ever stated. MEMORY_MAX_WORDS states it, and the prompt now says which details to keep when trimming: the ones a later scene could turn on. Re-run over the identical story, the same two blocks came back at 32 and 58 words, still named, still third person, still carrying the camp map, the strongbox behind the second tent, and the strap frayed near through. Variance is the real gain — 34..105 became 32..58. Overshooting 50 slightly is expected. Models exceed word budgets, which is why builder.length_hint already carries a buffer for the same reason. What this does not show: the run used a Claude model, and the app talks to an OpenAI-compatible endpoint whose weaker models are why worldstate/parse.py tolerates trailing commas. The prompt is followable and the brief supplies the missing information; a weaker model is not proven to comply as well. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
1b5e4c56cd |
Tell the summarizer who the characters are
`_create_due_memories` sent six actions of second-person prose and nothing else — no protagonist, no cast, no setting, and no instruction about what person to write in. So for `You push the door open. She grabs your arm.` the only honest memory was "You entered a room and she stopped you", which names nobody when it is retrieved forty turns later. The framing wandered too: with no rule, the model picked a person per call, and one bank ended up holding "You entered the crypt", "The player entered the crypt" and "He entered the crypt" for the same kind of event. Both prompts now carry a cast brief and a framing rule: third person, the protagonist by name, other characters named rather than left as bare pronouns. The rule states its reason, because a memory really is read in isolation much later and a model told why complies far more consistently than one handed a bare instruction. The cast comes from the story cards, not from `stat_schema`. Every schema NPC is already turned into a card at adventure creation, deduplicated against the hand-written ones by name, so the cards cover schema NPCs, an author's own cards, and an adventure with no RPG layer at all through one path instead of three. Keyword matching alone was not enough, and finding that out changed the design. Built that way first, the brief for "She grabs your arm" listed the protagonist and nobody else: the block that most needs a cast is exactly the one written in bare pronouns, and Gwen's trigger keys include "her" but the text says "she". So matched cards come first and the remaining slots are filled with the other character cards. Places and items are not topped up — an unmentioned tavern is not who "she" was — though a place that is mentioned still matches normally. The asymmetry with the turn prompt is deliberate: an untriggered card is wrong as lore and right in a roster, because the roster answers "who could these pronouns be" rather than "what is on stage". Fixed descriptions only, never live values. `Gwen: trust 40 (wary)` in the brief would make the same event summarized at two different times come out framed differently, which is the fault this removes. An adventure with no persona still gets the cast and the setting, and the model is told to write "the player". One with nothing to say sends byte for byte the prompt it sent before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
c0cd6fa7ce |
Stop a full memory bank from shutting itself
Eviction ranked on use_count first. A memory written this turn has never been used, so once every survivor in a full bank had been retrieved even once, the newborn was the lowest row in the bank and was marked forgotten by the same post-turn run that wrote it — one pass after the one that embedded it, before retrieval ever saw it. That state never ends, because counts only go up. The bank an adventure happened to hold when it first filled is the bank it keeps for good, and everything the story does afterwards is summarized, evicted and never ranked. A test now plays four turns against a full bank and asserts the survivors are the new memories, not the opening ones; before this it kept the opening three and forgot all four. Order on last-touch instead, with the count as the tiebreak. A new memory carries the newest timestamp there is, so it is the safest row in the bank rather than the most doomed, and it has until something outlives it to prove itself — the "protect the newest" behaviour falls out of the ordering rather than being a count of rows to spare. Little is given up: a memory that is genuinely used stays recently-used by being retrieved, so the two orders only disagree about memories that mattered once and have not been wanted since. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
4af6e17406 |
Answer the review, and keep the opening node's bank
Nine findings from a review of the phase-14 stack. The one about a retry withdrawing a memory is not a bug — a memory anchored to a node describes that node, and it goes when the node goes. The root is the exception, and it is the only one: migration 62 parked every memory written before memories had coordinates on depth 0, so withdrawing the opening node would retire a whole bank nobody attached there. A memory with no source range covers no story and now stays; a summary that genuinely ends there is still withdrawn. The rest are repairs. * The adventure list quoted whichever attempt was written last rather than the one the story tells, so switching back left the index disagreeing with the page. * A v1 import gave a typed memory no depth, rebuilding the NULL the migration exists to remove — invisible until the imported adventure forked. * The action cap counted a v1 file's turns, and a turn expands into a row per saved attempt, so a file inside the cap could write a multiple of it. * Forking a live node on a borrowed ancestor promoted a sibling on a branch the caller never named. It is a branch switch, and now says so. * Switching attempts left the state, status and memory panels reading the previous take: the story does not change length, so nothing keyed on its length noticed. Same class as the branch-switch bug this phase already fixed. * A retry after switching back numbered the new attempt into the middle of the group instead of the end. * The cursor backfill numbered every action in the table once per adventure; correlated to the adventure being updated, it is an index lookup instead. * Renaming a branch answered own_actions=0. And one behaviour change recorded rather than repaired: script-visible history and actionCount no longer count blank-text rows. That is the right shape and there is no reading compatible with both, so plan/14 says so. 409 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g |
||
|
|
2d38a162d4 |
Give every memory a node, and show only the ones on your path
A hand-written memory used to carry a NULL depth, described in the model as "belongs to the adventure rather than to a path". That sounds harmless and is not: a NULL is a coordinate no fork can cap, so a note typed on one line followed the reader onto branches whose story it never described. It takes the head now — the story you were reading when you wrote it — and obeys exactly the rule a summarised memory obeys. The unanchored escape clause in lineage.Path.clause existed for that single case and is deleted rather than left unused. Its docstring argued that a capped depth would drop a typed memory the moment its branch stopped being the newest entry; anchoring answers the same worry better, because the memory is not exempt from the path, it is on one. The drawer now shows the path being read and nothing else, filtered by the clause retrieval itself uses, so the bank you can see is the bank the model can see. Nothing is stranded: a memory lives on a branch, switching to that branch shows it, and deleting the branch deletes it. Pinning decides order, the path decides existence. Migration 62 lands existing NULL-depth memories at depth 0 of their branch rather than at the tip. 0 is at or before every fork point, so every memory stays visible from exactly the paths it is visible from today — nobody's bank loses a row on deploy. The tip is the tidier-sounding choice and would have emptied them out of every branch forked earlier than they were typed. This supersedes the on_path flag and the "another branch" badge from earlier today; anchoring makes them redundant, and they are removed. Four tests changed because they asserted the old contract, not because they broke. The one worth reading is the pair replacing test_a_hand_written_memory_is_not_lost_at_the_first_fork: typed on shared trunk it still survives a fork, and typed on ground the fork never travelled it no longer follows you. 402 tests. Verified on tools/branch_fixture.py: each branch's drawer holds its own memory and not the other's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g |
||
|
|
0a12d9cd47 |
Make a retry a node, not a rewrite
Every attempt at a turn is now its own row at the same (branch, depth), with `live` naming the one the story tells. The JSON repeating group on `actions.variants` is read one last time, by a migration that writes it out as the sibling rows it always described, and then goes unread. The snapshots turn around with it: an action carries the state it left behind rather than the state it started from, because attempts at one turn share a starting position and differ exactly in their outcome. Rolling back is "what the node in front left behind", one lookup on the path, and it is what undo and retry now both read. And the memory holdback goes. It existed because retry rewrote a row under a mark that had already moved past it; a retry writes a sibling now, and replacing what a coordinate says withdraws what was derived from it — the same repair undo and delete already made. The assembled prompt is still stored once per turn: it moves with the live flag, so a superseded attempt keeps only the few hundred bytes that were its own. Measured on the 600-action fixture: 700 rows for the same 600-turn story, prompt archive byte-identical at 0.50 MB, index 1.8 kB and page load 62.7 kB unmoved. 347 tests green. `tests/test_story_tree_baseline.py` and `tests/test_retry_variants.py` pass unmodified — SP4 was allowed to move the baseline for the variant-count semantics and did not need to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
c51531709d |
Mark the story with a node, not with a count
The memory bank and the story summary each kept a cursor: how many story actions they had already covered. A count is a position in a list, and this list moves — delete an action in front of the mark and every later one slides down a slot, so the mark now covers one it has never read. All the cursor bookkeeping existed to patch that up. Both marks are now (branch_id, depth): the node up to and including which the work is done. A depth is a coordinate along a path, not an offset into a list, so nothing in front of it can move it. That deletes rather than rewrites `position_of_index`, `note_action_removed`, `_rewind_cursors_to_index`, `prune_dangling_memories` and the every-pass clamp in `run_post_turn`. A memory hangs off the node its block ends on, so a fork inherits its ancestors' memories without copying any, and retrieval selects through the branch clause over the *whole* lineage — recall is long-range by definition and cannot be windowed. Measured: 1,807 B on a story forked twenty times against 1,823 B on a flat one of the same length. Migrations 53-56 translate the old counts into nodes. They rewrite `adventures` and not `actions`, so this one needs no VACUUM FULL. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
d3756abdaa |
Give every action a branch and a depth
Phase 14 SP1. The tree goes into the schema and nothing reads it yet: a `branches` table, `branch_id`/`depth` on actions and memories, a head pointer on adventures, migrations 46-52, and a server-side backfill that re-reads every existing adventure as a tree with one branch. `depth` holds the number `index` already held, gaps included, so no story changes — a linear story *is* a tree with one branch, which is what makes the SP0 baseline passing unmodified the pass condition rather than a hope. The writer had to come with it. No migration will ever visit a row written after it ran, so columns backfilled today and populated next subphase would leave a hole exactly the width of one deploy, and from SP2 on a row without a branch is a row no read can see. `app/tree.py` owns that: one module, because a node written without a branch fails by disappearing rather than by raising. Three things the schema itself insisted on: - `adventures.head_branch_id` is a plain integer, not a foreign key. Pointing both ways makes the two tables a cycle create_all cannot order, and its escape hatch needs an ALTER SQLite does not have. It is a cache, and a head naming a branch that is gone recovers onto the root. - `lineage` is NOT NULL, so the backfill inserts `'[]'` and fills it in a second pass guarded on `json_array_length(lineage) = 0` — not `= '[]'`, because Postgres `json` has no equality operator. - SQLite will not drop a column a foreign key names, which is how two existing tests broke: they simulated an old database by rewinding the stamp while leaving the new columns in place. Every ADD COLUMN migration is now idempotent, and `tests/test_tree_migration.py` builds a genuine schema 45 by rebuilding three tables from frozen DDL so the real ALTERs run. 297 tests green, 14 of them new. `branches` costs 0.1 kB of a 733.5 kB turn; page load and index are byte-identical to the recorded figures. The deploy that ships this needs one `VACUUM FULL actions;` on the direct endpoint afterwards — it rewrites every row. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
2c5909a268 |
Drop the JSON vector column, and fix what was hiding behind it
Migration 38 left memories.embedding in place so a rollback could still find the vectors. Production has since been verified reading from embedding_blob, so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone reads. Removing it surfaced a live bug. Changing your embedding model is supposed to throw the bank's vectors away and let the post-turn pass rebuild them, because two models' vectors are not comparable. The settings route did that by nulling memories.embedding -- correct until 38 moved the vectors, after which it cleared the dead column and left the blob intact with `embedded` still true. _embed_pending filters on `embedded IS FALSE`, so it never saw those rows and the bank went on ranking against the old model's vectors permanently. Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a different-width model scores every memory zero and retrieval returns whichever rows happen to sort first; a same-width model scores plausible garbage. The bulk clear now sets both columns. It stays a bulk UPDATE rather than going through set_vector -- loading the rows is the cost that whole path exists to avoid -- so set_vector's docstring now names it as the one caller that legitimately writes those columns by hand. No cache invalidation is added: clearing `embedded` drops the rows out of the catalogue query, and set_vector evicts each entry as the re-embed puts it back. test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests the backfill, since create_all no longer produces the column it converts from, and asserts 42 removes it at the end of a full bootstrap -- 38 reads that column and 42 drops it, so an upgrade that reordered them would arrive with an empty bank. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
b7e53ae581 |
Rank the memory bank without reading the memory bank
Retrieval walked adventure.memories, so every turn loaded every row of the
bank with its vector attached -- 3.1 MB, 96% of everything a turn read. It
now asks SQL which memories are in play (an id and a flag per row), ranks
against vectors held in process, and fetches text only for the five it picks.
Two more callers were doing the same thing and the production SQL could not
see them: _evict_over_capacity walked the bank to count it, and _embed_pending
walked it to find the rows with no vector. Both are counts and filters the
database can do without sending anything back.
one turn 3,258.7 kB -> 723.4 kB cold, 122.3 kB warm
run_post_turn 3,139.1 kB -> 0.7 kB
Insights 3,223.7 kB -> 117.9 kB
Memories drawer ~3.1 MB -> 23.6 kB
A played turn is turn plus post-turn work: 6.4 MB down to 123 kB.
The cache needs no invalidation callbacks, which is what makes it safe. A
vector can only change through set_vector, which drops that one entry;
anything that removes a memory from play leaves the catalogue query, and
entries missing from the catalogue are dropped on the next read. So eviction,
deletion and pruning have nothing to remember to call.
memories.embedded joins the blob, for the same reason actions.variant_count
sits beside actions.variants: with the vector deferred, every "is this
embedded?" check would otherwise be a 6 KB lazy load, once per row.
Capacity drops 200 -> 80, on retrieval quality as much as cost -- ranking two
hundred memories to pick five buries the five. Eviction was measured at scale
first: trimming 100 to 80 costs 0.8 kB and reads no vectors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
|
||
|
|
c56864877a |
Store embeddings as packed float32 instead of a JSON list
A 1536-dimension vector spelled out as JSON decimals is ~31 KB. The same
numbers packed as float32 are 6,144 bytes, and the whole bank is read on
every turn, so those bytes are paid over and over.
It is a format change, not a precision trade: the endpoints compute in
float32 and render that into JSON, so converting back recovers the original
bits exactly. Nothing is re-embedded and no API call is made -- migration 38
is a pure repack of what is already stored.
Unlike migrations 36 and 37 this backfill cannot be expressed in portable
SQL, so it comes through Python, batched, and pays a one-time read of every
vector to stop paying three megabytes a turn.
The JSON column stays, still written through set_vector, so a rollback finds
the vectors intact. Reading from the blob comes next; a follow-up migration
drops the old column once that is verified.
Migration SQL can now be a {dialect: sql} map -- BLOB and BYTEA have no
common spelling, and every Postgres deploy replays this one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
|
||
|
|
47e33fa311 |
Read a window of the story per turn instead of all of it
Two reads still grew without bound after the snapshot fix. `Action.variants` holds every discarded retry attempt, but a list response only needs how many there are — so each retry permanently added ~5 KB to every later load of that adventure. Defer the column and keep the count beside it (migration 37, backfilled server-side), with set_variants() as the one write path that keeps the two in step. `story_actions()` walked adventure.actions, then every caller threw almost all of it away: the builder concatenates the story and immediately cuts it back to the token budget, the NPC check looks at the last 6, retrieval at the last 4, the cursor clamp only wants a count. A turn on a 200-action adventure read 839 KB to use ~70 KB, and grew with every turn played. app/context/ history.py serves those shapes from SQL; window_covering() measures the actions it fetched and projects how many more it needs, fetching only the part it does not already hold. Memorybank cursors move to position_of_index() and settled_count()/settled_slice() — same arithmetic, no full list. The scripting pipeline still receives the whole history per AI Dungeon's API, and every helper reuses adventure.actions when it is already loaded, so a scripted adventure pays what it always did and never twice. Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete 115 KB -> 5 KB. Verified the window builds a byte-identical prompt to the full story across budgets from 1K to 100K tokens, with and without the retry exclusion - this is a cost change and nothing else. Cursor helpers checked against the old list arithmetic, including after deleting a middle action. Counts are real SELECT count(...): Query.count() wraps the entity select in a subquery, so the SQL named every deferred column and the egress guard could not tell it apart from a bulk fetch. 139 tests pass; the four new guards verified by sabotage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |
||
|
|
7c538c8235 |
Stop retried and deleted actions from corrupting story context and memories
Three fallout bugs from keeping the retried action row alive (
|
||
|
|
1dd31086c1 |
Add power-user AI Chat page; centralize demo-key model pinning
AI Chat is a plain scratchpad for talking to a model directly — no story context, scripts or world state — for poking at models, prompts and endpoints without starting an adventure. Power users only: the router 404s (rather than 403s) for everyone else and the nav link is hidden. The conversation lives in localStorage, so there's no new table or migration. is_power_user() now also returns True in local mode: it's the operator's own machine and their own key, the same reasoning that makes the provider debug log local-only. Alongside that, the rule keeping the shared demo key off paid models now lives in exactly one place. It had been duplicated into the chat router, which is how one copy eventually drifts: - resolve_provider_config() takes an optional model_override and is the only place the whitelist is applied, so turns, AI Chat and the connection test all inherit it. An override is a per-request preference, never a grant. - ProviderConfig.__post_init__ refuses to exist when api_key is the demo key and the model isn't whitelisted. It keys on the key itself rather than the using_demo flag, so a mislabelled config can't slip past, and it raises so a future path that bypasses the resolver fails loudly instead of billing. - The demo branch still pins endpoint_url too — a user-controlled endpoint would leak the key itself, which is worse than spending it. Provider gained chat(messages, ...) beside generate(), both delegating to a shared _stream(url, body); completion-mode endpoints get the messages flattened into a labelled transcript. Settings' /models fetch moved to list_endpoint_models() and is shared with /api/chat/config. Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub resolve_provider_config — the point is to exercise the real BYOK-vs-demo decision and assert on what the provider actually received: off-whitelist override pinned, off-whitelist Settings.model pinned, redirected endpoint pinned, BYOK passed through untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx |
||
|
|
dab1807118 |
Roll back script_state on undo/retry; fix retry double-apply
The shared per-adventure script_state ("scoreboard" scripts write to) was
never reverted by undo, and retry re-ran the output hook on top of the already-
mutated state, double-applying its changes (e.g. "+10 gold" became +20).
Each action now snapshots script_state as it was immediately before its own
hooks ran (new Action.state_before column, migration 25):
- undo restores the turn's first-action snapshot, prunes memories that
summarized the removed actions, and takes the turn lock against races.
- retry restores the AI action's snapshot before regenerating.
Story-card mutations are not reverted (documented limit). Adds the project's
first test suite: unit + full HTTP integration through the real scripting
engine (14 tests). See plan/11-state-revert-and-retry-fix.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
de4db373f2 |
Phase 8: optional accounts, per-user data, shared demo key
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me, register upgrades the guest in place, login/logout, per-IP rate limits. Every router scoped by user_id; Settings become per-user with the API key Fernet-encrypted at rest and write-only through the API. Users without a key get a server-funded demo key (OpenRouter free models, 20 turns/day, memory bank disabled on demo turns). Public read-only demo scenarios (seed_demo.py); debug log restricted to local mode. Frontend: auth modal + guest nudge, 401 re-establish/retry, demo banner and key management in Settings. Migrations 13-23 adopt existing data under a local user and encrypt stored keys. Verified: migration on a copy of real data.db, two-session isolation + register/login via curl and Chrome, demo cap 429, live OpenRouter turn through the encrypted-key path, vite build + oxlint. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |
||
|
|
253b533d3b |
Fix remaining code-review findings (15 bugs; #11 skipped as AID-compatible)
Backend: - provider: fall back to parsing a plain JSON body when a server ignores stream=true (was: silent empty turn); error if response has no text (#5) - memorybank: clamp cursors after undo/retry shrinks the action list, and translate summary_cursor (list position) to an Action.index boundary before comparing with Memory.source_end (#7) - memorybank: pinned memories now count toward the top_k budget (#8) - memorybank: cosine() returns 0.0 on dimension mismatch; changing the embedding model clears stored vectors so they re-embed (#9) - scripting: MAX_STORY_CARDS cap now counts cards inserted during the hook, so a script can't add unbounded cards in one turn (#10) - settings: /test tolerates non-dict JSON from /models (#13) - scenarios: import accepts worldInformation as a story-card source (#14) Frontend: - per-key debounce timers in PlotPanel and ScenarioEditor — editing two things within 600ms no longer drops the first save (#15, #16) - Continue button no longer discards typed input (#17) - failed retry resyncs actions from the server instead of leaving the removed action missing (#18) - Settings save/test surface errors instead of hanging on Testing… (#19) - InsightsPanel ignores stale responses from superseded requests (#20) - placeholder scan includes story-card trigger keys (#21) addStoryCard returning the 0-based index (falsy for the first card) matches real AI Dungeon per the scripting guidebook — kept, documented (#11). Statuses updated in CODE_REVIEW_FINDINGS.md; stale entries for previously fixed items (#1-4, #6, #12) corrected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |
||
|
|
db9f904222 |
Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |