c1a73b3d77e48491196e8887ee5abc2f818af05e
16
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
1b5e4c56cd |
Tell the summarizer who the characters are
`_create_due_memories` sent six actions of second-person prose and nothing else — no protagonist, no cast, no setting, and no instruction about what person to write in. So for `You push the door open. She grabs your arm.` the only honest memory was "You entered a room and she stopped you", which names nobody when it is retrieved forty turns later. The framing wandered too: with no rule, the model picked a person per call, and one bank ended up holding "You entered the crypt", "The player entered the crypt" and "He entered the crypt" for the same kind of event. Both prompts now carry a cast brief and a framing rule: third person, the protagonist by name, other characters named rather than left as bare pronouns. The rule states its reason, because a memory really is read in isolation much later and a model told why complies far more consistently than one handed a bare instruction. The cast comes from the story cards, not from `stat_schema`. Every schema NPC is already turned into a card at adventure creation, deduplicated against the hand-written ones by name, so the cards cover schema NPCs, an author's own cards, and an adventure with no RPG layer at all through one path instead of three. Keyword matching alone was not enough, and finding that out changed the design. Built that way first, the brief for "She grabs your arm" listed the protagonist and nobody else: the block that most needs a cast is exactly the one written in bare pronouns, and Gwen's trigger keys include "her" but the text says "she". So matched cards come first and the remaining slots are filled with the other character cards. Places and items are not topped up — an unmentioned tavern is not who "she" was — though a place that is mentioned still matches normally. The asymmetry with the turn prompt is deliberate: an untriggered card is wrong as lore and right in a roster, because the roster answers "who could these pronouns be" rather than "what is on stage". Fixed descriptions only, never live values. `Gwen: trust 40 (wary)` in the brief would make the same event summarized at two different times come out framed differently, which is the fault this removes. An adventure with no persona still gets the cast and the setting, and the model is told to write "the player". One with nothing to say sends byte for byte the prompt it sent before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok |
||
|
|
9ee052c51e |
Give the adventure a persona, so the protagonist has a name
The player had stats but no identity. `stat_schema.player` carried hp and
mana beside `npc.gwen.trust`, but where an NPC has a name and a
description the player had neither, so the block rendered as
`You: hp 100/100` and nothing in the prompt said who "you" was.
Three columns on `adventures`: name, pronouns, description. All
user-only, all optional, and an empty name means the app behaves exactly
as it did before — no backfill, no special case for an adventure that
predates the migration.
They are adventure columns rather than part of `stat_schema` for two
reasons. An adventure with no RPG layer still has a protagonist, and
that is the case this was added for. And `worldstate.schema._initials`
treats every dict inside a stat section as a stat definition, so a
persona placed there would be instantiated, rendered in the guide, and
handed an `initial` value as though it were one.
The paths do not change. `player.hp` stays `player.hp`; only the label
moves, to `Kaelen (player): hp 100/100`, the same way NPC lines already
print a display name beside the id. A path carrying the persona's name
would break the moment a player renamed their character, because
`_history_text` replays every past turn's stored delta into the prompt
and those blobs hold literal `player.hp` strings.
The section sits in the system block. Only the user can edit it, so it
never changes mid-story and stays inside the cached prefix. That is what
makes it free, and it is why the AI must not be able to move it — a
delta aimed at `persona.*` is already refused by `_resolve`, and there
is now a test holding that in place.
The modal that used to appear only for scenarios with `${Placeholder}`
tokens now always opens, and is where the character is named. Persona
and placeholders stay independent: a scenario asking for `${Name}` is
asking its own question. No scenario in the repo uses placeholders at
all, so the overlap is hypothetical.
Phase 2, which feeds the persona and the cast to the summarizer, is
written up in plan/18 and not started. That is where the memory-quality
problem actually gets fixed; this change is what gives it a name to use.
Not yet driven in a browser — plan/18 lists what to check by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
|
||
|
|
f1bebe18d0 |
Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`, `variant_count`, `state_before`, and `world_state_before`, plus `adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in SQLite, so migration 71 quotes it. Nothing outside the migrations read these. `models.py`, `schemas.py`, and `ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders by `id`, and `attempts.renumber`, `context.history.max_action_index`, and `nodes.next_index` are deleted. Two changes keep the migration replayable on a `create_all` database: - `_split_variants_into_siblings` wrote through the live ORM table, so it stopped compiling once migration 66 removed five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`. - Five data passes read columns these migrations drop. Each now calls `_has_columns` and returns early when the columns are absent. `bootstrap` takes a `through` version so a migration test can stop at the schema it asserts on. 555 tests pass, up from 549. Eight of the new cases assert each column is gone after a real schema-45 database migrates all the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK |
||
|
|
47c7800903 |
Report the world-state changes the engine refuses
`apply_delta` records three outcomes for every change the model sends: `applied`, `clamped`, and `rejected`. Everything downstream read only `applied`. A refused change reached the player as an ordinary chip, and reached the model on the next turn as a change that had succeeded. Five parts: - `Action.world_changes` reads `clamped` and `rejected` beside `applied`. Accepted stats carry a `clamped` flag; refusals become `kind: "rejected"` entries. The `fix` key is present only when the engine wrote one, because this property runs for every action of every list response. - The UI separates the three outcomes. A clamp to a standstill reads `no change - at its limit` on a dashed chip, a partial clamp is marked `(limited)`, and a rejection carries its reason. Dashed and dimmed rather than red: a refused change means the rules are working. - The goals line names the milestone id, as `milestones.<id>`. The ids appeared nowhere in the prompt before, so the model could not send one. - Each rejection, and each clamp that moved nothing, builds a `fix` string from the stat definition at the point of refusal. `render_refusals()` renders them into the next prompt above `EMIT_REMINDER`. - `_history_text` replays `applied_delta()` instead of the sent delta, so a past turn's state block shows only what the engine accepted. A clamp that reduced a change but still moved the value reports nothing. If you tell a model its 80 damage became 30, it can treat the shortfall as a debt and send the remaining 50 next turn, which is the swing `max_delta_per_turn` prevents. In the demo scenario, `pokemon_left` becomes `pokemon_fainted` (`type: counter`, `initial: 0`). Starting at the ceiling turned a wrong-signed delta into a silent no-op; counting up puts the wrong sign on the counter rule, which refuses it out loud. The instructions also now ask for `world.turn`, which sat at 0 for a whole playtest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
a408c7b6f7 |
Lay the prompt out so the endpoint can cache most of it
Prompt caching bills on a shared prefix: the endpoint reuses the request up to the first byte that differs from last time and no further. The live world-state block sat third from the top of the system message, so every turn re-priced the instructions, the plot essentials and the whole story history underneath it. The retrieved memories and the rewritten summary did it again. Everything fixed is emitted first now, and everything that moves goes after the history, ordered least-volatile first — which is also where recency serves it best, the reasoning that already put the emit reminder last. The three tail sections that are last for their own reasons stay last. The moved sections are still charged to the token budget; only their position changed. Two smaller halves of the same problem. OpenRouter serves a model from whichever upstream is free and each upstream holds its own cache, so a deepseek model now names deepseek as its preferred upstream — a preference, not a restriction, so a turn still runs if that upstream is down. And the endpoint's usage block is read back off the response and kept per attempt, so the hit rate shows up in Insights and the debug log instead of being assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
3b9e6b3d50 |
Give the length hint a floor, not just a wall
A ceiling alone is a one-sided instruction, and models read it in opposite directions. A verbose one is held back by it; a terse one has nothing to act on except "write only as much as the moment needs -- a typical turn is much shorter" and collapses to two paragraphs. Same prompt, wildly different turn lengths depending on which model is behind it. State a floor as well, so the guidance is a band. The two bounds are deliberately asymmetric -- "must not exceed" for the wall the endpoint enforces, "should not stop short of" for the floor -- so neither reads as a number to hit, which is the property the earlier A/B says decides whether this hint helps or hurts. "Prefer the lower end" inherits the anti-overshoot job the deleted "much shorter" line was doing, but now with a number under it, so a terse model lands on the floor instead of at forty words. Below MIN_LENGTH_FLOOR_WORDS the floor is dropped and the tight-cap wording is left byte-identical: at a tight cap a short turn is the correct turn, and that phrasing is the one measured to keep the state block alive (0/6 truncations at cap 250 against 2/6 unhinted). So this only moves loose caps. MAX_LENGTH_FLOOR_WORDS keeps the share from demanding 555 words minimum at cap 2400 -- a big cap means long turns are allowed, not compulsory. Shipped without an A/B run, deliberately. Two things to watch live: whether a stated range invites landing mid-range on verbose models (drop the share to ~0.25 if so), and whether the state block still survives -- nothing reads finish_reason yet, so truncation is silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
f295893204 |
Ask the model for a turn that fits inside the output cap
max_output_tokens is a hard wall the endpoint enforces mid-sentence. The ```state block is emitted after the narration, so a long turn hits the wall partway through the block and the deltas are lost — silently, since nothing reads finish_reason. builder.length_hint() derives a word limit from the cap ((cap - 50 headroom) * 0.75 words/token * 0.90 buffer) and injects it just above EMIT_REMINDER, which keeps the last slot it needs. Reserved in build_context like the reminder is. Phrased as a ceiling, not a budget. Measured against gemma-4-26b at cap 800, n=5 per arm: no hint 174 words, "keep this turn under about N words" 246, "hard limit ... a typical turn is much shorter" 170. A budget reads as a target to fill — every budget run was longer than every unhinted one, pushing turns toward the wall the hint exists to avoid. Ceiling phrasing still works at tight caps: at 250, unhinted hit finish_reason=length 2/6, hinted 0/6. tests/test_length_hint.py, 11 tests; each mechanism verified by sabotage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |
||
|
|
47e33fa311 |
Read a window of the story per turn instead of all of it
Two reads still grew without bound after the snapshot fix. `Action.variants` holds every discarded retry attempt, but a list response only needs how many there are — so each retry permanently added ~5 KB to every later load of that adventure. Defer the column and keep the count beside it (migration 37, backfilled server-side), with set_variants() as the one write path that keeps the two in step. `story_actions()` walked adventure.actions, then every caller threw almost all of it away: the builder concatenates the story and immediately cuts it back to the token budget, the NPC check looks at the last 6, retrieval at the last 4, the cursor clamp only wants a count. A turn on a 200-action adventure read 839 KB to use ~70 KB, and grew with every turn played. app/context/ history.py serves those shapes from SQL; window_covering() measures the actions it fetched and projects how many more it needs, fetching only the part it does not already hold. Memorybank cursors move to position_of_index() and settled_count()/settled_slice() — same arithmetic, no full list. The scripting pipeline still receives the whole history per AI Dungeon's API, and every helper reuses adventure.actions when it is already loaded, so a scripted adventure pays what it always did and never twice. Measured at production shape: retry tax 5.1 KB -> 0; turn 200 839 KB -> 129 KB and flat from ~turn 50; a 200-turn playthrough 84.5 MB -> 23.0 MB; a delete 115 KB -> 5 KB. Verified the window builds a byte-identical prompt to the full story across budgets from 1K to 100K tokens, with and without the retry exclusion - this is a cost change and nothing else. Cursor helpers checked against the old list arithmetic, including after deleting a middle action. Counts are real SELECT count(...): Query.count() wraps the entity select in a subquery, so the SQL named every deferred column and the egress guard could not tell it apart from a bulk fetch. 139 tests pass; the four new guards verified by sabotage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |
||
|
|
f1bd099ec8 |
Cut database egress 189x by deferring the prompt snapshot
The free-tier 5 GB/month network transfer allowance ran out, which blocks connections outright. The database is only ~55 MB, so 5 GB meant the whole thing was being pulled roughly 90 times over. Cause: actions is 39 MB of that 55 MB -- 541 rows at ~74 KB each, almost entirely context_snapshot, which stores the whole assembled prompt for a turn. Every adventure load and every turn fetched all of it in order to read two small things out of it: the world-change chips under an AI message (Action.world_changes) and the emit block re-attached when replaying history to the model (_history_text). The Insights viewer is the only consumer that wants the whole snapshot, and it asks for one action at a time. Lifts that slice into its own small actions.world_delta column (migration 36) and marks context_snapshot, state_before and world_state_before deferred, so they load only when something touches the attribute -- Insights, undo and retry, all single-action paths. The backfill runs server-side, dialect-specific (json_extract on SQLite, #> on Postgres), because pulling 39 MB of snapshots into Python to rewrite a slice of each would defeat the purpose. Measured at production shape (541 actions, 72 KB snapshots), one adventure load goes from 38.46 MB to 0.20 MB. The traffic that consumed 5 GB would now be about 27 MB. Deliberately not included: limiting the history query to recent actions, and removing the redundant db.refresh(adventure) calls. Both were sized against the old numbers; against a 0.20 MB load they would take ~27 MB a month down to ~10 MB, which is not worth the complexity. tests/test_egress.py hooks before_cursor_execute and asserts the emitted SQL never names the deferred columns during a bulk load, so this cannot regress silently. 123 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |
||
|
|
7c538c8235 |
Stop retried and deleted actions from corrupting story context and memories
Three fallout bugs from keeping the retried action row alive (
|
||
|
|
970d71a5b6 |
Keep world-state emit reliable across turns
The AI would sometimes stop emitting the `state` delta block once it missed a turn. Two compounding causes: the emit rule sat only in the system block (far from where the model generates), and the block was stripped before storage — so every replayed history turn looked blockless, biasing the model by imitation to stop emitting too. - EMIT_REMINDER: a one-line reminder appended last in the prompt (strongest recency slot), gated on has_ws and counted against the token budget. - render_delta_block + _history_text: re-attach each past AI turn's own delta block in replayed history (reconstructed from the stored snapshot delta), so the model always sees its emit format. Action.text stays clean, so UI, embeddings, and card/NPC trigger-matching are unaffected. History budgeting counts the augmented text so it can't overflow. Undo/retry untouched (read the separate world_state_before column). 46 tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UWVyFKvqJGjfbXdibLgkMe |
||
|
|
cf464a48b0 |
Make NPCs a dedicated section with per-NPC stats
Replace the single shared `npc` stat template (+ npc_card_types) with an `npcs` section: each NPC keyed by a stable id, carrying its own name, description, trigger keys, and its OWN stats block. The AI addresses NPCs as npc.<id>.<stat> (id shown in context), which also fixes the old card-id-guessing problem. On adventure creation each NPC auto-creates a story card (name/keys/desc) for lore + in-scene detection, unless a same-name card already exists. All NPCs instantiate up front. - engine: npcs instantiate/apply/render/reference, npc_name/npc_triggers - builder: _visible_npcs matches each NPC's own keys - create_adventure: auto-create story cards from npcs - WorldStateDrawer: render defined NPCs with their own stats + desc tooltip - demo seed: Gwen (health/trust) + Bandit Leader (health/aggression) - tests updated (34 pass); no new migration (npcs lives in stat_schema JSON) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4dac445f90 |
Add Phase 12: native RPG world state
Structured world/player/NPC stats, two-way flags, and sticky milestones per scenario (stat_schema). The AI proposes a per-turn delta; a Python engine referees it (clamp to min/max, per-turn cap, cooldown, counters). Band word-labels plus a fixed stat guide (descriptions + full ranges) keep the model grounded. World State drawer + Insights delta report; undo/retry roll it back via the Phase 11 snapshot pattern. - migrations 26-28 (scenarios.stat_schema, adventures.world_state, actions.world_state_before); all nullable, additive, safe on existing rows - migration 29 raises the default context budget 4096 -> 16384 (custom values preserved) - seeded demo scenario 04-rpg-world-state.json (Bandit Camp) - 19 new tests (33 total pass) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
db9f904222 |
Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |