144406cd481df567c4d96131002d907785045e9e
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1ce9972760 |
M8: the browser becomes the storyteller
The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.
All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.
So the shape now is one entry point and one screen:
Campaigns -> Campaign -> Story
State · Knowledge · Context · Save Points · Settings
Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.
Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.
Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.
The two defects worth the space:
A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.
And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.
Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.
`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.
Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.
Backend, and only what the browser could not otherwise reach:
AdventureCreate.opening a start action could only come from a Scenario, so
every campaign made in the new setup flow opened on
a blank page. Same node, same code path.
canon_rules campaign_canon has been the highest authority in a
campaign since M5, read by the prompt builder and
the state validator, and had no API at all — a
fixture had to write it with SQL.
a 401 and a 429 message the last user-facing text describing a hosted
deployment. One told the reader to check an API key
that has not existed since M2.
No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.
The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.
They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.
A verification pass over all of it then found three more, each by driving the
product rather than reading it:
Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.
The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.
And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.
Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.
`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.
Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:
The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.
Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.
The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.
The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.
Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
|
||
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
a408c7b6f7 |
Lay the prompt out so the endpoint can cache most of it
Prompt caching bills on a shared prefix: the endpoint reuses the request up to the first byte that differs from last time and no further. The live world-state block sat third from the top of the system message, so every turn re-priced the instructions, the plot essentials and the whole story history underneath it. The retrieved memories and the rewritten summary did it again. Everything fixed is emitted first now, and everything that moves goes after the history, ordered least-volatile first — which is also where recency serves it best, the reasoning that already put the emit reminder last. The three tail sections that are last for their own reasons stay last. The moved sections are still charged to the token budget; only their position changed. Two smaller halves of the same problem. OpenRouter serves a model from whichever upstream is free and each upstream holds its own cache, so a deepseek model now names deepseek as its preferred upstream — a preference, not a restriction, so a turn still runs if that upstream is down. And the endpoint's usage block is read back off the response and kept per attempt, so the hit rate shows up in Insights and the debug log instead of being assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
c500203270 |
Harden auth against a forwarded-header rate-limit bypass, and guard BYOK SSRF
The per-IP rate limits could be bypassed entirely: uvicorn ran with --forwarded-allow-ips "*", which trusts the leftmost X-Forwarded-For value (client-controlled), and Render forwards the inbound header rather than stripping it. Rotating the header handed out a fresh rate-limit bucket per request, so the login/register limit (10/5min) and guest-minting limit (30/5min) were no throttle at all — unbounded password guessing and guest-row creation. Confirmed live: fixed IP -> 429 after 10; rotating spoofed header -> no 429 across 14 attempts. Two-layer fix: - limits._client_ip now derives the client IP from the hop the trusted edge appends (rightmost of X-Forwarded-For), which a client can't spoof past; tunable via AIDND_TRUSTED_PROXY_HOPS. Dropped --forwarded-allow-ips "*". - New per-account login throttle (email-keyed, 8 fails / 15 min, cleared on success): stops distributed guessing against one account that a per-IP limit can't, since it can't be diluted across many source addresses. Also close an SSRF on the BYOK endpoint_url (hosted mode only): the connection test and turn/chat streams now refuse a URL that resolves to a non-public address (private/loopback/link-local metadata/reserved), checked at request time so it resists a DNS record flipping to a private IP. No-op locally, where reaching localhost Ollama is intended. Tests: test_ratelimit_hardening.py (8), test_netguard.py (13). 172 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7 |
||
|
|
d66fd1d6a2 |
Add a reasoning-off setting (-1 reasoning budget)
Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.
Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.
Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
|
||
|
|
1dd31086c1 |
Add power-user AI Chat page; centralize demo-key model pinning
AI Chat is a plain scratchpad for talking to a model directly — no story context, scripts or world state — for poking at models, prompts and endpoints without starting an adventure. Power users only: the router 404s (rather than 403s) for everyone else and the nav link is hidden. The conversation lives in localStorage, so there's no new table or migration. is_power_user() now also returns True in local mode: it's the operator's own machine and their own key, the same reasoning that makes the provider debug log local-only. Alongside that, the rule keeping the shared demo key off paid models now lives in exactly one place. It had been duplicated into the chat router, which is how one copy eventually drifts: - resolve_provider_config() takes an optional model_override and is the only place the whitelist is applied, so turns, AI Chat and the connection test all inherit it. An override is a per-request preference, never a grant. - ProviderConfig.__post_init__ refuses to exist when api_key is the demo key and the model isn't whitelisted. It keys on the key itself rather than the using_demo flag, so a mislabelled config can't slip past, and it raises so a future path that bypasses the resolver fails loudly instead of billing. - The demo branch still pins endpoint_url too — a user-controlled endpoint would leak the key itself, which is worse than spending it. Provider gained chat(messages, ...) beside generate(), both delegating to a shared _stream(url, body); completion-mode endpoints get the messages flattened into a labelled transcript. Settings' /models fetch moved to list_endpoint_models() and is shared with /api/chat/config. Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub resolve_provider_config — the point is to exercise the real BYOK-vs-demo decision and assert on what the provider actually received: off-whitelist override pinned, off-whitelist Settings.model pinned, redirected endpoint pinned, BYOK passed through untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx |
||
|
|
2953cd70ed |
Show a friendly message on 429 instead of raw JSON
Map HTTP 429 in _friendly_http_error: OpenRouter's free-models-per-day cap gets a "daily limit, resets 00:00 UTC" message; other rate limits get a generic "too many requests, try again" note. Avoids leaking the raw error JSON (incl. user_id) to players. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5 |
||
|
|
253b533d3b |
Fix remaining code-review findings (15 bugs; #11 skipped as AID-compatible)
Backend: - provider: fall back to parsing a plain JSON body when a server ignores stream=true (was: silent empty turn); error if response has no text (#5) - memorybank: clamp cursors after undo/retry shrinks the action list, and translate summary_cursor (list position) to an Action.index boundary before comparing with Memory.source_end (#7) - memorybank: pinned memories now count toward the top_k budget (#8) - memorybank: cosine() returns 0.0 on dimension mismatch; changing the embedding model clears stored vectors so they re-embed (#9) - scripting: MAX_STORY_CARDS cap now counts cards inserted during the hook, so a script can't add unbounded cards in one turn (#10) - settings: /test tolerates non-dict JSON from /models (#13) - scenarios: import accepts worldInformation as a story-card source (#14) Frontend: - per-key debounce timers in PlotPanel and ScenarioEditor — editing two things within 600ms no longer drops the first save (#15, #16) - Continue button no longer discards typed input (#17) - failed retry resyncs actions from the server instead of leaving the removed action missing (#18) - Settings save/test surface errors instead of hanging on Testing… (#19) - InsightsPanel ignores stale responses from superseded requests (#20) - placeholder scan includes story-card trigger keys (#21) addStoryCard returning the 0-based index (falsy for the first card) matches real AI Dungeon per the scripting guidebook — kept, documented (#11). Statuses updated in CODE_REVIEW_FINDINGS.md; stale entries for previously fixed items (#1-4, #6, #12) corrected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |
||
|
|
db9f904222 |
Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |