62a997f364e5387e4cc1dcc20f5a6618dee5f267
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
2c5909a268 |
Drop the JSON vector column, and fix what was hiding behind it
Migration 38 left memories.embedding in place so a rollback could still find the vectors. Production has since been verified reading from embedding_blob, so migration 42 drops it: 4 MB of a 99.6 MB database holding nothing anyone reads. Removing it surfaced a live bug. Changing your embedding model is supposed to throw the bank's vectors away and let the post-turn pass rebuild them, because two models' vectors are not comparable. The settings route did that by nulling memories.embedding -- correct until 38 moved the vectors, after which it cleared the dead column and left the blob intact with `embedded` still true. _embed_pending filters on `embedded IS FALSE`, so it never saw those rows and the bank went on ranking against the old model's vectors permanently. Nothing would have reported it. cosine returns 0.0 on a width mismatch, so a different-width model scores every memory zero and retrieval returns whichever rows happen to sort first; a same-width model scores plausible garbage. The bulk clear now sets both columns. It stays a bulk UPDATE rather than going through set_vector -- loading the rows is the cost that whole path exists to avoid -- so set_vector's docstring now names it as the one caller that legitimately writes those columns by hand. No cache invalidation is added: clearing `embedded` drops the rows out of the catalogue query, and set_vector evicts each entry as the re-embed puts it back. test_embedding_blob.py now rebuilds the pre-38 schema by hand where it tests the backfill, since create_all no longer produces the column it converts from, and asserts 42 removes it at the end of a full bootstrap -- 38 reads that column and 42 drops it, so an upgrade that reordered them would arrive with an empty bank. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7 |
||
|
|
c500203270 |
Harden auth against a forwarded-header rate-limit bypass, and guard BYOK SSRF
The per-IP rate limits could be bypassed entirely: uvicorn ran with --forwarded-allow-ips "*", which trusts the leftmost X-Forwarded-For value (client-controlled), and Render forwards the inbound header rather than stripping it. Rotating the header handed out a fresh rate-limit bucket per request, so the login/register limit (10/5min) and guest-minting limit (30/5min) were no throttle at all — unbounded password guessing and guest-row creation. Confirmed live: fixed IP -> 429 after 10; rotating spoofed header -> no 429 across 14 attempts. Two-layer fix: - limits._client_ip now derives the client IP from the hop the trusted edge appends (rightmost of X-Forwarded-For), which a client can't spoof past; tunable via AIDND_TRUSTED_PROXY_HOPS. Dropped --forwarded-allow-ips "*". - New per-account login throttle (email-keyed, 8 fails / 15 min, cleared on success): stops distributed guessing against one account that a per-IP limit can't, since it can't be diluted across many source addresses. Also close an SSRF on the BYOK endpoint_url (hosted mode only): the connection test and turn/chat streams now refuse a URL that resolves to a non-public address (private/loopback/link-local metadata/reserved), checked at request time so it resists a DNS record flipping to a private IP. No-op locally, where reaching localhost Ollama is intended. Tests: test_ratelimit_hardening.py (8), test_netguard.py (13). 172 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7 |
||
|
|
1dd31086c1 |
Add power-user AI Chat page; centralize demo-key model pinning
AI Chat is a plain scratchpad for talking to a model directly — no story context, scripts or world state — for poking at models, prompts and endpoints without starting an adventure. Power users only: the router 404s (rather than 403s) for everyone else and the nav link is hidden. The conversation lives in localStorage, so there's no new table or migration. is_power_user() now also returns True in local mode: it's the operator's own machine and their own key, the same reasoning that makes the provider debug log local-only. Alongside that, the rule keeping the shared demo key off paid models now lives in exactly one place. It had been duplicated into the chat router, which is how one copy eventually drifts: - resolve_provider_config() takes an optional model_override and is the only place the whitelist is applied, so turns, AI Chat and the connection test all inherit it. An override is a per-request preference, never a grant. - ProviderConfig.__post_init__ refuses to exist when api_key is the demo key and the model isn't whitelisted. It keys on the key itself rather than the using_demo flag, so a mislabelled config can't slip past, and it raises so a future path that bypasses the resolver fails loudly instead of billing. - The demo branch still pins endpoint_url too — a user-controlled endpoint would leak the key itself, which is worse than spending it. Provider gained chat(messages, ...) beside generate(), both delegating to a shared _stream(url, body); completion-mode endpoints get the messages flattened into a labelled transcript. Settings' /models fetch moved to list_endpoint_models() and is shared with /api/chat/config. Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub resolve_provider_config — the point is to exercise the real BYOK-vs-demo decision and assert on what the provider actually received: off-whitelist override pinned, off-whitelist Settings.model pinned, redirected endpoint pinned, BYOK passed through untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx |
||
|
|
4772171b6c |
Phase 9: production hardening
Config via env, abuse/resource limits, and production serving so the app is safe to expose publicly: - Fail-fast on missing SECRET_KEY when MULTI_USER=true - quickjs per-execution time/memory limits (while(true) can't hang server) - Per-user/per-IP rate limiting on turn/script/auth endpoints - Request body size limit + per-user row caps - Security headers (CSP, X-Frame-Options, nosniff, referrer-policy) incl. SSE - Debug router 403 and /docs disabled in multi-user mode - DATABASE_URL support (defaults to Neon Postgres) alongside SQLite - Documented all env vars in backend/.env.example Verified locally via uvicorn (MULTI_USER=1, SQLite); see plan/09-phase-hardening.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X |
||
|
|
de4db373f2 |
Phase 8: optional accounts, per-user data, shared demo key
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me, register upgrades the guest in place, login/logout, per-IP rate limits. Every router scoped by user_id; Settings become per-user with the API key Fernet-encrypted at rest and write-only through the API. Users without a key get a server-funded demo key (OpenRouter free models, 20 turns/day, memory bank disabled on demo turns). Public read-only demo scenarios (seed_demo.py); debug log restricted to local mode. Frontend: auth modal + guest nudge, 401 re-establish/retry, demo banner and key management in Settings. Migrations 13-23 adopt existing data under a local user and encrypt stored keys. Verified: migration on a copy of real data.db, two-session isolation + register/login via curl and Chrome, demo cap 429, live OpenRouter turn through the encrypted-key path, vite build + oxlint. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |
||
|
|
253b533d3b |
Fix remaining code-review findings (15 bugs; #11 skipped as AID-compatible)
Backend: - provider: fall back to parsing a plain JSON body when a server ignores stream=true (was: silent empty turn); error if response has no text (#5) - memorybank: clamp cursors after undo/retry shrinks the action list, and translate summary_cursor (list position) to an Action.index boundary before comparing with Memory.source_end (#7) - memorybank: pinned memories now count toward the top_k budget (#8) - memorybank: cosine() returns 0.0 on dimension mismatch; changing the embedding model clears stored vectors so they re-embed (#9) - scripting: MAX_STORY_CARDS cap now counts cards inserted during the hook, so a script can't add unbounded cards in one turn (#10) - settings: /test tolerates non-dict JSON from /models (#13) - scenarios: import accepts worldInformation as a story-card source (#14) Frontend: - per-key debounce timers in PlotPanel and ScenarioEditor — editing two things within 600ms no longer drops the first save (#15, #16) - Continue button no longer discards typed input (#17) - failed retry resyncs actions from the server instead of leaving the removed action missing (#18) - Settings save/test surface errors instead of hanging on Testing… (#19) - InsightsPanel ignores stale responses from superseded requests (#20) - placeholder scan includes story-card trigger keys (#21) addStoryCard returning the 0-based index (falsy for the first card) matches real AI Dungeon per the scripting guidebook — kept, documented (#11). Statuses updated in CODE_REVIEW_FINDINGS.md; stale entries for previously fixed items (#1-4, #6, #12) corrected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |
||
|
|
db9f904222 |
Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg |