9 Commits
Author SHA1 Message Date
JesseMarkowitz 8c65ae99de M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The
milestone is subtraction, and what is left is the single-user local
storyteller the specification describes.

Removed in full: campaign scripting and its QuickJS sandbox; multi-user
accounts, guest sessions, login, registration and the shared demo key;
the visitor-analytics tables, dashboard and page beacon; the access log
of sign-ins, addresses and devices; per-IP and per-user rate limiting
and quotas; Render deployment config; Postgres and psycopg; cloud
inference providers, the API-key field and the key encryption that
existed to store it; session-cookie signing. None of it was hidden
behind a flag — the routes are gone and answer 404.

Two things were kept that the brief allowed keeping. The `users` table
and its foreign keys stay as an internal ownership detail, because
rewriting them out means a migration across most of the schema to
delete a column that costs nothing; nothing creates a second user and
no request carries an identity. Five inert tables and four inert
columns stay for the same reason, so an M1 campaign database opens
unchanged.

The one addition is app/endpoints.py, which decides where a story may
be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an
explicit allowlist of networks, not a guess at what `ipaddress` means
by "private", which calls the documentation ranges private and IPv6
loopback reserved. Every address a hostname resolves to must be in it,
so a split answer does not squeak through, and the rule runs both when
the endpoint is saved and before every outbound request, because a name
that resolved to the LAN this morning can resolve elsewhere this
afternoon. Known cloud hosts are named in the refusal so the error says
why rather than looking like broken DNS. TLS is never traded against
it: M1's shared trust context is intact on all four clients and there
is no way to skip verification.

The hardcoded 120-second model timeout is now a setting. That was not
theoretical — on this GPU-less four-core host a cold load of
qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while
turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short
at 10s so a wrong address still fails fast; the read timeout defaults
to 300s and is bounded at 3600, because "wait longer" must stay a
number.

Two defects found while testing and fixed here. An unknown /api path
fell through the SPA catch-all and came back as HTML with status 200,
so a client asking for JSON parsed a web page instead of learning the
route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an
unauthenticated loopback API would hand every page on the Internet a
write handle on the campaign database; it now refuses to start.

Verified rather than assumed. Offline, on a network with no route out
and no DNS: five turns, retry with both takes retained, restart with an
identical transcript digest, a failed model call leaving the accepted
AI-turn count untouched, and a capture with zero non-loopback unicast
packets. Against a real second machine on the LAN over HTTPS with a
private CA: four turns, restart, and a capture showing 289 packets to
the approved host, 344 loopback, zero anywhere else, zero DNS queries.
Cloud and public endpoints refused with their reasons; no API key
settable; every removed route 404.

604 backend tests pass, down from 648 by the fifteen retired with the
subsystems they tested and up by the twenty-nine added for the endpoint
policy and the removed surface. The scripting tests were not deleted:
eight files used a JavaScript counter as instrumentation for the state
snapshot and rollback machinery, which M2 does not touch, so the
counter moved to the world-state engine and those tests still assert
what they always did. Frontend lint and build are clean; the image
builds, and its wheel-building stage is gone with quickjs.

No M3 work. Undo is still destructive and there is still no Redo.
2026-09-02 11:27:14 -04:00
parththakkar106andClaude Opus 5 b1772c6e21 Plan the readability refactor, and clear the tree for it
Phase 17 splits the four files that hold most of the code, finishes the
schema migration SP8 left half done, and stops the published guide from
drifting away from its Markdown source. `plan/17-refactor.md` carries the
plan and the progress table, and `plan/STATUS.md` points at it.

Stage 0 is hygiene only. Both abandoned worktrees are gone, which freed
about 104 MB. Removing `sp7-tree-ui` needed one extra step: a Vite dev
server had been running out of it since 2026-08-18, holding
`frontend/.vite` open and owning port 5173, and serving a tree 54 commits
behind `main`. The three stale `.db` files are deleted; `data.db` is not.

`AIDND_TRUSTED_PROXY_HOPS` is now documented. It was read at `limits.py:55`
and named in no `.env.example`, README, or blueprint. It sets how many
proxy hops the rate limiter trusts in `X-Forwarded-For`, so a deployment
that adds a hop without setting it gets the bucket-rotation bypass back.

The 19 squash-landed branches are still there. `git branch -D` is blocked
by the permission classifier; the verified command is in the plan file.

549 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 00:36:21 +05:30
parththakkar106andClaude Opus 5 041f9e25f3 Count the visits, and say whether anyone got anywhere
A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.

Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.

**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.

**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.

Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.

Three things came out of building it that a test would not have suggested.

**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.

**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.

**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.

Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.

No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.

497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-22 16:24:42 +05:30
parththakkar106andClaude Opus 5 bbcb07c6be Delete guest accounts left idle for five days
The app had no cleanup of any kind: in multi-user mode every first visit
mints a users row, so the demo has been accumulating one permanent account
per visitor along with everything they generated.

cleanup.py sweeps guests idle for AIDND_GUEST_RETENTION_DAYS (default 5),
once at startup and then every few hours. Startup is the load-bearing
trigger — the free tier sleeps after ~15 minutes, so a long timer rarely
gets to fire.

Idle is COALESCE(last_seen_at, created_at), not last_seen_at: _touch only
writes that column hourly, and a guest minted by /auth/me has it NULL until
its second request, so the simpler query would have deleted brand-new
visitors mid-session.

It's one Core DELETE rather than db.delete(user), which would SELECT every
adventure, action and memory into Python purely to delete them — the same
egress pattern as the 189x fix. Every FK from users down is ON DELETE
CASCADE, so the database does the whole graph and returns a count.

The filter requires is_guest AND email IS NULL, so registered users (who
upgrade in place) and local mode's implicit user are both out of reach, and
is_public is output-only so a guest can never own content another user can
see. Session cookies have no expiry and can outlive a swept row; that path
401s and the frontend's existing retry re-mints a session.

Guests are told: /auth/me serves guest_retention_days and the signup modal
states the window, sourced from the server so it can't drift from what is
enforced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CYEJKobJ2Re4Dv7qUoSA7
2026-08-15 20:14:12 +05:30
parththakkar106andClaude Opus 5 1dd31086c1 Add power-user AI Chat page; centralize demo-key model pinning
AI Chat is a plain scratchpad for talking to a model directly — no story
context, scripts or world state — for poking at models, prompts and endpoints
without starting an adventure. Power users only: the router 404s (rather than
403s) for everyone else and the nav link is hidden. The conversation lives in
localStorage, so there's no new table or migration.

is_power_user() now also returns True in local mode: it's the operator's own
machine and their own key, the same reasoning that makes the provider debug log
local-only.

Alongside that, the rule keeping the shared demo key off paid models now lives
in exactly one place. It had been duplicated into the chat router, which is how
one copy eventually drifts:

- resolve_provider_config() takes an optional model_override and is the only
  place the whitelist is applied, so turns, AI Chat and the connection test all
  inherit it. An override is a per-request preference, never a grant.
- ProviderConfig.__post_init__ refuses to exist when api_key is the demo key
  and the model isn't whitelisted. It keys on the key itself rather than the
  using_demo flag, so a mislabelled config can't slip past, and it raises so a
  future path that bypasses the resolver fails loudly instead of billing.
- The demo branch still pins endpoint_url too — a user-controlled endpoint
  would leak the key itself, which is worse than spending it.

Provider gained chat(messages, ...) beside generate(), both delegating to a
shared _stream(url, body); completion-mode endpoints get the messages flattened
into a labelled transcript. Settings' /models fetch moved to
list_endpoint_models() and is shared with /api/chat/config.

Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub
resolve_provider_config — the point is to exercise the real BYOK-vs-demo
decision and assert on what the provider actually received: off-whitelist
override pinned, off-whitelist Settings.model pinned, redirected endpoint
pinned, BYOK passed through untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
2026-07-26 20:16:56 +05:30
parththakkar106andClaude Opus 4.8 94db6e055b Add power-user email allowlist to bypass demo turn cap
Trusted testers listed in AIDND_POWER_USERS get unmetered turns on the
shared demo key: demo_turns_left reports the full cap and count_demo_turn
skips them. Registered accounts only, matched case-insensitively.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
2026-07-21 00:21:30 +05:30
parththakkar106andClaude Opus 4.8 4772171b6c Phase 9: production hardening
Config via env, abuse/resource limits, and production serving so the app
is safe to expose publicly:

- Fail-fast on missing SECRET_KEY when MULTI_USER=true
- quickjs per-execution time/memory limits (while(true) can't hang server)
- Per-user/per-IP rate limiting on turn/script/auth endpoints
- Request body size limit + per-user row caps
- Security headers (CSP, X-Frame-Options, nosniff, referrer-policy) incl. SSE
- Debug router 403 and /docs disabled in multi-user mode
- DATABASE_URL support (defaults to Neon Postgres) alongside SQLite
- Documented all env vars in backend/.env.example

Verified locally via uvicorn (MULTI_USER=1, SQLite); see plan/09-phase-hardening.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
2026-07-07 12:30:06 +05:30
parththakkar106andClaude Fable 5 de4db373f2 Phase 8: optional accounts, per-user data, shared demo key
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.

Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 23:04:03 +05:30
parththakkar106andClaude Fable 5 22630c8411 Phase 7: MIT license, portfolio README, Docker, cross-platform start
- LICENSE: MIT
- README: rewritten as portfolio-grade docs (features with code pointers,
  architecture, Docker/Windows/macOS quick starts, BYO-model table)
- Dockerfile (3-stage: SPA build, pip wheels, slim runtime) + compose with
  a /data volume; .dockerignore keeps secrets and local data out
- database.py: AIDND_DB_PATH env override so deployments can relocate the
  SQLite file; documented in backend/.env.example
- start.sh: macOS/Linux dev script with first-run setup

Verified locally: production SPA build served by the backend (deep links
OK), DB created at the override path. Docker image itself untested here
(Docker not installed); flagged in plan/07.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:51:48 +05:30