A campaign can import local .txt and .md files as Canon, Reference or
Inspiration, and the class is load-bearing rather than a label: it decides the
words a passage is framed with in the prompt, the weight it carries when
passages are ranked, and which budget it competes in when the context is tight.
This is a separate subsystem, which is the Phase 0B decision
(IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification,
provenance, content identity, chunking, an index or a lifecycle, and they were
not promoted into something that does. Nothing here reads or writes one.
The subsystem, in backend/app/knowledge/:
classes the three classes, their weights, and the prompt framing
chunking deterministic, heading-aware, 60-800 tokens, no overlap
fts SQLite FTS5 with porter stemming; scoped and bounded in SQL
importer validate, hash, store, chunk, index — in one transaction
embeddings local Ollama vectors through the shared provider
retrieval query construction, hybrid merge, rerank
inject the budgeted cut and the rendered prompt sections
Relevance admission is a separate stage from ranking, and that separation is
the milestone's most expensive lesson. An independent review found the first
implementation deciding relevance with a floor expressed as a share of the best
candidate — which the best clears by construction — so a passage was admitted on
every turn regardless of the scene. A query about tide tables and container
tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden
Canon among them.
So the pipeline is now:
candidate generation -> admission -> ranking -> class weighting -> budget
Admission reads raw, candidate-set-independent signals: the cosine the model
returned, and how many distinct meaningful query terms a passage contains.
Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero
is not zero. Normalization decides order among things that matched; it can never
decide whether anything matched. Authority is applied after admission, so a
class orders what matched and never rescues what did not.
Retrieval may therefore return nothing, and on a scene unrelated to the library
it does.
The other decisions that each replaced an obvious wrong one:
- The class multiplies relevance rather than adding to it. An additive bonus
satisfies "Canon outranks Reference" and makes "do not include irrelevant
Canon" impossible, because a large enough constant wins on its own.
- The semantic floor is measured, not guessed: 113 production-path pairs against
nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at
0.36-0.56, and 0.58 sits between them. Because it is a property of that model
and not of cosine similarity, it is keyed to the model rather than applied to
whatever is configured: an embedding model with no measured calibration in
this build does not borrow the number. Semantic admission is skipped, the
campaign retrieves lexically, and the reason is stated in the knowledge status
and in the turn's provenance. Degrading to lexical keeps the library usable;
lending the threshold to an unmeasured model is how the admitted-everything
defect would return.
- One lexical term is not evidence. Two distinct meaningful terms, or one that
is neither a standing campaign entity nor a negligible share of the query.
The stop list grew from 42 words to 261, all function words — no subject
matter, because a stop list that removes subject matter stops finding "The
Silver Key".
- Lexical retrieval is a production path, not a fallback. It finds the proper
nouns and invented terms a setting bible is made of, and the library is fully
usable with no embedding model configured.
Safety is structural rather than filtered. Imported text reaches the prompt
whole, inside a section that says what it is, under a rule stating the authority
order in words and refusing every instruction inside it. No endpoint accepts a
filesystem path, so H08 has no mechanism to escape from. Nothing renders
imported content as HTML, so a script tag is five visible characters and a
remote image is never fetched. Import, chunking, indexing, retrieval and a turn
open no socket at all; only embeddings do, through the endpoint allowlist the
memory bank already uses.
Provenance is the rendered text, not a foreign key: deleting a source cannot
turn a historical turn's evidence into dangling ids.
Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5
virtual table attached to knowledge_chunks as a DDL hook so it is created and
dropped with the table it indexes. Migration 92. A pre-M7 database opens
unchanged and needs no sources to play.
Bundle: the source content and the reader's judgements about it travel; the
passages, index rows and vectors are rebuilt on import, so a restored campaign
is searchable immediately without a reindex step.
One runtime dependency: python-multipart, Starlette's multipart parser. It is
what makes the upload surface possible, and the upload surface is why no
pathname is ever accepted.
The test doubles were the reason the defect shipped, so they were corrected too.
The retrieval stub scored unrelated text at 0.06-0.20 where the real model
scores it at 0.43-0.44, and its docstring said it had deliberately removed the
constant component that "would put a similarity floor under every pair" — which
is exactly the property real models have. The stub now has that floor, one test
fails if it is ever removed, and another reproduces the superseded rule and
asserts it is still fooled by the same fixture. Run against the pre-corrective
implementation, the new suite fails 13 of 18.
Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of
which mocks nothing between itself and Ollama and re-measures the similarity
separation on every run. 43/43 checks in a real Firefox, reproduced.
Docker build clean.
Four other defects found by review or by the browser run were fixed here rather
than carried: an unreachable relevance constant that appeared to enforce
something and did not; acceptance tests using the wrong fixture files, so G07's
trap was never exercised; a bidirectional override surviving into displayed
filenames; and, from the implementation pass, the Insights panel showing M5's
two state sections as raw keys and the source inspector refetching on every
keystroke.
M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK
REQUIRED. Both blocking findings are closed, and closeout resolved the
embedding-model calibration boundary the corrective pass had left as debt.
planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective
closeout and the closeout verification in sequence, none overwriting another.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
Phase 0B ran the upstream application on a network with no route out and
the first turn died in tiktoken, which downloads its BPE table the first
time anything counts a token. The browser separately fetched three font
families from Google on every page load. Neither is visible on a machine
that has been online once, which is why both now have tests.
The tokenizer table is vendored at
backend/app/context/vendor/cl100k_base.tiktoken and
backend/app/context/encoding.py builds the encoding from it directly,
verifying its SHA-256 against the digest tiktoken itself pins for that
URL. No code path in the tokenizer can reach the network any more —
not a warm cache, not an environment variable a deployment could forget.
The encoding was checked token for token against tiktoken's own.
The three font families are self-hosted as variable fonts under
frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in
frontend/src/styles/fonts.css, and re-vendored by
frontend/tools/vendor_fonts.py. Their OFL licences ship beside them.
With no remote asset left, the CSP drops both Google hosts and gains
object-src, base-uri and form-action; woff2 also gets its real media
type, which Python's table lacks on a slim image.
A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx
verifies against the certifi bundle, so an endpoint whose certificate
comes from a CA the user installed on their own machines — a StartOS
server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED
while curl and the browser on the same host accepted it.
app/tlstrust.py builds one context that unions the platform CA store
with certifi's, and all four outbound clients use it. A union rather
than a swap, so an image with an empty system store cannot start failing
on endpoints that worked before. Verification itself is untouched:
CERT_REQUIRED, hostname checking on, and no insecure escape hatch.
The storyteller listener is now loopback by explicit statement rather
than by inheriting uvicorn's default: start.sh, start.ps1, and
docker-compose.yml, which publishes to 127.0.0.1 rather than every
interface. Reaching an Ollama on another machine is outbound and needs
none of that inbound exposure.
backend/requirements.lock pins the exact tested closure;
requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the
same-host and trusted-LAN Ollama configurations, and how to re-run the
offline proof. PROVENANCE.md records the upstream commit, the MIT terms,
and both vendored assets.
Verified, not just compiled. On an --internal Docker network with
1.1.1.1 unreachable and no name resolving, a campaign was created and
played for six turns through same-host Ollama, restarted, and resumed.
A second run played ten turns through Ollama on a separate physical
machine on the LAN over verified HTTPS, summaries and embeddings
included, with the storyteller's default route deleted so the LAN was
reachable and the Internet was not. Its capture: 893 packets to the
approved host, 730 loopback, zero anywhere else, and zero DNS queries.
Two induced model failures left the accepted story bit-identical. The
inherited SPA was opened in a browser and a campaign read back from it.
Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the
findings that did not belong in this change.
648 backend tests pass, up from the inherited 632; frontend lint and
build are clean; the image builds. No M2 work is included: the hosted,
cloud, analytics, Postgres and scripting surfaces are untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL
`frontend/src/index.css` held 2865 lines. It is now an import list, and each
section is its own file under `frontend/src/styles/`.
The order is unchanged. Two comments in the file already recorded that the
cascade depends on it: several selectors in the tome section win only because
they come after the base card rules, and the 720px query overrides the whole
desktop design. The built CSS bundle is byte-identical before and after, at
56686 bytes.
The plan asked for each `@media` block to move next to the rules it overrides.
That is not done, and should not be. Moving a rule past a later rule of equal
specificity changes which one wins, and there is no frontend test that would
catch it. The 720px block is now `responsive.css`, imported last.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
`apply_delta` records three outcomes for every change the model sends:
`applied`, `clamped`, and `rejected`. Everything downstream read only
`applied`. A refused change reached the player as an ordinary chip, and
reached the model on the next turn as a change that had succeeded.
Five parts:
- `Action.world_changes` reads `clamped` and `rejected` beside `applied`.
Accepted stats carry a `clamped` flag; refusals become `kind: "rejected"`
entries. The `fix` key is present only when the engine wrote one, because
this property runs for every action of every list response.
- The UI separates the three outcomes. A clamp to a standstill reads
`no change - at its limit` on a dashed chip, a partial clamp is marked
`(limited)`, and a rejection carries its reason. Dashed and dimmed rather
than red: a refused change means the rules are working.
- The goals line names the milestone id, as `milestones.<id>`. The ids
appeared nowhere in the prompt before, so the model could not send one.
- Each rejection, and each clamp that moved nothing, builds a `fix` string
from the stat definition at the point of refusal. `render_refusals()`
renders them into the next prompt above `EMIT_REMINDER`.
- `_history_text` replays `applied_delta()` instead of the sent delta, so a
past turn's state block shows only what the engine accepted.
A clamp that reduced a change but still moved the value reports nothing. If
you tell a model its 80 damage became 30, it can treat the shortfall as a
debt and send the remaining 50 next turn, which is the swing
`max_delta_per_turn` prevents.
In the demo scenario, `pokemon_left` becomes `pokemon_fainted`
(`type: counter`, `initial: 0`). Starting at the ceiling turned a wrong-signed
delta into a silent no-op; counting up puts the wrong sign on the counter
rule, which refuses it out loud. The instructions also now ask for
`world.turn`, which sat at 0 for a whole playtest.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
A change chip like "bandit leader aggression +10" was nowrap inside the
story column, which has nothing to scroll, so it overran narrow screens.
The label may now break as a last resort while the value stays glued
together, and the analytics range buttons wrap onto a second row.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.
Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.
**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.
**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.
Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.
Three things came out of building it that a test would not have suggested.
**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.
**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.
**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.
Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.
No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.
497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
The Branches panel says which lines exist. It cannot say where they parted
or how much story each one is, because those are the two numbers `GET
/branches` already answers and a list has nowhere to put them. A ⌗ See the
tree button opens a map: one horizontal lane per branch, running from the
moment it left its parent to the moment it ends, joined to the parent by an
elbow at the fork. The horizontal axis is the story's own clock, so two
lanes at the same x are at the same moment and a short branch reads as
short.
This is a branch map, not the per-node map SP7 refused, and that is the
whole reason it was cheap. Lanes are bounded by branch count, not node
count, so the 600-node windowing problem never arrives. It reads the single
request the rail already made and nothing else.
`branches.js` holds the tree maths and `BranchMap.jsx` the drawing. Play.jsx
gives up its private copies of branchLabel and orderBranches: the panel and
the map now label and order a branch through the same functions, so a branch
cannot be called two things by the two views. The three operations stay in
BranchPanel and are passed down, and `run` answers whether it worked so
neither view clears a half-typed name on a refusal.
Three things came out of driving it, none of which a test could have seen.
`clientWidth` counts the canvas padding the ResizeObserver leaves out, so
the first paint drew an svg 24px wider than its box — and the observer's
initial observation never arrived here, so dropping the seed left the map
never drawing at all. Both are needed and the seed subtracts the padding.
Only the name was being clipped, not the meta line under it, so a late fork
ran its text off the right edge; both are clipped now, and a lane starting
in the right third hangs its labels back over the fork, where its own band
guarantees nothing to collide with. And the delete rule lived in the server
and in the map but not in the list, which offered Delete on a branch the
head was forked from and answered with a toast from the server's 400.
`headLineage` is the client's copy of that rule and both views use it; the
server stays the authority.
`tools/tree_fixture.py` is the counterpart to `tools/branch_fixture.py` —
four branches at three fork depths, one forked off a fork. With no frontend
test runner it is the whole of the map's coverage, and it exists to be
looked at. Nothing in the backend changed; the 440 tests pass unmoved.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
SP7 replaced the pager with chips, on the grounds that a chip could also offer
"take this path" while a pager could only step. Driving it by hand said
otherwise, and the reason is worth keeping: the chip meant two different
things depending on where the reader was standing -- a real switch at the tip,
a preview needing a second button above it further back. Two meanings in one
control is what made the tree unusable.
So: one control that does one thing. Stepping reads a take and nothing else,
and it tells the server nothing, because reading is not a decision. The
transcript below a take that is not live simply ends -- such a take is a leaf
by construction, since whatever was played after the turn was played after the
take that *is* live. The decision is made by writing, and `after_id` carries it.
One step does reach the server and is still not a fork: a take with a story of
its own lives on its own branch, so going there is a branch switch and only the
server can say what is underneath. `branch_id` on the take is what tells the
two apart without asking first.
And a fork button on every turn but the opening. On the AI's it regenerates; on
your own it opens the text so you can say something else. What the story made
of the old take is kept, on the line it was written on.
`selectVariant` and `forkFromAttempt` leave the client. Both endpoints stay --
tested, and `stand_on` is shared with the write path -- but the pager needs
neither.
426 backend tests; lint and build clean. Not yet driven by hand: the frontend
still has no test runner, so this needs the `--keep` fixture and eyes, exactly
as SP7 did.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
A hand-written memory used to carry a NULL depth, described in the model as
"belongs to the adventure rather than to a path". That sounds harmless and
is not: a NULL is a coordinate no fork can cap, so a note typed on one line
followed the reader onto branches whose story it never described. It takes
the head now — the story you were reading when you wrote it — and obeys
exactly the rule a summarised memory obeys.
The unanchored escape clause in lineage.Path.clause existed for that single
case and is deleted rather than left unused. Its docstring argued that a
capped depth would drop a typed memory the moment its branch stopped being
the newest entry; anchoring answers the same worry better, because the
memory is not exempt from the path, it is on one.
The drawer now shows the path being read and nothing else, filtered by the
clause retrieval itself uses, so the bank you can see is the bank the model
can see. Nothing is stranded: a memory lives on a branch, switching to that
branch shows it, and deleting the branch deletes it. Pinning decides order,
the path decides existence.
Migration 62 lands existing NULL-depth memories at depth 0 of their branch
rather than at the tip. 0 is at or before every fork point, so every memory
stays visible from exactly the paths it is visible from today — nobody's
bank loses a row on deploy. The tip is the tidier-sounding choice and would
have emptied them out of every branch forked earlier than they were typed.
This supersedes the on_path flag and the "another branch" badge from
earlier today; anchoring makes them redundant, and they are removed.
Four tests changed because they asserted the old contract, not because
they broke. The one worth reading is the pair replacing
test_a_hand_written_memory_is_not_lost_at_the_first_fork: typed on shared
trunk it still survives a fork, and typed on ground the fork never
travelled it no longer follows you.
402 tests. Verified on tools/branch_fixture.py: each branch's drawer holds
its own memory and not the other's.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
Retrieval has been path-scoped since SP3: a memory on a branch this story
never travelled is never sent to the AI. The drawer listed every memory
alike, so on a fork you read "Fell down the cellar stairs; badly hurt" and
reasonably concluded the model knew it. It does not. That is worse than
either hiding the row or retrieving it — it is the screen claiming
something the engine contradicts.
Hiding them is not the answer either; that was the original comment's
point, and it stands. A memory nobody can list is a memory nobody can
delete, in a phase whose rule is that nothing is removed automatically.
So the whole bank still lists, and the rows off the current path are set
back, dashed, and labelled "another branch". MemoryOut.on_path carries it,
computed from the predicate retrieval itself uses rather than a second
spelling of the same idea — two spellings drift, and the failure mode here
is a badge that says the opposite of what the model gets. Pinning does not
override it: the path clause runs before pinning is considered.
The relationship is asymmetric and there is now a test that says so. A fork
borrows its ancestors, so a memory written on the parent is on the fork's
path too; the reverse never is. Worth pinning before somebody makes it
symmetric on the grounds that it looks wrong.
399 tests, three new, egress ceilings intact — the flag costs one id-only
query. Verified in a browser on tools/branch_fixture.py.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
The pager could only step between attempts, and stepping has nothing to say
about the thing the tree exists for: taking a path the story moved past and
keeping both. Every attempt has been its own node since SP4, so a chip is a
node now, and "take this path" forks — or simply switches, when the turn is
still the tip and its attempts are leaves nobody has built on.
Beside it, a Branches panel: every line the story has taken, with where each
left its parent and where it ends, and switch, rename and delete-with-confirm.
It sits with Plot/Memory/Scripts/Insights rather than inventing a new place to
put a rail. An unnamed branch is drawn from its fork depth, never from its
position in the list — a position shifts the moment a branch above it goes.
A spatial per-node map was considered and deliberately not built. At the size
this has to be verified against it is a second windowing problem, and it can be
added later without a new endpoint, since the rail and a map read the same
GET /branches. VariantOut grows an id because a fork is addressed by the node
being taken, not by an ordinal in a group that renumbers.
Driven by hand against the 602-action fixture, which found one bug that no test
could: the panel refreshed on actions.length, and a fork swaps a 60-action
window for another 60-action window, so it went on drawing a one-branch tree
while the story was already on the second. It keys off the counter adoptWindow
bumps now.
The scroll path was driven at the same time — three prepends of ~16,200 px, the
same node holding viewport top 792 to 787, never thrown to the end. That closes
the standing gap in this project. Console clean.
396 tests, build clean, no new lint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015H5qiyiR7gtFQaoDphHZ3g
Opening a finished adventure fetched every action in one response: 589.5 kB on
production's longest, and nothing about that curve bends on its own, because a
story only ever gets longer. The page load now brings the newest 60 actions
and the reader pages up from there. On the harness's 600-action fixture that
is 606.0 kB down to 62.6 kB, and -- the part that matters -- it no longer
depends on how long the story is.
Paged by anchor, not by offset. `before_id` is the oldest action the caller
holds; the server returns what precedes it. An offset counted back from the
newest would shift every older position the moment a turn lands, which is
exactly when someone is likely to be scrolling, and the reader would get one
action twice and never see another. It also keeps working when the story stops
being a flat list: comparing indices to order a branch survives the story tree,
treating them as positions does not.
`has_more` comes from fetching one row past the window rather than from
counting. A deleted anchor -- undo, mid-scroll -- reports the end rather than
guessing and serving a page the reader already has.
GET /{id}/actions and POST /{id}/undo now return {actions, total, has_more}
instead of a bare list. Undo is the action most likely to be repeated several
times running, so having it re-fetch the whole story would have undone the
paging on the worst case.
The adventure payload gets its window through set_committed_value rather than
by assignment: the actions relationship cascades delete-orphan, so assigning a
60-item list to it would delete everything outside the window on the next
flush.
In Play.jsx the prepend is followed by a useLayoutEffect that restores the
scroll position, before paint, so the story does not jump. Loading starts 400px
from the top rather than at it, guarded by a ref because scroll fires far
faster than React re-renders. There is a button as well as the scroll trigger:
on a short viewport the transcript may not be tall enough to scroll at all, and
a reader who cannot scroll must still be able to reach the beginning.
Verified against a running backend and a 220-action adventure: the page load
returns 60 of 220 ending on the newest, and walking back from an anchor returns
exactly the actions before it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Dvvqn9ZDR4ixeFPHNbww7
Both composers are position:sticky;bottom:0, which pins them to the bottom
of the *layout* viewport. A phone keyboard does not shrink that viewport --
it only covers it -- so the first tap on the input left the box sitting
behind the keyboard. Tapping back and focusing again appeared to work only
because the page had been scrolled during the first attempt, so the
browser's scroll-into-view landed somewhere else.
Two engines, two halves:
- Chromium honours interactive-widget=resizes-content on the viewport meta,
which makes the layout viewport shrink when the keyboard opens, so 100dvh
and sticky bottoms account for it on their own.
- iOS Safari ignores that flag and only shrinks the visual viewport, so
measure the covered height from visualViewport and publish it as
--kb-inset; the composers offset their sticky bottom by it. On browsers
that already resized the layout viewport the measurement is ~0, so the
same rule is a no-op there instead of a double lift.
The pages also gain a matching bottom margin so the last beat can still be
scrolled clear of a lifted composer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BC7Dsg4gSQm3NLcZcG5SK
The panel reported one number for the whole prompt and one flat fill bar,
so "3412 / 4000 tokens" never said which section was spending it. Replace
the fill with a stacked bar — one segment per context section, scaled to
the budget so the leftover width is the remaining headroom — plus a legend
ordered biggest-first with each section's share and token count. Hovering
either the bar or a legend row dims the rest; clicking a row jumps to that
section's text below, and each section header now carries its own share.
Section colours move out of CSS into a single SECTION_COLORS map applied
inline, so the bar, the legend and the section headers cannot drift apart,
and related sections keep their hue family without sharing an exact shade
(two identical colours read as one segment in a stacked bar). The narrow
side panel only fits one legend column, so the long tail folds behind a
"+N smaller sections" toggle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011BC7Dsg4gSQm3NLcZcG5SK
The global input styling in index.css enumerates types, and `email` was
never in the list, so the Email box fell through to browser defaults while
the Password box right below it got the app treatment. Added it to both
copies of that selector list -- the base rule and the 16px iOS-zoom rule in
the mobile block; missing the second would revert the fix on phones.
While in there, the modal itself: log in / sign up are now segmented tabs
instead of a link crammed into the button row, fields carry autoComplete so
password managers work at all, a show/hide toggle on the password, Esc to
dismiss, and the error is a real role=alert box rather than the shared
.test-error text.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
Retry used to delete the last AI action and generate a replacement, so the
discarded narration was simply gone. The row survives now: each attempt is
appended to actions.variants with variant_index naming the live one, and a
ChatGPT-style pager under the message browses them.
Action.text still mirrors the active variant, so the context builder, memory
bank, summarizer and export needed no changes. A variant carries only what
differs between attempts -- the text, the reasoning, and the state it
produced -- never the assembled prompt, which is identical across attempts of
one turn and is the bulk of context_snapshot.
Only the last message can be switched, restoring the script/world state that
attempt produced; earlier turns were written as a continuation of whatever is
active there, so theirs are read-only previews.
Three things that would otherwise bite:
- generate_turn now wraps _generate_turn and watches for a save sentinel. If
the generator ends without it (provider error, empty reply, script stop,
client hangup) it re-applies the previous variant -- otherwise a failed
retry leaves rolled-back stats under un-rolled-back text.
- Retry reuses the turn's own index rather than next_index, or the clock the
world-state cooldowns run on advances on a re-run of the same turn.
- Editing a message rewrites the active variant too, or paging away and back
silently reverts the edit.
Migrations 34/35 verified as an upgrade against a populated database, not
just a fresh schema. test_state_revert's retry test asserted the old delete
behaviour and was rewritten.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
NPCs are still part of stat_schema, but editing them inside the World State
panel meant a nested dashed sub-box among four flat stat lists. They now get
a top-level "Cast & NPCs" section: one card per NPC (avatar, name, npc.<id>
address, triggers, description, its own stats) in a responsive grid, saved
through the same schema path via the new NpcEditor/addNpc exports.
The panel's real inconsistency was CSS, not layout: the global field rule
keys off input[type="..."], which never matches the editor's typeless key and
description inputs, so those rendered with browser defaults beside properly
styled number boxes. One scoped base rule now covers the whole editor and
every entry is the same tile with a caption over each control.
Two bugs fell out of that: .se-band-n (0,1,0) always lost to
input[type="number"] (0,1,1), so band bounds rendered full-width; and the
exclusion has to be :not(:where([type="checkbox"])) because a plain :not()
inherits its argument's specificity and swallows every override.
In Play, NPCs group under a single "Cast" heading as compact plates instead
of each becoming another top-level group in the rail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
An adventure copies its scenario's plot text and story cards at creation so
later authoring never disturbs a story in progress. This is the explicit
opt-out, alongside the existing per-script "Sync from library".
GET /adventures/{id}/refresh returns a plan (per-field old/new diff, card
add/update/remove, world-state added/removed paths, and any ${...} answers
still needed); POST applies it under the turn lock so it can't race a
generating turn. The Plot panel shows the plan in a confirm modal first.
Overwrites the plot fields and scenario-derived cards. Deliberately left
alone: the opening `start` action (the story is built on it, and it is baked
into memories and the summary), the adventure's own title and summary,
player-authored story cards, and the live value of every stat the schema
still defines.
Two enablers were needed:
- adventures.placeholders (migration 32). ${...} answers were used once at
creation and discarded, so re-copying scenario text would have re-injected
a literal ${Hero}. Adventures predating the column re-prompt once via the
existing modal, then the answers are saved.
- story_cards.source_ref (migration 33), "card:<id>" / "npc:<key>", NULL for
player-authored. Adventure cards had no link back to their source, so a
rename read as delete-plus-add and player cards would have been clobbered.
Legacy cards name-match once, then adopt the ref.
World state goes through a new worldstate.reconcile(): keep values the schema
still defines, add missing ones at their initial, drop removed ones and their
cooldown bookkeeping. Not instantiate(), which would heal the player to full
and wipe their milestones.
The confirm modal is portalled to <body>: .side-panel's panel-in animation
has fill mode `both`, which makes it the containing block for position:fixed
descendants, so an overlay rendered in place was trapped in the 420px panel
and clipped by its overflow.
14 new tests in backend/tests/test_scenario_refresh.py; 85 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XzBCyXH4hBEVHqEaertcq4
Editing a beat opened a fixed 110px box. Player actions are one line, but AI
beats are several paragraphs, so editing one meant working through a keyhole.
New AutoTextarea sizes to its content on mount and on every change; CSS keeps a
110px floor so a short player action doesn't collapse, and a 65vh ceiling that
scrolls internally so Save/Cancel can't be pushed off the screen.
On phones both composers are one flex row, which left almost nothing to type
into: measured at a 390px composer, the Play textarea got 138px (Do/Say/Story
eat ~180px, Send ~75px) and the chat textarea 310px. Both now wrap, with the
buttons on their own line and the textarea taking the full 390px — Play keeps
Do/Say/Story and Send together on the top row, chat drops Send beneath. Play
stays two rows, so nothing gets taller. Desktop is unchanged; every rule is
inside the existing max-width:720px block.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
AI Chat is a plain scratchpad for talking to a model directly — no story
context, scripts or world state — for poking at models, prompts and endpoints
without starting an adventure. Power users only: the router 404s (rather than
403s) for everyone else and the nav link is hidden. The conversation lives in
localStorage, so there's no new table or migration.
is_power_user() now also returns True in local mode: it's the operator's own
machine and their own key, the same reasoning that makes the provider debug log
local-only.
Alongside that, the rule keeping the shared demo key off paid models now lives
in exactly one place. It had been duplicated into the chat router, which is how
one copy eventually drifts:
- resolve_provider_config() takes an optional model_override and is the only
place the whitelist is applied, so turns, AI Chat and the connection test all
inherit it. An override is a per-request preference, never a grant.
- ProviderConfig.__post_init__ refuses to exist when api_key is the demo key
and the model isn't whitelisted. It keys on the key itself rather than the
using_demo flag, so a mislabelled config can't slip past, and it raises so a
future path that bypasses the resolver fails loudly instead of billing.
- The demo branch still pins endpoint_url too — a user-controlled endpoint
would leak the key itself, which is worse than spending it.
Provider gained chat(messages, ...) beside generate(), both delegating to a
shared _stream(url, body); completion-mode endpoints get the messages flattened
into a labelled transcript. Settings' /models fetch moved to
list_endpoint_models() and is shared with /api/chat/config.
Tests: 10 new in tests/test_chat.py (70 total). These deliberately do not stub
resolve_provider_config — the point is to exercise the real BYOK-vs-demo
decision and assert on what the provider actually received: off-whitelist
override pinned, off-whitelist Settings.model pinned, redirected endpoint
pinned, BYOK passed through untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
Rework the home page into a single landing surface and give the app a
consistent visual language, chosen as "illuminated tome" over two other
pitched directions because it builds on the existing Cinzel + gold identity
instead of replacing it.
Home is now Continue (up to 4 in-progress stories, each showing where you
left off) over a scenario shelf, each section with a "See all" link. The
full adventure list moves to /adventures.
Scenario cover art has three tiers, in precedence order: an uploaded picture
(downscaled client-side to 400px WebP before storing), an emoji, or gradient
art generated from a hash of the title so no card is ever an empty box.
Adventures inherit their scenario's art. Images live in the row rather than
on disk because Render's free tier has no persistent volume, and it keeps
export bundles self-contained; list responses carry a cacheable
/api/scenarios/{id}/image URL rather than the base64.
Also: ambient drifting motes behind the app, loading skeletons, staggered
card entrance, ornamental scene breaks and a drop cap in the story, a
"Weaving" thinking indicator, and a toast system replacing every alert().
Two fixes found along the way:
- Importing a scenario bundle with no "tags" key returned a 500. Column
defaults are not applied until flush, so the attribute was still None
when the width clamp sliced it.
- Anything meaning "the story's latest narration" was missing action type
"start", which is the only text a freshly created adventure has, so new
adventures looked empty. Collected as NARRATION_TYPES.
Migrations 30 and 31 add scenarios.image and scenarios.icon; both are
additive with a '' default and were verified against a database stamped at
29. vite.config.js now reads AIDND_API_PORT so the recurring port-8000
clash with another local app needs no file edit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014FGY1yvzSeKgTtRfeVtDmx
The stylesheet had no real mobile design (one media query that only hid a
nav hint), so on phones the Play screen's four side-by-side columns
overflowed horizontally. Add a max-width:720px layer that collapses the
Play layout to a single story column, turns the World/Script State rails
into left-edge tabs that open as slide-in overlays, and makes the
Plot/Memory/Scripts/Insights side-panel a full-screen overlay. Also add a
hamburger dropdown for the top nav, switch the play-layout heights from vh
to dvh (fixes the mobile address-bar jump), and bump inputs to 16px to stop
iOS zoom-on-focus.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UWVyFKvqJGjfbXdibLgkMe
Players/authors can now directly correct the live values of stats the
schema already defines (health, trust, flags, milestones, etc.) without
waiting for the AI to emit a delta. New apply_override() sets values
absolutely rather than adding deltas, and — unlike the AI-facing
apply_delta() — bypasses cooldown/max_delta_per_turn and lets
milestones be un-set, since this is a deliberate correction rather
than a turn to police. Exposed via PUT /adventures/{id}/world-state
and an edit toggle in the World State drawer. Also fixes the
schema-editor stat-kind dropdown and free-text initial-value input
to size consistently with the numeric fields next to them.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The new kind <select> fell through to the generic input/select CSS
rule instead of the compact .se-num sizing, making it visually
inconsistent with the min/max/initial number boxes next to it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Each AI action now carries a compact world_changes summary (derived from its
stored snapshot), rendered as small chips beneath the message: numeric stats
show a signed delta (green up / red down), flags show on/off, milestones show
a check. Gives an at-a-glance "what changed" without opening Insights.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON view no longer parses/saves on every keystroke — you type freely
and click "Save JSON" to apply. Invalid JSON (or a non-object) shows the
parse error inline instead of auto-rejecting; valid JSON is pretty-printed,
applied to the form/preview, and saved. The form editor still auto-saves.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Scenario editor now has a "World State (RPG)" section with an Editor/JSON
toggle. The Editor is a form UI to add/edit/remove world & player stats
(min/max/initial/±per-turn/cooldown/counter/desc/bands), NPCs (name, keys,
description, and their own stats), flags, and milestones — no JSON required.
Both views edit the same parsed schema and stay in sync; empty clears the
RPG layer. Keeps the read-only structured preview under the JSON view.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Structured world/player/NPC stats, two-way flags, and sticky milestones
per scenario (stat_schema). The AI proposes a per-turn delta; a Python
engine referees it (clamp to min/max, per-turn cap, cooldown, counters).
Band word-labels plus a fixed stat guide (descriptions + full ranges)
keep the model grounded. World State drawer + Insights delta report;
undo/retry roll it back via the Phase 11 snapshot pattern.
- migrations 26-28 (scenarios.stat_schema, adventures.world_state,
actions.world_state_before); all nullable, additive, safe on existing rows
- migration 29 raises the default context budget 4096 -> 16384
(custom values preserved)
- seeded demo scenario 04-rpg-world-state.json (Bandit Camp)
- 19 new tests (33 total pass)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Swap the single-line action <input> for a <textarea> that grows with its
content (capped at ~4 lines, then scrolls) and shrinks back after send.
Enter still sends; Shift+Enter inserts a newline. Buttons stay bottom-aligned
as it grows.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
The State drawer stringified non-string values with JSON.stringify, so
nested objects/arrays showed as raw JSON. Add a recursive StateValue /
StateTree renderer: primitives are typed and coloured, objects/arrays are
collapsible and indented so deep state shows its structure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
The panel only showed a name + on/off toggle, so demo (server-owned)
scripts — which never appear on the personal Scripts page — had no
in-app way to inspect their code. AdventureScriptOut already ships the
hook sources, so add a "View code" expander that renders the non-empty
hooks (library / onInput / onModelContext / onOutput) read-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Exposes the scripting `state` object (every variable scripts read/write
via state.x, persisted per-adventure) in a collapsible left rail that
refreshes after each turn. New owner-scoped GET
/api/adventures/{id}/script-state endpoint; the drawer only fetches
while open and shows an empty-state hint until a script sets a variable.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
- Pin the play header (Plot/Memory/Scripts/Insights toggles) just below
the top nav so they stay reachable without scrolling back up to the
first message.
- Fix streaming autoscroll: snap to the real document bottom instantly
instead of smooth-scrolling to storyEndRef. That ref sits above the
sticky composer, so block:'end' stopped short — hiding the latest line
behind the input bar and yanking the reader back up when they scrolled
down. Smooth behavior also never settled at per-token speed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.
Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg