Commit Graph
23 Commits
Author SHA1 Message Date
JesseMarkowitzandClaude Opus 5 1ce9972760 M8: the browser becomes the storyteller
The interface was AI-DnD's with this product's features bolted into it. The
navigation read Home · Adventures · Scenarios · Settings · AI Chat; starting a
story meant first picking a *world*, and making a world meant a JSON stat-schema
form, a story-card table and an art picker. The play screen had a Branches tab.
The input had three modes. Sixteen of the sixteen controls on a two-turn story
had no accessible name — they were single glyphs with a tooltip.

All of that was measured in a real browser before anything was changed, and the
measurements are in planning/reports/M8-IMPLEMENTATION-REPORT.md §C. Almost
nothing underneath was wrong: the play loop, the history controls, the takes,
the Save Points, the state correction and the knowledge library all worked. What
was wrong was what a reader was asked to understand in order to use them.

So the shape now is one entry point and one screen:

  Campaigns -> Campaign -> Story
                           State · Knowledge · Context · Save Points · Settings

Everything that is not the story lives in a panel that starts closed. The
top navigation bar is hidden on the story screen entirely, because on that one
screen the story is the interface.

Play is one natural-language field. An action and a piece of quoted dialogue are
both just what the reader wrote, and B01/B02 confirmed against a real narrator
that the model reads the quotes without being told which kind of turn it is.
What survives from the old Story mode is a Story direction toggle, which is not
a fourth mode: it changes who is being spoken to, not what kind of action is
taken, and the box is visibly marked while it is on.

Branch, fork, node, merge and head appear nowhere a reader can see them. The
branch panel and the tree overlay are gone from the browser. The mechanism is
untouched — takes, divergence, retained futures and Save Points all still work,
and their endpoints are still tested. This is a decision about what a reader is
asked to understand, not a reduction of what the product can do.

The two defects worth the space:

A player action is stored with AI Dungeon's "> You " prefix. That was right when
the Do mode asked for a bare verb phrase. With one field the spec tells the
reader to write "I enter the tavern", and the result was "> You I enter the
tavern." — in the transcript, in the replayed history, and therefore in the
narration, where a small model imitates it and writes "You I thank her". M8's
own design surfaced it, so M8 fixed it: the prefix is added only when the reader
has not already written a subject. The ">" marker, which is what actually
identifies a player turn in the prompt, is unchanged in every case.

And a stale `.input-bar { display: flex }` in play.css overrode the new
composer, because that sheet is imported after the new one. The direction row
and the input row laid out side by side and the box was unusably narrow. Found
by opening the product in a browser, not by reading the CSS — which is the
argument for having done that first.

Failures now have the taxonomy the spec asked for rather than one toast: model,
generation, state, knowledge, server, each with the thing to do about it. A
failed turn leaves the reader's words in the box and says so. The classification
reads backend strings, so it is a fallback ladder rather than a lookup — an
unrecognised message still classifies, still shows the server's own words and
still offers Retry.

`Settings.model` could be empty with nothing saying so until the first turn
failed with a provider error. The header now reports Ollama in five states, and
an unconfigured or missing model offers the models actually installed on the
endpoint, from the connection test that already knew them. Nothing is chosen
automatically: an endpoint's first model may be an embedding model, which cannot
narrate at all.

Narrator prose is rendered as safe Markdown — headings, emphasis, lists,
blockquotes, code. The safety is structural rather than filtered: every node is
a React element built from parsed text, and there is no dangerouslySetInnerHTML
in the file. A sanitizer is not needed to make markup safe if markup is never
produced from input. Link schemes are checked with the URL parser rather than a
pattern, because the bypasses are all in the parsing. A remote image is a
placeholder naming the blocked address; the knowledge and context panels
deliberately do not use this renderer at all, because they exist to show a
reader exactly what is in their file.

Backend, and only what the browser could not otherwise reach:

  AdventureCreate.opening   a start action could only come from a Scenario, so
                            every campaign made in the new setup flow opened on
                            a blank page. Same node, same code path.
  canon_rules               campaign_canon has been the highest authority in a
                            campaign since M5, read by the prompt builder and
                            the state validator, and had no API at all — a
                            fixture had to write it with SQL.
  a 401 and a 429 message   the last user-facing text describing a hosted
                            deployment. One told the reader to check an API key
                            that has not existed since M2.

No schema change and no migration: proved by building a database with a server
running the M7 commit's own code and opening it with this one.

The project had no frontend tests. It has 132 now, across ten files, running
in about six seconds — the enabled state of every history control, the take
selector, the confirmations, the panels, the five model states, the failure
taxonomy, the focus trap, accessibility, and that the reserved dictation control
never touches the microphone. Writing them found a real defect: the focus trap
filtered candidates with offsetParent, which is null inside the fixed-position
ancestor the dialog has and which jsdom never computes — it would have behaved
differently in the tests from the browser.

They do not replace the real-browser runs, and both kinds of evidence are in the
report. The browser suites drive the production build served by the real backend
with a real local narrator, including a genuine process restart.

A verification pass over all of it then found three more, each by driving the
product rather than reading it:

Stepping between alternate takes did nothing. The pager asked whether a take
lived on another line by comparing `target.branch_id !== action.branch_id`, and
`ActionOut` has never carried `branch_id` — so the comparison was permanently
`number !== undefined`, always true, and every step took the branch-switch path.
For two takes of an ordinary retry, which share a line until one is written
below, that meant switching to the line already being read: the same window came
back and nothing moved. D07 is a required v1 acceptance test. The fix needed no
new field — the variants list already carries every attempt's branch and marks
the live one.

The first regression test for that passed against the broken code, because its
fixture gave the action a `branch_id` the real payload never sends. That is the
exact failure M7's review was about, so the fixture was corrected, the tests were
re-run against the reverted code and failed for the right reason, and the
fixture now carries a docstring saying why the field must never come back.

And the knowledge panel pointed readers at an "embedding model" while the
setting is called "Model for meaning-based search" — a reader sent looking for a
field that does not exist by that name.

Campaign canon was measured rather than assumed. Editing it after play is a
configuration change: every turn already played keeps the canon it was actually
given, in its own context snapshot, and the accepted story, the state document
and the state audit log are byte-identical across an edit. It is not routed
through M5's state audit, because canon is not narrative state and doing so
would create the second representation the spec forbids. What the editor does
now is say so, once a campaign has moments.

`BROWSER-UX-SPEC.md` §38 asked for a "Show Hidden Story State" toggle. There is
no hidden story state — a secret lives in a narrator-only knowledge source and
never enters the state document. The section is rewritten to require what it
actually meant: ordinary surfaces must not carry narrator-only information,
advanced inspection must withhold it by default behind an explicit warned
choice, and no second store may be invented to give a toggle something to
reveal. The protection is stricter than before, not weaker.

Closeout. An independent review returned M8 IMPLEMENTATION: PASS subject to
evidence and documentation cleanup, and this commit carries that cleanup:

The report named two frontend bundles as the artifact behind its acceptance
evidence. The saved run logs settle it. index-Ii-lARp9.js, built at 18:53:02
from this tree, is the one final frozen artifact behind all 157 browser checks;
index-C6E5Uvtu.js is superseded — it predates the D09 fix and its acceptance
suite ended 54/55 on exactly that defect. No tracked file under backend/app or
frontend/src has a modification time after the freeze, so the whole final
campaign describes one build. §P sets the two side by side.

Finding 14 — the app budgets 16,384 prompt tokens while an Ollama that sees no
VRAM enforces 4,096 — is resolved operationally, with no application change.
The OpenAI-compatible endpoint this app speaks accepts num_ctx and ignores it,
and reloads the model at its own default, so a native call cannot prime it
either. A model derived with POST /api/create carries the parameter, is honoured
through the app's own OpenAI-compatible path, and appears in /v1/models — which
is the listing the Settings model picker already reads. Measured end to end.
The procedure is in DEVELOPMENT.md; nothing in the repository depends on any
particular derived model existing. Adding provider code to work around this was
declined deliberately: it would mean either a second native request path,
against ADR 011, or a parameter the endpoint provably ignores.

The §38 rewrite is ratified as a requirement clarification aligned with the
implemented architecture, and the spec gains the clause finding 3 was really
about: withheld material must be absent from the rendered DOM, not merely
collapsed in it.

The report's §U carries the M9 handoff — what a portable campaign has to include,
whether historical context snapshots belong in the bundle, what happens to
inherited story cards, and that a restored campaign may meet a different context
window than the one that wrote it. None of it is implemented here.

Final: backend 950 passed / 14 skipped; frontend 132 passed; lint, production
build and Docker build clean; 157 browser checks across six suites, zero
failures. M8 is implemented, verified, reviewed and accepted (2026-09-06).
M9 has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 23:31:45 -04:00
JesseMarkowitzandClaude Opus 5 480414efe0 M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or
Inspiration, and the class is load-bearing rather than a label: it decides the
words a passage is framed with in the prompt, the weight it carries when
passages are ranked, and which budget it competes in when the context is tight.

This is a separate subsystem, which is the Phase 0B decision
(IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification,
provenance, content identity, chunking, an index or a lifecycle, and they were
not promoted into something that does. Nothing here reads or writes one.

The subsystem, in backend/app/knowledge/:

  classes      the three classes, their weights, and the prompt framing
  chunking     deterministic, heading-aware, 60-800 tokens, no overlap
  fts          SQLite FTS5 with porter stemming; scoped and bounded in SQL
  importer     validate, hash, store, chunk, index — in one transaction
  embeddings   local Ollama vectors through the shared provider
  retrieval    query construction, hybrid merge, rerank
  inject       the budgeted cut and the rendered prompt sections

Relevance admission is a separate stage from ranking, and that separation is
the milestone's most expensive lesson. An independent review found the first
implementation deciding relevance with a floor expressed as a share of the best
candidate — which the best clears by construction — so a passage was admitted on
every turn regardless of the scene. A query about tide tables and container
tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden
Canon among them.

So the pipeline is now:

  candidate generation -> admission -> ranking -> class weighting -> budget

Admission reads raw, candidate-set-independent signals: the cosine the model
returned, and how many distinct meaningful query terms a passage contains.
Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero
is not zero. Normalization decides order among things that matched; it can never
decide whether anything matched. Authority is applied after admission, so a
class orders what matched and never rescues what did not.

Retrieval may therefore return nothing, and on a scene unrelated to the library
it does.

The other decisions that each replaced an obvious wrong one:

- The class multiplies relevance rather than adding to it. An additive bonus
  satisfies "Canon outranks Reference" and makes "do not include irrelevant
  Canon" impossible, because a large enough constant wins on its own.
- The semantic floor is measured, not guessed: 113 production-path pairs against
  nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at
  0.36-0.56, and 0.58 sits between them. Because it is a property of that model
  and not of cosine similarity, it is keyed to the model rather than applied to
  whatever is configured: an embedding model with no measured calibration in
  this build does not borrow the number. Semantic admission is skipped, the
  campaign retrieves lexically, and the reason is stated in the knowledge status
  and in the turn's provenance. Degrading to lexical keeps the library usable;
  lending the threshold to an unmeasured model is how the admitted-everything
  defect would return.
- One lexical term is not evidence. Two distinct meaningful terms, or one that
  is neither a standing campaign entity nor a negligible share of the query.
  The stop list grew from 42 words to 261, all function words — no subject
  matter, because a stop list that removes subject matter stops finding "The
  Silver Key".
- Lexical retrieval is a production path, not a fallback. It finds the proper
  nouns and invented terms a setting bible is made of, and the library is fully
  usable with no embedding model configured.

Safety is structural rather than filtered. Imported text reaches the prompt
whole, inside a section that says what it is, under a rule stating the authority
order in words and refusing every instruction inside it. No endpoint accepts a
filesystem path, so H08 has no mechanism to escape from. Nothing renders
imported content as HTML, so a script tag is five visible characters and a
remote image is never fetched. Import, chunking, indexing, retrieval and a turn
open no socket at all; only embeddings do, through the endpoint allowlist the
memory bank already uses.

Provenance is the rendered text, not a foreign key: deleting a source cannot
turn a historical turn's evidence into dangling ids.

Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5
virtual table attached to knowledge_chunks as a DDL hook so it is created and
dropped with the table it indexes. Migration 92. A pre-M7 database opens
unchanged and needs no sources to play.

Bundle: the source content and the reader's judgements about it travel; the
passages, index rows and vectors are rebuilt on import, so a restored campaign
is searchable immediately without a reindex step.

One runtime dependency: python-multipart, Starlette's multipart parser. It is
what makes the upload surface possible, and the upload surface is why no
pathname is ever accepted.

The test doubles were the reason the defect shipped, so they were corrected too.
The retrieval stub scored unrelated text at 0.06-0.20 where the real model
scores it at 0.43-0.44, and its docstring said it had deliberately removed the
constant component that "would put a similarity floor under every pair" — which
is exactly the property real models have. The stub now has that floor, one test
fails if it is ever removed, and another reproduces the superseded rule and
asserts it is still fooled by the same fixture. Run against the pre-corrective
implementation, the new suite fails 13 of 18.

Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of
which mocks nothing between itself and Ollama and re-measures the similarity
separation on every run. 43/43 checks in a real Firefox, reproduced.
Docker build clean.

Four other defects found by review or by the browser run were fixed here rather
than carried: an unreachable relevance constant that appeared to enforce
something and did not; acceptance tests using the wrong fixture files, so G07's
trap was never exercised; a bidirectional override surviving into displayed
filenames; and, from the implementation pass, the Insights panel showing M5's
two state sections as raw keys and the source inspector refetching on every
keystroke.

M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK
REQUIRED. Both blocking findings are closed, and closeout resolved the
embedding-model calibration boundary the corrective pass had left as debt.
planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective
closeout and the closeout verification in sequence, none overwriting another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 15:40:13 -04:00
JesseMarkowitzandClaude Opus 5 b7005e6fdd M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.

This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:

    visible active transcript position == stored head == authoritative state

Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)

  A narrator edit no longer rewrites a row. It returns to the state before the
  turn, takes the reader's exact text as the accepted narration, re-derives the
  state that text implies, and becomes a new active continuation — while the
  original narration keeps its words, its live flag and its whole future as
  retained history. At the tip the correction is another take; with story below
  it, it forks. No new history machinery: this is the existing fork/take/head
  path with the reader's text in place of a generated reply. The §14A refusal
  is therefore gone for narrator turns, and remains only for player input.

Pre-M5 positions

  Migration 88 backfills the empty narrative document onto every action written
  before M5, and a missing snapshot now restores the empty document instead of
  leaving the previous position's state standing. Restoring to an old Save
  Point no longer leaves a later position's entities and facts on screen.

Narrator context

  Replayed history carries prose only; the machine-readable block is no longer
  reconstructed into past turns, where it contradicted the authoritative state
  in the same prompt. A fact withdrawn by a manual correction is now named as
  no longer true, with the reader's reason, rather than silently dropped.

Also

  - state_changes joins the action-list bulk read, removing one query per row.
  - Extraction takes only the application's own protocol payload: an ordinary
    ```json or ```python block in a story survives, and a mangled proposal
    still does not reach the reader.

Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-05 07:01:50 -04:00
JesseMarkowitzandClaude Opus 5 62a997f364 M4: close out Save Points, with browser verification
Closes M4. The review's three findings are fixed, the durability rule the
specification always implied is now enforced, and M3's and M4's browser
behaviour has been verified in a real browser for the first time.

B-1 -- the Save Point list was an N+1 that loaded whole Action rows,
narration included, to answer "does a row exist here". It is now one bulk
two-column coordinate query plus one lineage: 53 SELECTs for 25 Save Points
became 5, and the count no longer grows with the list. The clause is an OR
of exact (branch, depth) pairs rather than two IN lists, because the cross
product would report a Save Point resolved on the strength of another one's
depth existing on this one's branch. A test builds exactly that trap.

B-2 -- reclassified during closeout from "missing warning" to a behaviour
defect, and fixed as one. STORY-BRANCH-SEMANTICS §19 says a named checkpoint
remains until explicitly deleted, and §28 already required future cleanup to
retain checkpoint-referenced paths; a cascade that silently removed Save
Points with a branch violated both, and a warning would only have documented
the violation. A branch a Save Point names can no longer be deleted. The
request is refused with the offending Save Points named, the user deletes
them explicitly -- which deletes no story -- and the branch then goes. The
scope is the subtree, because deleting a branch takes its descendants. Both
delete controls disable and explain. Recorded as a new §19.1; models.py,
TECHNICAL-DESIGN §8.8 and DATA-MODEL §8 had all recorded the cascade as the
rule and now record the refusal.

An earlier pass in this same closeout had kept the cascade and added a
warning. That was the wrong fix and its tests were replaced rather than left
standing, since they pinned the defect.

B-3 -- the D11/L03 automation never left one process, so it could not
distinguish durable state from a live Python object. It now spawns real
server processes, kills the first, and reads the campaign back with the
second.

C-5 -- creating a Save Point takes the campaign's turn lock. "Save where I
am" has to name one committed position, and the head is what a turn in
flight is about to move. Rename and Delete deliberately do not take it.

The architecture is untouched: a Save Point is still name + note +
(branch, depth), and restore is still coordinate -> head.move_to_node ->
head.move_to -> attempts.restore_state. No second restore path, no state
copied into a checkpoint, no fork on restore.

Browser verification -- the first in this project, and it covers both
milestones. Firefox 154.0.1 through geckodriver over the W3C WebDriver
protocol, driving the rendered DOM: 47/47 checks, twice, on independent
databases, no console errors. M3's Undo/Redo enable states, transcript
movement, Retry and the take pager, divergence retiring Redo; M4's whole
Save Point lifecycle, both confirmations, and the new branch-delete refusal
including its recovery. No dependency was added: the WebDriver client is
stdlib HTTP.

No application defect was found by the browser. Four failures occurred, all
in the harness -- a wrong SPA route, a wait comparing transcript length when
the empty-story placeholder is longer than the first turn, a fixture
deleting the branch it was reading, and a reload assertion that sampled
once instead of waiting. The last was checked against the app before being
called a harness bug.

Tests: 698 backend pass (was 680), 60 M4, 94 M3 history, 66 export/
migrations, 93 security/local-only. Frontend lint and build clean, Docker
build clean, loopback binding unchanged. No assertion weakened, no skip
added.

Planning: STORY-BRANCH-SEMANTICS §19.1 is the only behavioural change and it
strengthens §19. V1-ACCEPTANCE-TESTS records D11-D14, I04, L03 and the
E-series, keeping automated, live-runtime and browser evidence distinct, and
weakens no pass condition. DATA-MODEL records the coordinate with the retry
measurement that settles it. BROWSER-UX-SPEC rules for Moment over Turn.
BUILD-MILESTONES marks M4 COMPLETE, closes M3's browser condition, and lists
what M5 inherits. VERSION adds v2.6.

No new ADR: ADR 005 already decides that history is preserved rather than
overwritten, and §19.1 is that decision applied to checkpoint-referenced
history.

M4 is closed. M5 may now be briefed; it has not been started.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-04 06:34:56 -04:00
JesseMarkowitzandClaude Opus 5 e08d49c3eb M4: add durable named Save Points
A Save Point is a name for a story position, and restoring one is head
movement. That is the whole architecture, and it is what ADR 012 and
BUILD-MILESTONES' note on M4 asked for: M3 made the head a stored
(branch, depth) and made arriving at one a row lookup plus a state restore,
so a Save Point needs no restore machinery of its own.

What the user gets:

- Name the moment they are reading, keep playing, restart the app, and come
  back to it. Restoring moves the story back and deletes nothing: the later
  turns stay, Redo still walks forward into them, and writing something
  different is what starts a new line while the old one is kept.
- Rename, delete, and a list, in a Save Points panel beside the branch panel,
  with a Save Point button next to Undo and Redo. Both confirmations say what
  is *not* destroyed, because that is the part the screen cannot show.
- Save Points survive export and import.

What was deliberately not built:

- No second restore path. `head.move_to_node` is the only new movement: its
  depth half is M3's `head.move_to` unchanged, and its branch half is the
  single assignment `switch_branch` already makes. No head field is written
  in the checkpoint router, nothing reconstructs state, nothing prunes a
  memory, nothing copies or deletes a turn, and restore never forks — the
  first write below the restored head does, through `fork_if_behind_head`.
- No automatic cleanup. A Save Point behind the head, or naming a line the
  story left, is doing its job (STORY-BRANCH-SEMANTICS §19). The one removal
  is a cascade: deleting a branch takes its Save Points, as it takes its
  memories, because the story they named went with it.
- No new ADR. ADR 012 already decides the architecture, and a table is not a
  decision.

The one call the planning package did not already make: restore moves the
branch half of the head only when the coordinate is off the path being read.
Doing it unconditionally would quietly hand back an abandoned continuation
whenever a Save Point in a shared prefix was restored; never doing it would
make a Save Point on a departed line unrestorable, which contradicts §19.
TECHNICAL-DESIGN §8.8 records it.

Schema: a `checkpoints` table holding a name, an optional note and a
(branch, depth) coordinate — no copy of any story. `create_all` builds it as
it did `memories` and `branches`; migration 80 adds the index. No backfill,
because nobody had named a position before M4.

The coordinate is deliberately not an action id: one coordinate holds every
attempt at a turn and exactly one is live, so a coordinate follows a retry
where a row id would pin a take the story no longer tells.

Tests: 680 pass (638 before). 42 new in tests/test_save_points.py covering
D11-D14, I04, L03, E-series lineage and memory isolation after restore and
divergence, the edge cases, and an M3-database migration. One pre-existing
fixture in test_tree_migration.py needed `checkpoints` added to its drop
list — SQLite refuses to drop a table another table references.

Not verified: the browser. No session has had a usable one, so the Save
Point panel's DOM behaviour is unobserved — as M3's Redo control still is.
The twenty-step sequence was driven over HTTP against a live server with a
real process restart instead, and all seventeen checks pass. M4 is
implemented, not accepted: no review has been written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-03 18:48:54 -04:00
JesseMarkowitzandClaude Opus 5 3c8e91f644 Docs: correct post-M3 status and Ollama configuration
Three stale claims found in active documentation after the post-M3
consolidation:

- BUILD-MILESTONES.md still opened "M1 and M2 complete; M3 next", which
  contradicted its own M3 "Status: COMPLETE" block, planning/README.md and
  VERSION.md. It now states M1, M2 and M3 complete and accepted, M4 next.
- DEVELOPMENT.md described the settings row as "endpoint, model and (unused)
  API key" and sent "api_key":"" in its curl example. M2 removed the field;
  schemas.SettingsUpdate has no api_key. The prose and the example now match
  the real request shape. No application code was changed.
- README.md listed LM Studio as a supported local endpoint. ADR 002 and ADR 011
  make Ollama the only v1 backend; LM Studio is a rejected alternative there.
  The row is removed and the surrounding wording now says Ollama is the
  supported backend, same-host is the default, trusted-LAN Ollama is supported,
  public/cloud is prohibited, and the OpenAI-compatible adapter is an
  implementation detail rather than a support promise. The claude_shim section
  stays, relabelled "(development only)".

Documentation only; no code, schema or test changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-03 16:39:29 -04:00
JesseMarkowitzandClaude Opus 5 d27ee34901 Docs: consolidate active planning and archive historical material
The planning package had grown to where a new agent could not tell what was
authoritative. Phase 0 execution prompts sat beside the specification; four
completed milestone reports sat beside the current one; and upstream AI-DnD's
own `plan/` build log and `docs/` project site still described a hosted,
scripted, multi-user product with accounts — every screenshot in it showed a
Scripts tab and a Sign up button, none of which has existed since M2.

`planning/archive/` now holds the history and says so in its own README:
`phase0/` for the research that chose AI-DnD, `milestone-reports/` for M1 and
M2, `decisions/` for ADR 008, the Phase-0-before-build gate Phase 0 satisfied.
`planning/reports/` holds only the current milestone's report, because that is
the one M4 planning has to read; it moves to the archive when M4's replaces it.

Deleted rather than archived: the Phase 0B execution prompts and the
handoff/status/summary documents, the Phase 0A discovery and triage reports,
upstream's `plan/` and `docs/` trees, and `frontend/README.md`, which was Vite's
template boilerplate. All of it is in Git history, and the two recommendation
reports carry every conclusion the deleted research reached.

Archived documents are kept verbatim. Paths written inside them point at where
those files were when the document was written, which is the point: an evidence
record that has been quietly edited is no longer evidence.

Active documentation is corrected where it pointed at the removed trees or
described removed capability as present. `DEVELOPMENT.md`'s "things M1 did not
touch" list had gone stale at M2 and claimed QuickJS scripting was still tested;
its test count was 604 against an actual 638. `README.md` loses the upstream CI
badge, which reported upstream's pipeline rather than this fork's, and a
reference to `backend/app/worldstate/engine.py`, a file that does not exist.
`planning/README.md` is rewritten as the documentation index.

New: `planning/PROJECT-SOURCES.md` and `planning/project-sources.txt`, the
manifest of what belongs in the ChatGPT project's Sources.

Source comments referring to the deleted trees are reworded; no behaviour
changes. 638 backend tests pass, frontend lints and builds, and a reference scan
over all 48 tracked Markdown files reports no unresolved path in active
documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NCbwH7yLGKsj1rhXXzKSCu
2026-09-03 14:33:07 -04:00
JesseMarkowitzandClaude Opus 5 c8755c21c2 Planning: close M3 and record active-head architecture
M3's review recommended planning changes and, following the M2 pattern,
reported rather than applied them. This applies them, and adds the ADR the
review asked for.

ADR 012 records the architecture rather than the requirement. ADR 005
already says that going backward must preserve abandoned history and that
the user sees Undo/Redo/Retry rather than branch management; it names a
movable active head as the direction and stops. What M3 settled is the
shape: the head is stored rather than derived, every read of the story is
capped at it in one place, one mechanism moves it, the state of a position
comes off the node rather than from a replay, the first write below a
moved-back head is the divergence, and whether Redo exists is decided by
the lineage rather than by a flag that could be stale. The last of those
is the property worth keeping — a flag can be wrong and make the story
wrong; a lineage cannot.

Two semantics are ratified in STORY-BRANCH-SEMANTICS.md, both of them
reversals or narrowings that a reader would otherwise take for bugs. Undo
now crosses fork points and continues to the campaign opening, because
refusing at the fork was a consequence of deleting rows the parent line
was also reading, and nothing is deleted any more. And the system refuses
to switch which take is live while a later story is off screen, because
doing it quietly would leave retained history continuing from words the
story no longer says.

A new §14A covers editing in place. §14-15 describe the finished
behaviour — the edit becomes authoritative, the state it implies is
re-evaluated, a new continuation is created, the original is retained —
and that requirement is intact and explicitly not weakened here. It is
also not built, because re-evaluating state from prose a user typed needs
M5's extraction pass. §14A says what exists in the meantime and why
refusing is the minimum that holds the invariant rather than the
destination.

TECHNICAL-DESIGN.md gains §8.7 and §9.1, recording the implemented model
and the bundle behaviour as fact in the way §5.2 records M1 and M2. §10.4
gains a constraint that is easy to lose: the snapshot half of the hybrid
state model is a requirement, not an optimization. Head movement is a row
lookup plus a restore, which is why Undo, Redo and Save Point restore cost
the same at any distance into a campaign; a state model recoverable only
by replaying from the opening would make all three proportional to
campaign length, on exactly the long campaigns this product is for.

DATA-MODEL.md records the head as stored on the campaign rather than
derived from its newest turn — two campaigns holding identical turns can
be read at different places, and nothing about the turns can tell them
apart — and the branch disposition as implemented: the depth a divergent
write left the branch at, deliberately advisory, and carried through
export because every row of an abandoned line is exported either way.

BUILD-MILESTONES.md marks M3 complete and states the one condition still
open. M4 is told a Save Point is a durable pointer and that restoring one
is head movement with a bounds check, not a restore system: a second
mover is the specific failure to avoid, because the two paths would
silently disagree about what restore means. M5 gets three constraints —
keep state efficiently recoverable, move the test instrumentation rather
than the assertions when the world-state protocol goes, and finish the
narrator edit §14A defers.

V1-ACCEPTANCE-TESTS.md clarifies ownership without lowering a bar. D10
keeps all three pass conditions and is explicitly recorded as *not*
satisfied at the end of M3; what changed is that the document now says
which milestone delivers which condition. D03's result is recorded as a
full pass rather than the partial the text allowed for, I07 gains the
pre-M3 bundle clause, and L01 gains the note that resolves its apparent
conflict with A05 — a failed turn does advance the head by one, onto the
player's retained input, and that is A05 working rather than L01 failing.

README.md described a different application: a hosted demo, guest
accounts, cloud providers, Postgres, a Render blueprint, an analytics
dashboard, a QuickJS scripting engine, and 549 tests. M2 removed all of
that and the README was never updated — a gap M2's own debt table missed.
It now describes what this fork is, including the endpoint policy and the
TLS behaviour, and the numbers in it are the current ones.

M3's report is included here as its own evidence record: no separate
baseline report was produced, so it carries the raw counts and runtime
observations as well as the review, and §W records this closeout.

SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged. M3 altered no
product requirement and touched no path in the threat model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u
2026-09-03 13:54:44 -04:00
parththakkar106andClaude Opus 5 b1772c6e21 Plan the readability refactor, and clear the tree for it
Phase 17 splits the four files that hold most of the code, finishes the
schema migration SP8 left half done, and stops the published guide from
drifting away from its Markdown source. `plan/17-refactor.md` carries the
plan and the progress table, and `plan/STATUS.md` points at it.

Stage 0 is hygiene only. Both abandoned worktrees are gone, which freed
about 104 MB. Removing `sp7-tree-ui` needed one extra step: a Vite dev
server had been running out of it since 2026-08-18, holding
`frontend/.vite` open and owning port 5173, and serving a tree 54 commits
behind `main`. The three stale `.db` files are deleted; `data.db` is not.

`AIDND_TRUSTED_PROXY_HOPS` is now documented. It was read at `limits.py:55`
and named in no `.env.example`, README, or blueprint. It sets how many
proxy hops the rate limiter trusts in `X-Forwarded-For`, so a deployment
that adds a hop without setting it gets the bucket-rotation bypass back.

The 19 squash-landed branches are still there. `git branch -D` is blocked
by the permission classifier; the verified command is in the plan file.

549 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r
2026-08-29 00:36:21 +05:30
parththakkar106andClaude Opus 5 9398c13da5 Delete seeded scenarios no seed file claims any more
`previous_titles` stops a rename stranding the row it left behind, but the rows
already stranded still had to be deleted by hand on every deployment. The
seeder now removes them on the next boot, which retires the stale "Road to the
Champion" demo without a database console.

Only rows with a NULL owner and `is_public` are considered, and a player's own
scenario is neither, so nothing anybody created is reachable. An adventure
started from a deleted demo survives: `adventures.scenario_id` is `ON DELETE
SET NULL`, and the adventure holds its own copies of the cards and scripts, so
it loses only the inherited cover art.

Two cases skip the sweep, because neither is an instruction to remove live
content: a seed file that fails to parse claims no title, and an empty seed
directory reads as a packaging failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 19:19:23 +05:30
parththakkar106andClaude Opus 5 ef4da7fda2 Keep the local Claude test rig, and record what it found
`backend/tools/claude_shim.py` serves an OpenAI-compatible endpoint backed by
the local `claude` command line tool, so a demo can be played against a real
model with no API key. Each request spawns one `claude --print` process, which
suits the turn engine: the app assembles the whole prompt every turn and
expects a stateless endpoint.

Playing the Pokemon demo through it made all five of the world-state changes in
`plan/16` work, and the refusal loop ran end to end for the first time: turn 3
clamped to nothing, turn 4's assembled prompt carried the correction verbatim,
and the model's next delta was right. Two failures previously blamed on the
code were the demo model. One is a schema fault and is still open: Milo's three
Pokemon share one `active_hp` stat, so a switch leaves the newcomer at 0 HP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PacdRuPXSkQQy4ZYdH32hF
2026-08-28 18:42:24 +05:30
parththakkar106andClaude Sonnet 5 446ceb6cab Rewrite docs and READMEs in Google developer documentation style
Trim em dashes, convert first-person plural to second person, and
tighten sentences in README.md, docs/GUIDE.md, docs/self-review.md,
and frontend/README.md, matching the style already applied to code
comments. No technical content, numbers, or code blocks changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7vUFpYkLrJnuSgwcRjqSx
2026-08-26 16:07:34 +05:30
parththakkar106andClaude Opus 5 041f9e25f3 Count the visits, and say whether anyone got anywhere
A hosted demo raises a question a local app never does: is anyone using it,
and do they reach the part that matters? `/analytics` answers it — visitors,
pages, referrers, countries, devices, which shared scenarios get played, turns
and demo-key spend, API and turn errors, and a funnel from visited to played a
turn to signed up.

Not a third-party script, for reasons specific to this one. The CSP allows
`script-src 'self'`, so a tracker means loosening it; adblockers eat the
popular ones, which silently biases exactly the technical audience this
project gets shown to; and none of them can see the measurement that actually
matters here, which is a turn, not a pageview.

**A visit is a write and never a read.** After the 189x egress fix it would be
perverse to add a feature that reads rows per request, so counts accumulate in
a process-local dict and flush every 60s as UPSERTs. Storage is a generic
`(day, metric, label) -> hits` counter, so measuring something new later costs
a constant rather than a migration, plus one row per visitor per day for the
funnel flags. Every dashboard query is a GROUP BY returning tens of rows
however much traffic sits behind it; a month reads back in a few kilobytes.
The buffer's cost is that a hard restart can lose up to a minute — the flusher
also runs on shutdown, and a tier that sleeps when idle sleeps on an empty
buffer anyway.

**The counters are anonymous; the access log beside them is not, on purpose.**
A visitor is `HMAC(secret, "visitor:<user id>")` truncated to 32 chars —
one-way, so `analytics_daily` and `analytics_visitor_days` cannot be joined
back to `users`, and keyed, so no client can compute one. Story content never
reaches that module, and the only content it ever names is a seeded public
scenario's title; a player's own titles are theirs. `accesslog.py` is the
identifying half and is a separate module writing a separate table so that
separation is a property of the code rather than a convention: `access_events`
records sessions, sign-ins, registrations and failed attempts with address,
email and device, read on a second tab of the same page behind the same gate.

Both halves are gated on `AIDND_ANALYTICS_EMAILS`, not `POWER_USERS`. An
unmetered tester is not automatically someone who should see the traffic. The
route 404s and the nav link is absent for everyone else, the same treatment
AI Chat gets; unset in a hosted deploy means nobody sees it, including me.

Three things came out of building it that a test would not have suggested.

**A failed turn is an HTTP 200 with a bad ending.** The status-code middleware
cannot see one, so a demo whose model had started refusing every request would
look perfectly healthy from outside. All five SSE error paths in
`_generate_turn` now go through a `turn_error()` helper that counts on the way
out. Error buckets elsewhere are labelled by the matched route template rather
than the requested path — one bucket per endpoint instead of one per adventure
id, and, the reason it isn't merely tidier, an unmatched path is entirely
attacker-chosen, so labelling by it would let anyone mint rows.

**The funnel counts people, not clicks.** A player who starts six adventures
is one person who started an adventure. That is the whole reason the
per-visitor-day table exists; its flags only ever turn on, and `is_new` is
settled by the first write of a visitor's first day.

**The tests run on SQLite and production is Neon.** A flush that raises is
caught and logged, so a dialect mistake in the UPSERTs would have stayed
invisible until the dashboard quietly never filled.
`test_the_upserts_compile_for_postgres` compiles both statements against the
Postgres dialect without connecting to one.

Two things this leans on elsewhere. `limits._client_ip` is now public
`client_ip`: the access log needs the same answer, and two functions both
deciding which hop is the caller's is how one of them ends up trusting a
header it shouldn't. And the cleanup sweeper now starts if *either* job has
work — a deployment can keep every guest forever and still want its
visitor-day rows aged out.

No migration. Both tables are new and `bootstrap()` calls `create_all` on
existing databases too, the route `branches` took in Phase 14, so
`LATEST_VERSION` is still 64.

497 tests green, frontend lint and build clean, driven by hand against a
synthetic 90-day fixture at 1568px. The narrow-screen layout follows the
existing 720px block but is unverified: `resize_window` is ignored on a
maximized Chrome and `frame-ancestors 'none'` rules out checking it in a sized
iframe. Also repaired here: a rename in test_ratelimit_hardening.py had run
through the test names themselves, leaving `testclient_ip_*` — still collected
by pytest, which is why it passed unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-22 16:24:42 +05:30
parththakkar106andClaude Opus 5 40d2555f84 Say what the app is now, everywhere it is published
The README, the project page and the engineering guide all describe a linear
story. The tree shipped two days ago. Every published surface is a phase
behind, and the guide is not merely behind — it is wrong in a way that costs a
reader time.

Its 2.2 was "Two coordinate systems, and the bug class they create", and it
explained the codebase through position_of_index, note_action_removed and
settled_story_actions. All three were deleted in SP3. 2.3 explained retry
through Action.variants and state_before. Somebody reading either would go
looking for machinery that is not there, which is worse than a gap.

So 2.2 is now "The story is a tree", written at the depth 1.2 and 1.3 are
written at: the seven bugs that turned out to be one bug, the lineage clause
and the two properties that make fork count free, why takes group by parent_id
rather than by coordinate, cursors becoming anchors, and a closing list of what
the design is honest about. 2.3 is rewritten around state_after and takes, and
1.1 and 1.5 follow, because the pipeline no longer snapshots before the call
and the memory bank no longer holds an action back.

The numbers were simply old: 151 tests where there are 440, 37 migrations where
there are 64, twelve phases where there are fourteen. They appear in four
places across the README, the project page's stat tiles and the guide's results
table. The measured branch cost — 103 B, and 1.007x the page load of the same
story flat — is added beside the egress and turn-cost figures it belongs with,
since it is the number that answers "what does branching cost me".

Three screenshots, on a new tools/shots_fixture.py: the Bandit Camp demo driven
through eight written turns with written deltas, three discarded takes forked
onto branches of their own, one off a branch so the map has to nest. Same
reason tree_fixture.py is committed — the shots have to be reproducible and the
frontend still has no test runner. play-world-state.jpg is reshot because it
predates the entire tree UI; the map and the branches panel are new.

Note for next time: docs/guide.html is hand-written, not generated from the
Markdown, so every guide edit is two edits in two vocabularies. Both files were
checked for tag balance and both pages rendered locally before this landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-20 04:19:53 +05:30
Claude 8857757642 Route the README's live-demo link through the project page
The hero linked at the Render hostname directly. Pointing it at the project
page keeps the deploy URL out of the README and folds in the "prefer a tour
first?" line, which pointed at the same place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PoBAfwRzHozF2bhZumxjPk
2026-08-12 18:57:51 +00:00
Claude a0d4db7661 Document the engine's design decisions
The README says what the project does; nothing said why any of it works the
way it does. This adds design notes written to be read end to end: each
section states the decision, the reasoning, and what it cost.

docs/GUIDE.md is the readable source. docs/guide.html is the same material as
a self-contained reading page for the project site — no webfonts, no scripts
beyond a progress rail, so it also works saved to disk and opened offline.

Covers the context budget allocator, the propose-and-referee world-state
engine, the measured length-hint result, the memory bank's settled-action and
cursor rules, the two coordinate systems behind the summarization bugs, the
retry variant machinery, the egress fix, and the demo-key pinning. Closes with
the measured numbers, the known limitations, and a pointer to the cleanup
backlog in self-review.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PoBAfwRzHozF2bhZumxjPk
2026-08-11 15:30:22 +00:00
parththakkar106andClaude Opus 5 93204ce1b0 Add a project page so the resume link loads instantly
The demo sleeps on Render's free tier, so a cold link looks broken to
anyone who won't wait 30-60s. A static page on GitHub Pages is never
asleep: it shows the screenshots immediately and sets the expectation
before the visitor clicks through to the demo.

Served from main:/docs, reusing the screenshots already committed there.
Palette and type match the app so the two read as one product.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015sygnUWH8WnaoKPDp7hXM1
2026-08-10 19:21:26 +05:30
parththakkar106andClaude Opus 5 ebb81b8f9a Make the repo legible to someone seeing it for the first time
The README claimed screenshots were "coming" and stopped describing the
project around phase 10, so the world-state engine, the retry/undo state
revert and the egress work — the most interesting parts — were invisible.

- Screenshots of the play screen, Insights, the schema editor and the
  script editor, captured from the running app.
- Document the world-state engine, undo/retry rewind, and the two measured
  performance fixes; correct the context assembly order to match
  context/builder.py; drop the stale "phases 1-6 complete" note.
- Add CI (backend pytest, frontend lint + build, Docker image build) and
  its badge. Nothing ran the 151 tests but me.
- Move CODE_REVIEW_FINDINGS.md to docs/self-review.md and say up front
  that every correctness finding is resolved.
- Drop a dead import that the newly-wired lint flagged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015sygnUWH8WnaoKPDp7hXM1
2026-08-10 18:52:02 +05:30
parththakkar106andClaude Opus 4.8 82fb6c8535 Reasoning-model UX: better empty-response error, higher default budget, live link
- When a model streams reasoning but no story text, return a specific,
  actionable error (raise Max output tokens / set a reasoning cap / use a
  non-reasoning model) instead of the opaque "empty response".
- Bump default max_output_tokens 400 -> 800: 400 truncated scenes and
  left reasoning models with no room after thinking. Affects new settings
  rows; existing users keep their value.
- README: add a "Try it live" link to the Render deployment.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
2026-07-20 17:05:29 +05:30
parththakkar106andClaude Opus 4.8 68c165585c Phase 10: Render deploy blueprint
- render.yaml: single Docker web service (SPA + API same-origin), free tier,
  Neon Postgres via AIDND_DATABASE_URL, generated AIDND_SECRET_KEY,
  /api/health check, us-east region, auto-deploy on main.
- README: "Deploy (Render)" section.
- Record verified Postgres path (real Neon, PG 18.4) and Phase 10 decisions
  (Neon, free tier, cloud Docker build) in plan/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017e6tQuojBLYPetUfmhit4X
2026-07-07 14:07:08 +05:30
parththakkar106andClaude Fable 5 de4db373f2 Phase 8: optional accounts, per-user data, shared demo key
Guest-first multi-user mode behind AIDND_MULTI_USER (local installs
unchanged): signed-cookie guest sessions bootstrapped by /api/auth/me,
register upgrades the guest in place, login/logout, per-IP rate limits.
Every router scoped by user_id; Settings become per-user with the API
key Fernet-encrypted at rest and write-only through the API. Users
without a key get a server-funded demo key (OpenRouter free models,
20 turns/day, memory bank disabled on demo turns). Public read-only
demo scenarios (seed_demo.py); debug log restricted to local mode.
Frontend: auth modal + guest nudge, 401 re-establish/retry, demo
banner and key management in Settings.

Migrations 13-23 adopt existing data under a local user and encrypt
stored keys. Verified: migration on a copy of real data.db, two-session
isolation + register/login via curl and Chrome, demo cap 429, live
OpenRouter turn through the encrypted-key path, vite build + oxlint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 23:04:03 +05:30
parththakkar106andClaude Fable 5 22630c8411 Phase 7: MIT license, portfolio README, Docker, cross-platform start
- LICENSE: MIT
- README: rewritten as portfolio-grade docs (features with code pointers,
  architecture, Docker/Windows/macOS quick starts, BYO-model table)
- Dockerfile (3-stage: SPA build, pip wheels, slim runtime) + compose with
  a /data volume; .dockerignore keeps secrets and local data out
- database.py: AIDND_DB_PATH env override so deployments can relocate the
  SQLite file; documented in backend/.env.example
- start.sh: macOS/Linux dev script with first-run setup

Verified locally: production SPA build served by the backend (deep links
OK), DB created at the override path. Docker image itself untested here
(Docker not installed); flagged in plan/07.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:51:48 +05:30
parththakkar106andClaude Fable 5 db9f904222 Initial commit: AI Dungeon clone (FastAPI backend + React frontend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KFsGHju9szibJJa2YJcdbg
2026-07-06 16:08:19 +05:30