c8755c21c2bd9399a21e70413e7784d19b0d7df7
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c8755c21c2 |
Planning: close M3 and record active-head architecture
M3's review recommended planning changes and, following the M2 pattern, reported rather than applied them. This applies them, and adds the ADR the review asked for. ADR 012 records the architecture rather than the requirement. ADR 005 already says that going backward must preserve abandoned history and that the user sees Undo/Redo/Retry rather than branch management; it names a movable active head as the direction and stops. What M3 settled is the shape: the head is stored rather than derived, every read of the story is capped at it in one place, one mechanism moves it, the state of a position comes off the node rather than from a replay, the first write below a moved-back head is the divergence, and whether Redo exists is decided by the lineage rather than by a flag that could be stale. The last of those is the property worth keeping — a flag can be wrong and make the story wrong; a lineage cannot. Two semantics are ratified in STORY-BRANCH-SEMANTICS.md, both of them reversals or narrowings that a reader would otherwise take for bugs. Undo now crosses fork points and continues to the campaign opening, because refusing at the fork was a consequence of deleting rows the parent line was also reading, and nothing is deleted any more. And the system refuses to switch which take is live while a later story is off screen, because doing it quietly would leave retained history continuing from words the story no longer says. A new §14A covers editing in place. §14-15 describe the finished behaviour — the edit becomes authoritative, the state it implies is re-evaluated, a new continuation is created, the original is retained — and that requirement is intact and explicitly not weakened here. It is also not built, because re-evaluating state from prose a user typed needs M5's extraction pass. §14A says what exists in the meantime and why refusing is the minimum that holds the invariant rather than the destination. TECHNICAL-DESIGN.md gains §8.7 and §9.1, recording the implemented model and the bundle behaviour as fact in the way §5.2 records M1 and M2. §10.4 gains a constraint that is easy to lose: the snapshot half of the hybrid state model is a requirement, not an optimization. Head movement is a row lookup plus a restore, which is why Undo, Redo and Save Point restore cost the same at any distance into a campaign; a state model recoverable only by replaying from the opening would make all three proportional to campaign length, on exactly the long campaigns this product is for. DATA-MODEL.md records the head as stored on the campaign rather than derived from its newest turn — two campaigns holding identical turns can be read at different places, and nothing about the turns can tell them apart — and the branch disposition as implemented: the depth a divergent write left the branch at, deliberately advisory, and carried through export because every row of an abandoned line is exported either way. BUILD-MILESTONES.md marks M3 complete and states the one condition still open. M4 is told a Save Point is a durable pointer and that restoring one is head movement with a bounds check, not a restore system: a second mover is the specific failure to avoid, because the two paths would silently disagree about what restore means. M5 gets three constraints — keep state efficiently recoverable, move the test instrumentation rather than the assertions when the world-state protocol goes, and finish the narrator edit §14A defers. V1-ACCEPTANCE-TESTS.md clarifies ownership without lowering a bar. D10 keeps all three pass conditions and is explicitly recorded as *not* satisfied at the end of M3; what changed is that the document now says which milestone delivers which condition. D03's result is recorded as a full pass rather than the partial the text allowed for, I07 gains the pre-M3 bundle clause, and L01 gains the note that resolves its apparent conflict with A05 — a failed turn does advance the head by one, onto the player's retained input, and that is A05 working rather than L01 failing. README.md described a different application: a hosted demo, guest accounts, cloud providers, Postgres, a Render blueprint, an analytics dashboard, a QuickJS scripting engine, and 549 tests. M2 removed all of that and the README was never updated — a gap M2's own debt table missed. It now describes what this fork is, including the endpoint policy and the TLS behaviour, and the numbers in it are the current ones. M3's report is included here as its own evidence record: no separate baseline report was produced, so it carries the raw counts and runtime observations as well as the review, and §W records this closeout. SPECIFICATION.md and SECURITY-THREAT-MODEL.md are unchanged. M3 altered no product requirement and touched no path in the threat model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u |
||
|
|
7f082b61d8 |
M3: complete non-destructive history and active-head export
The head-cursor model landed in
|
||
|
|
903fa7a74f |
M3: move the story's head instead of deleting its turns
Undo deleted. It removed the trailing AI action and the player action in front of it, pruned the memories covering them, and let the tip fall back to whatever survived. That made it the one operation in the application that destroyed accepted story, and it was why there was no Redo: the turns to move forward into no longer existed. Phase 0B demonstrated the head-cursor alternative in a disposable spike; this is that concept as production code. backend/app/head.py is the whole of it. Three questions that used to be one — where the story is being read, how far it is retained, and where it opens — are now three functions, and every caller that moves the head or asks about it goes through this module. The spike put the fork check in the write path and left Retry and Add-take on the old one; sharing the rules is what stops that divergence coming back. lineage.Path now caps every entry at the head, so hiding the retained future costs nothing at the call sites: the transcript, the context builder, attempts.preceding and memory retrieval already funnelled through path_of and narrow together. Path.uncapped() is the deliberate exception, and only Redo and the fork check may use it. The memory bank needs no pruning for the same reason — a memory carries the coordinate of the node its block ends on, so one derived past the head falls outside the capped clause and becomes retrievable again on Redo without having been deleted and re-embedded. Undo alone does not fork. Moving the head is not a decision to abandon anything, since the user may be reading or about to Redo; the first write below the head is where the story states which continuation it means. A head already at the tip forks nothing, so a story that is never undone forks exactly as often as it did before and the branch table does not fill up with one branch per turn. Redo follows the lineage rather than choosing among branches, which is what invalidates it after a divergence with no flag to set or clear. Migrations 78 and 79 give a branch superseded_at and superseded_depth. Nothing reads them to decide behaviour — Redo is decided by the lineage, so a stale or hand-edited value here cannot make the story wrong. They exist so the cleanup and discarded-history features left to a later version have something to select on, and so a divergence is observable in a test. Deleting an action no longer drags a moved-back head forward to the recomputed tip, which would have silently redone the story. can_undo and can_redo ride on AdventureOut and ActionPage because the client can work out neither for itself: the campaign opening may be off the top of the loaded window, and the retained future is never sent to it. This is a checkpoint, not the finished milestone. 601 backend tests pass. Five still assert the destructive contract — they count rows after an undo and expect the story to be shorter — and need rewriting against the new one; the world-state assertions inside them already pass. Export and import do not yet carry the head coordinate, so a bundle still reopens at the deepest node and can silently redo an undone story, which is the Phase 0B finding this milestone exists to close. The browser has no Redo control yet. None of the M3 acceptance coverage (D01-D10, E01-E04, I01-I03, I07, L01-L02) is written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QF5TcoB86QADgjHz1GZe8u |
||
|
|
2fdd2547f0 |
Planning: record M2 closeout decisions
M2's review reported six planning recommendations rather than applying them, three marked before M3. All six are applied here, plus three additions drawn from the same evidence. No implementation file is touched. The endpoint policy was the gap that mattered. It is the most consequential setting in the application — the storyteller sends the player's prose, the context, the memories and the embedding inputs to whatever address it names — and it existed only as a module docstring. It is now ADR 011 and a new §10A in the threat model, which also retires the assumption in §71A that the inherited guard was a starting point. It was not: AI-DnD's SSRF guard blocked private addresses to stop a hosted server reaching its own internal network, which is the exact opposite of what a local storyteller needs. It was removed, not adapted. Both documents state the rule as implemented — an allowlist of explicit local-network CIDRs, every resolved address checked, enforced on save and again before every outbound request, TLS never traded against it — and both state the two residual limits plainly rather than implying they are covered: a hostile host already on the trusted LAN is inside the permitted boundary, and a rebinding interval exists between the policy's resolution and the client's connection. Accepted risks, not M3 work. The CIDRs are spelled out rather than derived from is_private/is_reserved, and the ADR records why: is_private is true of the documentation ranges and 0.0.0.0/8, and is_reserved is true of IPv6 loopback, so a rule built on it refuses an ordinary same-host Ollama on [::1]. TECHNICAL-DESIGN §5.1 items 3 and 4 are marked done, closing all five hardening items. A new §5.2 records the M1/M2 architecture as fact rather than intention, so later milestones inherit what the code does. A new §18.1 carries the lesson of M2's two regressions: when removing a setting, test a real consumer construction path; when adding one, prove it reaches the component that uses it. Both defects hid behind a green suite because the tests at that boundary were mocks. BUILD-MILESTONES records M2 complete, with the capabilities later milestones inherit and the debt carried forward. Two notes go to milestones that would otherwise misread what M2 left them. M5 is told that eight rollback tests now use the world-state engine as instrumentation and not as endorsement — the instrumentation moves when the protocol does, and those tests are reworked rather than deleted. M6 is told that the memory bank died silently under a green suite, so background failure must be observable and at least one real provider-construction path must be tested. The security contract gains what M2 demonstrated. H10 now names the two conditions that were defects during M2: a wildcard origin must be refused at startup, and an unknown /api path must 404 rather than returning the SPA with 200. New H12 covers endpoint enforcement, and its fourth pass condition is the one that matters — a public endpoint written into the database behind the settings API must still be refused at the wire. A build passing the first three and failing that one has configuration validation only. SPECIFICATION.md is deliberately unchanged. M2 altered no product requirement; it removed capability the specification never asked for. The two M2 reports gain appended closeout notes rather than edits. Their original wording about an uncommitted working tree was true when written, and the note records what happened afterwards: the six-file correction is |
||
|
|
8652fe7cd8 |
M2 review: two regressions the green suite hid, and the reports
The post-implementation review of M2, plus the three corrections it took
to make the evidence true. Reports:
planning/reports/M2-BASELINE-REPORT.md 868 lines, the measurements
planning/reports/M2-IMPLEMENTATION-REPORT.md 758 lines, the reading of them
Verdict is PASS, accept with non-blocking debt, proceed to M3. Every M2
requirement is met and the ones that matter were tested by running the
build rather than reading it: a cloud endpoint written straight into
SQLite with sqlite3, behind the API's back, still refused at the wire;
trusted-LAN HTTPS against the real second machine with verification on;
captures showing zero packets outside loopback and the approved host.
Three defects, all found by running the shipped image.
The memory bank was dead. M2 removed Settings.api_key_plain with the API
key, and memorybank's two provider factories still read it. It failed
inside a fire-and-forget task, so no user error, no log anyone would
read, and no test — every memory test stubs those factories. All 604
tests passed with summaries and embeddings silently not happening.
The configurable model timeout never reached the turn engine. Stored,
validated, exposed in the API, rendered in the UI, and not passed to the
provider. M2's own exit criterion was half met: the constant had moved
but the setting did nothing.
And requirements.lock still pinned quickjs, psycopg and cryptography, so
the setup path DEVELOPMENT.md gives a new developer would have
reinstalled all three.
Both code defects now have the test that would have caught them: one
constructs every provider factory from a real Settings row, one drives
the turn endpoint, the chat endpoint and the summariser and asserts the
configured timeout arrives at each. That is the lesson worth keeping from
this milestone — after removing an attribute, build each consumer from a
real object; after adding a setting, prove it lands. Both failures were
in background or plumbing paths, which is exactly where a subtractive
change cannot see itself.
606 tests pass, up from 604. Lint, build and image are clean. Every
runtime result in the baseline report came from an image built after
these fixes; the reports say plainly that commit
|
||
|
|
8c65ae99de |
M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The milestone is subtraction, and what is left is the single-user local storyteller the specification describes. Removed in full: campaign scripting and its QuickJS sandbox; multi-user accounts, guest sessions, login, registration and the shared demo key; the visitor-analytics tables, dashboard and page beacon; the access log of sign-ins, addresses and devices; per-IP and per-user rate limiting and quotas; Render deployment config; Postgres and psycopg; cloud inference providers, the API-key field and the key encryption that existed to store it; session-cookie signing. None of it was hidden behind a flag — the routes are gone and answer 404. Two things were kept that the brief allowed keeping. The `users` table and its foreign keys stay as an internal ownership detail, because rewriting them out means a migration across most of the schema to delete a column that costs nothing; nothing creates a second user and no request carries an identity. Five inert tables and four inert columns stay for the same reason, so an M1 campaign database opens unchanged. The one addition is app/endpoints.py, which decides where a story may be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an explicit allowlist of networks, not a guess at what `ipaddress` means by "private", which calls the documentation ranges private and IPv6 loopback reserved. Every address a hostname resolves to must be in it, so a split answer does not squeak through, and the rule runs both when the endpoint is saved and before every outbound request, because a name that resolved to the LAN this morning can resolve elsewhere this afternoon. Known cloud hosts are named in the refusal so the error says why rather than looking like broken DNS. TLS is never traded against it: M1's shared trust context is intact on all four clients and there is no way to skip verification. The hardcoded 120-second model timeout is now a setting. That was not theoretical — on this GPU-less four-core host a cold load of qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short at 10s so a wrong address still fails fast; the read timeout defaults to 300s and is bounded at 3600, because "wait longer" must stay a number. Two defects found while testing and fixed here. An unknown /api path fell through the SPA catch-all and came back as HTML with status 200, so a client asking for JSON parsed a web page instead of learning the route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an unauthenticated loopback API would hand every page on the Internet a write handle on the campaign database; it now refuses to start. Verified rather than assumed. Offline, on a network with no route out and no DNS: five turns, retry with both takes retained, restart with an identical transcript digest, a failed model call leaving the accepted AI-turn count untouched, and a capture with zero non-loopback unicast packets. Against a real second machine on the LAN over HTTPS with a private CA: four turns, restart, and a capture showing 289 packets to the approved host, 344 loopback, zero anywhere else, zero DNS queries. Cloud and public endpoints refused with their reasons; no API key settable; every removed route 404. 604 backend tests pass, down from 648 by the fifteen retired with the subsystems they tested and up by the twenty-nine added for the endpoint policy and the removed surface. The scripting tests were not deleted: eight files used a JavaScript counter as instrumentation for the state snapshot and rollback machinery, which M2 does not touch, so the counter moved to the world-state engine and those tests still assert what they always did. Frontend lint and build are clean; the image builds, and its wheel-building stage is gone with quickjs. No M3 work. Undo is still destructive and there is still no Redo. |
||
|
|
1a28a9a708 | Apply post-M1 corrections to the planning package | ||
|
|
645f07f06d | Add the M1 implementation review report | ||
|
|
c1a73b3d77 |
M1: make the first story turn work with no Internet
Phase 0B ran the upstream application on a network with no route out and the first turn died in tiktoken, which downloads its BPE table the first time anything counts a token. The browser separately fetched three font families from Google on every page load. Neither is visible on a machine that has been online once, which is why both now have tests. The tokenizer table is vendored at backend/app/context/vendor/cl100k_base.tiktoken and backend/app/context/encoding.py builds the encoding from it directly, verifying its SHA-256 against the digest tiktoken itself pins for that URL. No code path in the tokenizer can reach the network any more — not a warm cache, not an environment variable a deployment could forget. The encoding was checked token for token against tiktoken's own. The three font families are self-hosted as variable fonts under frontend/public/fonts/ (343 KiB, Latin and Latin Extended), declared in frontend/src/styles/fonts.css, and re-vendored by frontend/tools/vendor_fonts.py. Their OFL licences ship beside them. With no remote asset left, the CSP drops both Google hosts and gains object-src, base-uri and form-action; woff2 also gets its real media type, which Python's table lacks on a slim image. A trusted-LAN Ollama turned out not to work at all over HTTPS. httpx verifies against the certifi bundle, so an endpoint whose certificate comes from a CA the user installed on their own machines — a StartOS server's Ollama, for one — was refused with CERTIFICATE_VERIFY_FAILED while curl and the browser on the same host accepted it. app/tlstrust.py builds one context that unions the platform CA store with certifi's, and all four outbound clients use it. A union rather than a swap, so an image with an empty system store cannot start failing on endpoints that worked before. Verification itself is untouched: CERT_REQUIRED, hostname checking on, and no insecure escape hatch. The storyteller listener is now loopback by explicit statement rather than by inheriting uvicorn's default: start.sh, start.ps1, and docker-compose.yml, which publishes to 127.0.0.1 rather than every interface. Reaching an Ollama on another machine is outbound and needs none of that inbound exposure. backend/requirements.lock pins the exact tested closure; requirements.txt keeps the ranges. DEVELOPMENT.md covers setup, the same-host and trusted-LAN Ollama configurations, and how to re-run the offline proof. PROVENANCE.md records the upstream commit, the MIT terms, and both vendored assets. Verified, not just compiled. On an --internal Docker network with 1.1.1.1 unreachable and no name resolving, a campaign was created and played for six turns through same-host Ollama, restarted, and resumed. A second run played ten turns through Ollama on a separate physical machine on the LAN over verified HTTPS, summaries and embeddings included, with the storyteller's default route deleted so the LAN was reachable and the Internet was not. Its capture: 893 packets to the approved host, 730 loopback, zero anywhere else, and zero DNS queries. Two induced model failures left the accepted story bit-identical. The inherited SPA was opened in a browser and a campaign read back from it. Evidence is in planning/reports/M1-BASELINE-REPORT.md, along with the findings that did not belong in this change. 648 backend tests pass, up from the inherited 632; frontend lint and build are clean; the image builds. No M2 work is included: the hosted, cloud, analytics, Postgres and scripting surfaces are untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017foPNqFjAJa2Ngebf5mEfL |
||
|
|
7f182a86e9 |
Fork AI-DnD at d72f7c1 as the production base
Establishes the production fork lineage decided in ADR 009. Upstream
AI-DnD is merged with --allow-unrelated-histories so this repository
carries the real upstream history alongside the planning package, and
future upstream commits can still be fetched and cherry-picked against
matching paths.
Upstream: https://github.com/parththakkar106/AI-DnD
Commit:
|
||
|
|
717670afe0 | Update planning package after Phase 0B | ||
|
|
ba737de9b4 |
Add Phase 0B local validation findings and recommendation
Validates the three finalists by clone, build, test run and live local Ollama inference, then answers the fork question with measurements rather than static review. Recommendation: fork AI-DnD, confidence high. The Phase 0A call holds, but it was wrong that AI-DnD's undo is non-destructive — retry preserves the replaced take, undo hard-deletes it. A follow-up spike fixed that in 3 files (+130/-31): undo now moves a head cursor, redo round-trips, writing below a moved-back head forks and keeps the abandoned line, branch-scoped memory isolation survives, suite 627/632 with all 5 failures asserting the deleted-row behaviour that was replaced. Findings that change the plan: - AI-DnD cannot take a turn air-gapped as shipped; tiktoken fetches its encoding from a CDN. Proven on an internal Docker network, proven fixed by vendoring the file. - ai-adventure needs zero code for Ollama — two config lines — and its turn/head/checkpoint schema is the target model to build to. - Open Dungeon has zero automated tests and a positional summary watermark, making its branch retrofit larger than Phase 0A costed. - The world-state referee takes relative deltas; a 3B model sent absolute values under full context, so a wounded player ended at full health. Validation cannot catch this, so prefer ai-adventure's typed-event vocabulary when generalising narrative state. - Export/import recomputes head depth, so a round-trip silently undoes an undo. Must be fixed alongside the undo work. Docs only; no production code. Working tree from the runs stays untracked under phase0b/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015gUPLuxLs8wypxZPEmccJu |
||
|
|
f011362494 | Add initial planning files from ChatGPT research here |