ac465ed867cec2157532ea18a21c092e7dd04a6b
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a6e9c7a32b |
M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against
|
||
|
|
b7005e6fdd |
M5: genre-neutral authoritative narrative state, with review corrections
Replaces AI-DnD's RPG relative-delta world state with the genre-neutral typed
narrative state of ADR 010: explicit, absolute, allowlisted events proposed by
the model, validated by the application, applied to one authoritative document,
and snapshotted per position so restore stays a row read.
This commit includes the corrective pass that followed the independent review
in planning/reports/M5-IMPLEMENTATION-REPORT.md. The invariant it exists to
hold is:
visible active transcript position == stored head == authoritative state
Narrator editing (D10, STORY-BRANCH-SEMANTICS §§14-15)
A narrator edit no longer rewrites a row. It returns to the state before the
turn, takes the reader's exact text as the accepted narration, re-derives the
state that text implies, and becomes a new active continuation — while the
original narration keeps its words, its live flag and its whole future as
retained history. At the tip the correction is another take; with story below
it, it forks. No new history machinery: this is the existing fork/take/head
path with the reader's text in place of a generated reply. The §14A refusal
is therefore gone for narrator turns, and remains only for player input.
Pre-M5 positions
Migration 88 backfills the empty narrative document onto every action written
before M5, and a missing snapshot now restores the empty document instead of
leaving the previous position's state standing. Restoring to an old Save
Point no longer leaves a later position's entities and facts on screen.
Narrator context
Replayed history carries prose only; the machine-readable block is no longer
reconstructed into past turns, where it contradicted the authoritative state
in the same prompt. A fact withdrawn by a manual correction is now named as
no longer true, with the reader's reason, rather than silently dropped.
Also
- state_changes joins the action-list bulk read, removing one query per row.
- Extraction takes only the application's own protocol payload: an ordinary
```json or ```python block in a story survives, and a mangled proposal
still does not reach the reader.
Planning: ADR 013 records the authoritative document shape; §§14-15/14A, D10,
C04 and BUILD-MILESTONES are updated to describe what exists. Debt is recorded
against M8 (scenario editor UX) and M9 (export of the audit trail).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
|
||
|
|
f1bebe18d0 |
Drop the eight legacy columns SP8 left behind
Migrations 66 to 73 drop `actions.index`, `variants`, `variant_index`, `variant_count`, `state_before`, and `world_state_before`, plus `adventures.memory_cursor` and `summary_cursor`. `index` is a keyword in SQLite, so migration 71 quotes it. Nothing outside the migrations read these. `models.py`, `schemas.py`, and `ACTION_LIST_COLUMNS` lose the same eight fields, `Adventure.actions` orders by `id`, and `attempts.renumber`, `context.history.max_action_index`, and `nodes.next_index` are deleted. Two changes keep the migration replayable on a `create_all` database: - `_split_variants_into_siblings` wrote through the live ORM table, so it stopped compiling once migration 66 removed five of its columns. It now writes through `_ACTIONS_AT_60`, a frozen `Table` with its own `MetaData`. - Five data passes read columns these migrations drop. Each now calls `_has_columns` and returns early when the columns are absent. `bootstrap` takes a `through` version so a migration test can stop at the schema it asserts on. 555 tests pass, up from 549. Eight of the new cases assert each column is gone after a real schema-45 database migrates all the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0198qDK3gmgSo7EtQ4GTPqqK |
||
|
|
32cd7c1077 |
Give the tests one setup instead of thirty-five
Every test module carried the same eight-line prologue redirecting the database to a temp file. Only the first one to be imported ever took effect: `app.database` reads `AIDND_DB_PATH` at import and builds `engine` from it once, so by the time the second module ran the engine already existed. The other 34 copies created a temp file that nothing opened and nothing deleted, and leaked one per module per run. `conftest.py` now does it once, which is early enough because pytest imports conftest before any test module. It also deletes the file when the run ends. The tests still share one database, exactly as they already did: each `client` fixture calls `create_all` on setup and `drop_all` on teardown, so no test sees another test's rows. `tests/fakes.py` holds the one `ScriptedProvider`. Nine modules each had a copy, and the copies had drifted into four feature sets, so a test that needed to raise a provider error had to be written in one of the files whose copy supported that. The shared one is the superset. The two `FakeProvider` copies were the same class with a fixed reply, so they use it too. `test_chat.py` keeps its own, which implements `chat` rather than `generate` and records what it was constructed with. An autouse fixture resets the fake's class state between tests, so a stale reply list can no longer reach the next test. 435 lines out of the suite. 549 tests pass. Verified live by sabotage: breaking the shared fake fails 13 tests across four modules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Dix4oGV3njgWRdu7P9t6r |
||
|
|
e7d75c3b05 |
Rewrite Python comments in Google developer documentation style (#12)
* Rewrite comments in Google developer documentation style Rewrite the comments and docstrings across the backend core modules so they read plainly. The previous prose was accurate but dense and figurative, which made it slow to skim. Applies the Google developer documentation style guide: short sentences, active voice, present tense, American spelling, and no metaphors, idioms, or rhetorical asides. Replaces em-dash chains with separate sentences. |
||
|
|
3b9e6b3d50 |
Give the length hint a floor, not just a wall
A ceiling alone is a one-sided instruction, and models read it in opposite directions. A verbose one is held back by it; a terse one has nothing to act on except "write only as much as the moment needs -- a typical turn is much shorter" and collapses to two paragraphs. Same prompt, wildly different turn lengths depending on which model is behind it. State a floor as well, so the guidance is a band. The two bounds are deliberately asymmetric -- "must not exceed" for the wall the endpoint enforces, "should not stop short of" for the floor -- so neither reads as a number to hit, which is the property the earlier A/B says decides whether this hint helps or hurts. "Prefer the lower end" inherits the anti-overshoot job the deleted "much shorter" line was doing, but now with a number under it, so a terse model lands on the floor instead of at forty words. Below MIN_LENGTH_FLOOR_WORDS the floor is dropped and the tight-cap wording is left byte-identical: at a tight cap a short turn is the correct turn, and that phrasing is the one measured to keep the state block alive (0/6 truncations at cap 250 against 2/6 unhinted). So this only moves loose caps. MAX_LENGTH_FLOOR_WORDS keeps the share from demanding 555 words minimum at cap 2400 -- a big cap means long turns are allowed, not compulsory. Shipped without an A/B run, deliberately. Two things to watch live: whether a stated range invites landing mid-range on verbose models (drop the share to ~0.25 if so), and whether the state block still survives -- nothing reads finish_reason yet, so truncation is silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY |
||
|
|
f295893204 |
Ask the model for a turn that fits inside the output cap
max_output_tokens is a hard wall the endpoint enforces mid-sentence. The ```state block is emitted after the narration, so a long turn hits the wall partway through the block and the deltas are lost — silently, since nothing reads finish_reason. builder.length_hint() derives a word limit from the cap ((cap - 50 headroom) * 0.75 words/token * 0.90 buffer) and injects it just above EMIT_REMINDER, which keeps the last slot it needs. Reserved in build_context like the reminder is. Phrased as a ceiling, not a budget. Measured against gemma-4-26b at cap 800, n=5 per arm: no hint 174 words, "keep this turn under about N words" 246, "hard limit ... a typical turn is much shorter" 170. A budget reads as a target to fill — every budget run was longer than every unhinted one, pushing turns toward the wall the hint exists to avoid. Ceiling phrasing still works at tight caps: at 250, unhinted hit finish_reason=length 2/6, hinted 0/6. tests/test_length_hint.py, 11 tests; each mechanism verified by sabotage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet |