Give a memory a word ceiling, after measuring one

Ran the memory prompt end to end against a real model as a controlled
A/B: one story generated through the app's own build_context a turn at
a time, then both prompts run over the same blocks, so the story is held
constant and the prompt is the only variable. Fresh process per call, so
neither arm sees the other and the model is never told what is being
tested. The control is the exact prompt from 9cdcb55.

The reported fault reproduced. Two consecutive memories written from one
story minutes apart came back in two different persons — "You crept low
through the mist" and "The player asked Gwen to". With the cast brief
both named Kaelen. The control also inverted who acted on a move whose
player text was "grab her wrist and pull her down", which is the failure
the brief predicts: with no cast there is nothing to say whose wrist
"her wrist" was.

It also found something the prompt review had not. "1-2 plain sentences"
is not a length, and the same model wrote 34 words for one block and 105
for the next. A 105-word memory is a paragraph, and `memory_top_k`
injects five every turn, so the bank's running cost was set by a number
nobody had ever stated.

MEMORY_MAX_WORDS states it, and the prompt now says which details to
keep when trimming: the ones a later scene could turn on. Re-run over
the identical story, the same two blocks came back at 32 and 58 words,
still named, still third person, still carrying the camp map, the
strongbox behind the second tent, and the strap frayed near through.
Variance is the real gain — 34..105 became 32..58.

Overshooting 50 slightly is expected. Models exceed word budgets, which
is why builder.length_hint already carries a buffer for the same reason.

What this does not show: the run used a Claude model, and the app talks
to an OpenAI-compatible endpoint whose weaker models are why
worldstate/parse.py tolerates trailing commas. The prompt is followable
and the brief supplies the missing information; a weaker model is not
proven to comply as well.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NPyQN926gkZTAYfgugcaok
This commit is contained in:
Claude
2026-08-31 15:03:20 +05:30
committed by Parth
parent 1b5e4c56cd
commit 07192767d8
2 changed files with 73 additions and 14 deletions
+62 -13
View File
@@ -5,9 +5,8 @@ along with the cast, to fix the memories. Phase 1 is worth shipping on its own;
Phase 2 depends on it and is much smaller once it lands.
**Both phases are built and green (610 backend tests). Phase 1 was driven in a
browser (21/21 checks). Phase 2's prompts were verified against the real seeded
scenario — see "Phase 2 as built" — but no memory has been generated by a real
model yet, because that needs an API key.**
browser (21/21 checks). Phase 2 was run end to end against a real model, as a
controlled A/B on one story — see "Run with a real model".**
**Last updated: 2026-08-31.**
@@ -459,15 +458,65 @@ person and named. Embeddings handle paraphrase well, so this is speculative.
If retrieval quality visibly dips, prepending the same brief to the query text
is the first thing to try.
## Still to check with a real model
## Run with a real model, and what it changed
Everything above is the prompt, not the output. Nobody has yet run a real
provider over it and read the memories that come back. That needs an API key,
and it is the only way to know whether the framing rule actually holds across a
whole adventure. What to look for:
Run end to end against a real model, using the `claude` CLI as the provider: one
story generated through the app's own `build_context` a turn at a time, then
**both** memory prompts run over the **same** blocks, so the story is held
constant and the prompt is the only variable. Every call is a fresh process, so
no arm can see the other's output and the model is never told what is being
tested. The control is the exact prompt from commit `9cdcb55`, not a paraphrase.
1. Every new memory names the protagonist and uses no bare pronouns.
2. The framing is the same across memories written many turns apart.
3. The rewritten story summary inherits it.
4. Memories do not get noticeably longer — the brief is context, not content to
be repeated back.
### The reported fault reproduced, and the fix held
| | words | framing |
|---|---|---|
| memory 1, before | 34 | second person — "**You** crept low through the mist…" |
| memory 2, before | 105 | third person — "**The player** asked Gwen to…" |
| memory 1, after | 72 | "**Kaelen** and Gwen crouched at dawn…" |
| memory 2, after | 68 | "**Kaelen** and Gwen infiltrated the bandit camp…" |
Two consecutive memories in one bank, written from the same story minutes apart,
in two different persons. That is the complaint, reproduced under controlled
conditions rather than argued from the prompt text.
The control also got a fact wrong that the treatment did not. The player's move
was `grab her wrist and pull her down behind the woodpile`; the before-memory
recorded "The player asked Gwen to grab her wrist and pull her down behind the
woodpile", inverting who acted. One sample, so this is an observation rather
than a claim — but it is the failure mode the brief predicts, since without a
cast there is no way to tell whose wrist "her wrist" is.
### It also found a real problem, which is now fixed
"1-2 plain sentences" is not a length. The same model wrote 34 words for one
block and 105 for the next. A 105-word memory is a paragraph, and
`memory_top_k` injects five of them every turn, so the bank's running cost is
set by a number nobody had ever stated.
`MEMORY_MAX_WORDS = 50` now states it, and the prompt says which details to keep
when trimming: the ones a later scene could turn on. Re-run over the identical
story, the same two blocks came back at **32 and 58 words** (mean 45, down from
70), still naming Kaelen, still third person, and still carrying every
load-bearing fact — the camp map, the strongbox behind the second tent, the
strap frayed near through. Overshooting 50 slightly is expected: models exceed
word budgets, which is why `builder.length_hint` carries a `LENGTH_BUFFER` for
the same reason.
The variance is the real gain. Before: 34 to 105. After: 32 to 58.
### What this does not show
The model behind the run is a Claude model. The app talks to an
OpenAI-compatible endpoint, and the parser in `worldstate/parse.py` exists
because weaker free models emit trailing commas and leading `+`. So this shows
the prompt is followable and that the brief supplies the missing information.
It does not show that a weaker production model complies as well. An explicit
framing rule is usually worth *more* on a weaker model, but that is an
expectation, not a measurement.
Two side observations from the same run, both pre-existing behavior working
correctly: the model sent `{"milestones.strongbox_found": false}`, `apply_delta`
refused it ("a milestone is sticky, so only true is accepted"), the refusal
reached the model through `render_refusals`, and the next turn sent `true`. The
Phase 1 persona also held across all six generated turns.