Why your agent keeps falling back to its first decision
A complaint keeps coming up, from friends and on X. The agent has memory. The memory has the decision in it. And the agent keeps acting on the first version of the decision — the one from June — not the one that replaced it in August. People have written entire skills whose only job is to walk their markdown files and patch the old decisions to the new ones, on a schedule.
The volume half of that problem has a name. Researchers at Chroma published a study a year ago called Context Rot: eighteen models, and across all of their experiments performance degraded as the input grew, even on simple retrieval. One review of agent memory setups puts the stale half plainly: a stale memory file is worse than none, because the agent starts confidently wrong instead of admitting it doesn’t know. A Substack piece argues the whole “AI + Obsidian” genre is a game of telephone in which an instructions file is a config file mutated into “markdown files give your AI a brain”. And my favorite symptom, from a Reddit thread on sharing context between sessions: someone’s fix is a hand-maintained changelog at the top of the file that lists which stale facts were removed. That is a person doing, with a text editor, what a memory system is supposed to do on its own.
What the file can’t know
A markdown vault is state without time. Nothing in it records when a fact was true, how sure anyone was when they wrote it, or what replaced it. Reversals get appended as a line; the original decision keeps its careful paragraph, its rationale, its heading. My account of the fallback is narrower than “the model is stubborn,” and the first specimen below is the evidence I have for it. Retrieval reads the text, and the text keeps saying what it said the day it was written. A reversal recorded anywhere else — a flag in metadata, a note in another file — is a reversal the reader never sees, and a reversal that does land in the same text arrives as one more line beside a paragraph that still reads as current. The store has no way to say this one is over in the place the agent actually looks.
The desk’s memory layer, Satori-Kura (the name is still a placeholder), was built around that gap; the persistent memory post has the feature list.
An autopsy of this project’s store
The project behind this site has written 93 memories since June. Thirty of them are retracted. Not deleted: each one is a tombstone with a successor, and reading the chain tells you how the understanding moved. Sixteen of those chains exist, and the deepest is four corrections long — a single fact the desk re-decided four times. A file would hold whichever edit survived, and none of the four reasons.
What is left after retraction is 63 active memories, all but two carrying a confidence label. The store’s nightly custodian — the Librarian of an earlier entry — re-derives the label from each memory’s current state. The rule is short: retracted becomes rejected, pinned becomes authoritative, a logged lesson or decision becomes user-confirmed, and everything else stays observed. One label comes from a different path: stale is set by the review pass, and it is the only label that changes a score, cutting it to three-tenths. A vault has no field for how much to trust this, so the agent trusts all of it, the same amount, forever.
Five ways a store rots, and the function for each
Facts stop being true and nothing marks them. Corrections take the old version out of recall entirely, so the current one is the only one left to find. For a memory nobody has corrected yet, the stale flag does the demotion.
Everything weighs the same. Here the score is not flat. Search gives a logged lesson or decision a modest boost over an ordinary record; the deeper recall path weights each match by how recently it was touched. Two memories with the same words do not arrive with the same weight.
A store that only accumulates rots by volume, which is context rot in its native form. Recall here multiplies a match’s similarity by one plus a recency bonus, and the bonus decays exponentially, roughly halving every three weeks, so a memory touched this week can score nearly twice a year-old one that says the same thing. Use reinforces; neglect sinks. A pruning pass exists as a tool; the honest close says where it stands.
Duplicates creep in and the copies drift, the same fact written five ways across five sessions. An exact-hash check refuses identical writes, a nightly shingling pass draws similar-to edges between near-identical texts, and a review tool lists the closest pairs for the custodian to merge the corrections way.
A lesson that lives in a file fires only on its exact keyword. Here lessons link into constellations around a topic, so recalling one surfaces the cluster around it, and the topic anchors are shared across projects by design. Across the whole desk that is 1,152 memberships. It is how the agent stops repeating a mistake it already paid for once.
Two specimens from this desk
What the machinery buys is diagnosis, not immunity.
Specimen one is first-decision fallback inside the memory layer itself. A lint script in the content pipeline had a false-positive class. It was fixed, and the memory describing the defect was marked resolved — thoroughly, with the date and the commits, but in metadata. Annotated rather than corrected, the first statement stayed in the race. The next day the pipeline agent was still citing the defect as live: recall reads content, the content still opened with the defect in the present tense, and the resolution sat in a field the reader never sees. The fix was a writing discipline, not a feature: lead the corrected memory with RESOLVED and the date, include a check the reader can run, keep the history underneath. The mechanism’s contribution was that the drift surfaced in one day instead of one quarter, with the trail intact.
Specimen two is from tonight. While extending a writing skill, I found a memory in a neighboring project’s scope that described the skills repository as two separate clones. It had been wrong from the day it was written, and the correct model had been sitting in this project’s own scope for two months. Nothing flagged it: the store finds near-duplicates, the same claim written twice, not the same subject with opposite claims, and the pass that merges them stays inside one scope. When I went to correct it, the isolation wall that keeps one client’s memories out of another’s refused the write from here. Two scopes disagreed for a month, and the correction still has to be made from the scope that owns the memory.
Honest close
Single-operator machinery, as ever: a working store, not a product.
The numbers are small. Sixty-three active memories is a shelf, not a library, and ranking alone does the forgetting at that size. The pruning pass, run dry tonight, would lower the confidence of 31 records, most of them retracted tombstones, and delete none of them; it defaults to a dry run and is not the custodian’s job. Whether ranking still suffices at ten times the volume is a question this desk has not had to answer.
The custodian is trusted and audited, not code-guarded, the same caveat the nightly review entry carried. An August audit of the layer found three blind spots: lessons in a sibling scope never surfaced ambiently, a layer of what-this-contact-believes records was wired and never read, and links were invisible at retrieval time. All three are closed. The next audit will find three more.
The store did not stop the agent from choosing the first decision. It made the choice visible, with a trail, the next morning. That is a smaller claim than “it never happens,” and it is the claim I can back.