My agent's memory disagreed with itself
The hub of this series named four places design intent can live and left one unwritten. Memory: what the agent learned yesterday, written down for tomorrow.
My agent published a post nobody had reviewed. It was following a note I wrote, and contradicted twice afterwards. Nothing noticed the three disagreed.
TL;DR:
- Three notes answer one question, and they disagree. May the agent publish? All three load every session, and nothing flags the conflict.
- The note that says never publish is filed under a name that says always. Nothing checks that a name, an index line and a body agree.
- 94 memory notes for this repository, 49 of them corrections to behaviour. Every failure here is in that half.
- Everything else in this repo is enforced by code. 100 tests, 16 lint rules, six commit guards. Memory was the one layer that only advises.
- What worked was a check. The publish decision became a date the build reads, not a note the agent weighs.
Everything here is checked, except memory
This blog is written with an agent and shipped by one. The posts, the LinkedIn hooks, the newsletter: a git push puts all of it in front of people.
So the repository is not short of checks. A hundred tests run on every push. Sixteen lint rules read the posts themselves: a long sentence, an em-dash, a hook that lost its link. Six guards run before every commit, and the deploy gate reads the built HTML.
Every one of those is code. The rule about when a post may go public was prose in a memory note, weighed against whatever I had just asked for. There were three such notes.
So it published, and nothing refused, because nothing had been asked to.
A note written after an incident records that incident’s side of the argument.
The note is named wrongly now
The never-push note is still filed as always-push-blog-repo. It was named when the rule was the opposite, then reversed in place. The index line is right. The filename is not, and the filename is what six other notes link to.
Anthropic’s documentation describes the layout: an index line per note, loaded every session, plus one file per note. So one memory has three parts, written at three different times, and nothing checks they agree.
Written most, followed least
For this repository, June to September 2026: 94 notes on disk. By their own tags, 49 correct behaviour and 31 record project state.
Every failure above comes from the behaviour half. The facts half has none I can point to. Which port the browser listens on is a fact the agent looks up. Never push is a judgement, weighed against whatever you just asked for.
The notes are not written by hand. Every night a job reads back the day’s sessions and proposes what memory should learn. Since 4 August: 92 proposals, 85 taken, 7 turned down.
That is half the loop Google’s SRE book describes for people, in its chapter on blameless postmortems. Write down what happened, without blame, then follow the actions through. Mine only writes. It records a lesson well and has no way to make one stick.
A note is advice the agent may take. A check is a decision it cannot make.
The obvious answer is to write less of it. That is the one thing memory will not let you do.
The diet that could not touch memory
Everything the agent reads before I type anything came to 123,039 characters on 4 August. So the rules went on a diet.

The diet took 83% off the repository's rules and 4% off the memory index.
The rules lost 83%, moved into files that load only when asked for. The memory index lost four lines and stopped, because an index line is both the summary and the key that finds the note.
Anthropic calls what you pay for the extra text context rot: the more is loaded, the worse the recall of any one line.
What I cannot rule out
- One operator, one machine, 14 weeks. Counts, not rates, and nothing memory-free to compare against.
- I never counted the failures the notes prevented. A note that works leaves no trace.
- The check is young. A few weeks without an incident is not evidence.
What I would do on Monday
- Tag every note as fact or behaviour. Review only the behaviour half, monthly. That is the half that fails.
- Count repeats, not notes. The second time you write a lesson, link it to the first. The third time, build a check instead.
- When two notes disagree, the tie-breaker is a check. A third note joins the argument. A check ends it.
- Sweep memory on a schedule, before it asks you to. Read the behaviour notes as a set each month and look for pairs that pull against each other. Contradictions never announce themselves. They pile up quietly and surface as one wrong decision, months later.
- Check that a note’s name, index line and body agree, automatically, on every change.
- Move any decision with an outside consequence out of memory. Publishing, sending, spending, deleting: each becomes a field the build reads.
Conclusion
Memory recorded what I told it, faithfully, including the three times I contradicted myself in a week. A better note would not have fixed that. Moving the decision into the build did.
What would change my mind: a post of mine going public without a dated decision before 2027. Then the check was not the fix either.
This closes the series. The hub has the four places side by side. Next: the metric I never took, and why it could not be backfilled.
Methodology & limitations (click to expand)
Prior art
- How Claude remembers your project, Anthropic: both memory systems load every conversation as “context, not enforced configuration”, a hook being the way to block an action anyway.
- Effective context engineering for AI agents, Anthropic, 2025: context as “a finite resource with diminishing marginal returns”.
- Postmortem Culture, Google SRE book, chapter 15.
Data sources
- This repository’s memory folder, counted 14 September 2026: 94 note files plus a 96-line index; by front-matter type 49 feedback, 31 project, 9 reference, 5 other. Only the index loads unasked.
- The enforcement counts, read 22 September 2026: 100 tests across 7 files, 16 rules in
scripts/markdownlint-rules/, 6 guards in.githooks/pre-commit. - The diet, measured 4 August 2026 and recorded in this repository’s own playbook, in characters. The chart and its figures come from
diagrams/di4-context-diet.py. - The nightly loop’s decision log, 4 August to 14 September 2026.
What this does not show
- No archive receipt. The platform the rest of this series measured had no memory tier.
- The hub’s “grows without bound” was right about the files and wrong about the bill.
Get new posts by email
One email per post, about two a month. The numbers and the caveats, same as here. No sequence, no pitch, unsubscribe in one click.
Almost there. Check your inbox and click the link to confirm.
That did not go through. Try again, or email [email protected] and I will add you by hand.
No tracking pixels. I never pass the address on. How this is handled.