Every LLM session starts at zero. Anyone who has worked with a coding agent for more than a day knows the feeling: yesterday's decision is gone, yesterday's mistake is about to be repeated, and the first ten minutes of every session are spent re-explaining context.
I have been running a small multi-agent setup for about nine months — one human operator, a coordinating agent, and a few worker agents. Over time I stopped asking "how do I make the model smarter" and started asking a different question: how do I make the system remember? This post describes what I ended up building. It is not a product and not a proposal — it is a description of a working setup, including the parts that failed and one honest null result.
The stack
The system uses four memory layers, each answering a different question:
- Semantic store (SuperMemory): "What did we do before?" — vector search over past sessions, decisions, and lessons.
- Entity graph (Memory MCP): "What do we know?" — entities and observations as an explicit graph, so facts have a place and relations can be traversed.
- Affective layer (Physarum): "How is the operator doing?" — a probabilistic model of operator state (fatigue, focus, trust), because how an agent should respond depends on more than the text of the request.
- Lifecycle layer (Digital Brain): "What gets kept, consolidated, forgotten?" — a process that encodes episodes at session end, consolidates them into semantic knowledge, occasionally reflects, and lets low-value observations decay exponentially.
The lifecycle layer borrows openly from the literature: a write-time surprisal filter (Friston's free-energy work, via the mnemos line of research), an epistemic router that classifies writes as new/update/skip (Dual-Layer Agentic Memory), multi-faceted retrieval that returns current, historical, supporting and conflicting evidence (MemoryLACE), recurrence-based consolidation triggers (RecMem), Hebbian hub distillation (HeLa-Mem), and semantic-aware spreading activation with lateral inhibition (SYNAPSE).
There is a fifth component that turned out to matter more than I expected: procedural memory — a working agreement of 18 rules, each with a literal source quote and a deterministic checker that fails loudly if a rule breaks.
Where this sits in the literature
I want to be precise about novelty, because the field moved fast in 2026. The individual mechanisms are established: several surveys now map this design space (agent-memory taxonomies in arXiv 2602.06052 and 2603.07670, plus ACL Findings' "From Storage to Experience"), and comparable systems exist — mnemos implements five of the same module concepts, SYNAPSE (ACL 2026 Findings) does spreading activation, EverMemOS (ACL 2026) does self-organizing consolidation, and drift-memory runs a much larger 10-layer architecture. My contribution is not invention — it is integration under discipline: the rituals, the provenance tags, the deterministic checks, and the willingness to publish the failures.
Notably, the field's own meta-analyses report that agentic memory systems routinely underperform their theoretical promise — benchmark saturation, judge sensitivity, metric misalignment ("Anatomy of Agentic Memory", arXiv 2602.19320). Which makes honest small numbers worth publishing.
The part that actually makes it work
The layers are less interesting than the discipline around them. Two rituals carry the whole system:
- Session start: a fixed 10-step read protocol. Memory stats, decay candidates, last session brief, surprisal and recurrence counters, semantic recall, entity search, operator state, last retrospective. The next instance does not recall the previous one — it inherits a verified state.
- Session close: a fixed write ritual. Encode episodes, persist them, verify they landed, consolidate, reflect, record co-activations, spread activation, decay, save.
Every observation carries a provenance tag — who wrote it, when, with what confidence. That sounds like bookkeeping. It is actually the difference between knowledge and hearsay.
What I got wrong
The most instructive failure: the encode step returned beautifully formatted episode strings — and wrote nothing. Thirteen episodes from one session were silently lost. The fix was a rule that now sits in the project documentation: a call whose success is reported but whose effect is never verified is not a write. Stored strings are now grepped in the memory file before the close-out counts as done.
A second failure was subtler: the self-reflection step stored its own reasoning trace as high-importance insights — the system was contaminating its own memory with its inner monologue. That required a filter, tests, and a retroactive audit of all stored observations.
And one measurement that didn't go my way: the surprisal gate was supposed to reject most writes as noise. Measured rejection rate: 1.54%, not the 90–95% the motivating claim implied. The gate stays, but the reason it exists is now the honest one — it is a cheap sanity filter, not a compressor.
Does it help? The honest number
I ran an LLM-as-judge evaluation over 97 stored sessions, 50 questions, baseline vs. full memory stack. After three rounds of fixing real bugs the evaluation itself exposed (a retrieval join that actually hurt, d = −0.685), the current result is:
Baseline 0.7601, full system 0.7906 — Cohen's d = +0.087, not significant.
Read plainly: the stack doesn't hurt, helps a little, and the effect is negligible at this sample size. I am publishing that number anyway because it is the truth, and because a null result with honest methodology is worth more to me than an impressive claim I can't defend. The next step is a larger question set and live retrieval instead of simulated.
Numbers, and limits
As of this week: roughly 1,050 entities, about 9,000 observations, 660+ co-activation pairs, 514 tests, and 287 deterministic consistency checks that run at every session start. The affective layer calibrates reasonably (ECE 0.036, AUROC 0.94 on 349 pairs) — on paper.
The limits are real: this is one operator, one setup, no control group. Some of it is probably over-engineered; parts are certainly under-validated. I built it because the forgetting was costing real work every day, and each layer exists because a specific failure demanded it — not because a reference architecture said so.
Why I'm writing this
I'm an independent researcher, not a lab. If you work on agent memory, evaluation, or governance and see something here worth testing, reusing, or falsifying — I'd genuinely like to hear from you. The most useful thing a reader could tell me is where this is wrong.
The code is not public yet. If there's interest, that's a reason to clean it up.