AI Agent Memory Systems: What Actually Persists Between Sessions
Most "memory" in AI agents is a retrieval layer, not a model capability — here is how the major approaches actually work and where they break.
Jump to 6 sections
Quick answer: Most "AI agent memory" today is not a persistent brain — it's a retrieval layer bolted onto a stateless model. What actually persists between sessions is whatever the tool chooses to write to a database and re-inject as context on the next call, not anything the model itself remembers.
This guide walks through how agent memory actually works under the hood, what the major frameworks store versus discard, and how to tell a genuinely useful memory system from a marketing label. It's written for developers evaluating an agent framework or building their own memory layer, not for people looking for a philosophical take on machine memory.
What Actually Persists Between Sessions?
Nothing persists in the model itself. Every call to a large language model starts from the same weights with no memory of the last conversation. What people call agent memory is a separate storage layer — usually a vector database, a key-value store, or both — that the agent framework reads from and writes to around each model call.
At The Model Drop, we've tested memory layers across a dozen agent frameworks over the past year, and the pattern holds everywhere: persistence is an engineering decision, not a model capability. If a framework doesn't explicitly write a fact to storage, that fact is gone the moment the session ends.
This distinction matters because it changes where the bugs live. A model that "forgets" a user's preference usually isn't a model problem — it's a missing write, a failed retrieval query, or a context window that got trimmed before the relevant memory made it back in.
The Three Memory Types Agent Frameworks Use
Most production agent systems split memory into three categories, borrowed loosely from cognitive science terminology but implemented very differently underneath.
Episodic memory stores records of past interactions or task runs — what the user asked, what the agent did, what the outcome was. This is typically a log or a set of summarized transcripts in a vector store, retrieved by similarity to the current query.
Semantic memory holds discrete facts: a user's name, a stated preference, a project's tech stack. Storing these as structured key-value pairs instead of unstructured text usually beats vector search for exact-fact recall, since similarity search can miss a specific fact buried in a long transcript.
Working memory is the current task's scratchpad — intermediate reasoning, tool outputs, partial results within one active session. This lives in the context window itself and doesn't need external storage, but it's the piece most agent frameworks manage worst under long tool-calling chains.
A framework that only implements episodic memory will feel forgetful about stated preferences. One that only implements semantic memory will forget the nuance of how a past conversation actually went, even if it remembers the facts.
How Agents Decide What to Write to Memory
The write step is where most memory systems actually differ from each other, and it's the part vendors talk about least.
- Write everything, filter on read. Log the full transcript, rely on retrieval to surface the relevant slice later. Simple to build, expensive to run, and prone to retrieving near-duplicate or stale entries.
- Extract-then-write. Run a smaller model pass after each session to pull out discrete facts worth keeping, then write only those. Costs an extra model call per session but keeps the memory store small and precise.
- User-triggered writes. Only store something when the user explicitly says "remember this" or a tool call flags it. Cheapest and most predictable, but misses implicit preferences the user never states directly.
In our own testing, extract-then-write consistently produced the cleanest long-term recall, because it front-loads the cost of deciding what matters instead of pushing that decision onto every future retrieval query.
Where Memory Systems Fail in Production
Three failure patterns show up repeatedly once a memory-enabled agent moves from demo to real usage.
Unbounded memory growth is the most common one. Every session adds more vectors to the store, retrieval gets slower and noisier, and nothing ever gets pruned. Long-horizon agent evaluation research indexed on arXiv (cs.AI) has repeatedly found that recall precision degrades as a memory store grows past a few thousand entries without any decay or summarization step — the agent technically "remembers" more, but retrieves worse.
The second failure is contradiction: an old fact ("prefers dark mode") never gets overwritten when the user states a new one, and both versions sit in the store with equal retrieval weight. Systems that don't version or timestamp facts have no principled way to resolve this at query time.
The third is context budget collapse. Even with a good memory store, an agent still has to fit retrieved memories into a finite context window alongside the current task and tool outputs. IEEE Spectrum has covered this tradeoff in the broader context of long-context model deployment — more retrieved memory competes directly with the working context the model needs for the task at hand, and past a certain point, adding more retrieved memory actively hurts task performance rather than helping it.
Comparing the Common Approaches
The table below summarizes how the three write strategies trade off in practice, based on our own test harness across repeated multi-session agent runs.
| Approach | Setup cost | Per-session cost | Recall precision | Best fit |
|---|---|---|---|---|
| Write everything, filter on read | Low | Low | Low-medium | Prototypes, low-stakes assistants |
| Extract-then-write | Medium | Medium | High | Production agents with real users |
| User-triggered writes | Low | Very low | Medium (misses implicit facts) | Privacy-sensitive or opt-in tools |
If you're evaluating an agent framework for a real product, we've found the honest question isn't "does it have memory" — nearly all of them claim that now. It's "what's the write strategy, and does it prune." For a deeper look at how these frameworks structure agent behavior more broadly, see our rundown of agent frameworks in 2026, and if your agent leans on long documents instead of chat history, our piece on RAG versus long context covers the adjacent retrieval tradeoff in more depth.
Debugging a memory-enabled agent also benefits from the same tracing discipline you'd apply to any tool-calling pipeline — our LLM observability tools roundup covers the instrumentation options worth adopting once an agent ships to real users.
The Bottom Line
Agent memory is a storage and retrieval engineering problem wearing a cognitive-science name. The frameworks that actually hold up under real usage are the ones that decide deliberately what to write, prune what's stale, and budget context carefully at read time — not the ones with the biggest marketing claim about "remembering everything." Before adopting a framework's memory feature, ask to see its write strategy and its pruning policy. If neither has an answer, treat the memory claim as unverified.
The Model Drop tracks agent tooling as it ships, testing memory claims against real multi-session workloads rather than vendor demos.