Key takeaways
- Long-term memory is what an agent keeps after the conversation ends and uses to change a later decision. Retrieval-augmented generation looks up knowledge the agent could fetch again.
- If the context that changes behavior fits in a few thousand tokens, load a compact profile before building a retrieval system.
- Test the write path, a later session that should use the memory, and a case where remembered context should lose to the current instruction.
Long-term memory sounds simple: an agent learns something in one session and uses it again in the next. But things quickly become challenging when you need to decide what deserves to survive, how to make it available later, and what to do when that information changes.
A user may have already explained how their project is structured or corrected something three sessions ago. For that matter, the agent might have discovered where ground truth lives while debugging a trace.
Those are different kinds of information, but they create the same problem. The useful part of one interaction has to survive after the conversation itself is gone.
Most published memory architectures solve this with retrieval: extract facts from interactions, put them in a vector store or knowledge graph, and search for the relevant ones later. That can work. It also creates another system with its own failure modes. The right memory can be stored and never retrieved. A stale fact can outrank the current one. The agent can even receive the correct memory and fail to use it.

This guide covers the main ways teams are solving those problems, how to choose between them, and what we learned about storing, updating, forgetting, and evaluating memory while building it for our own agent.
Trying to build your own agent memory system?
We built long-term memory for Alyx, the AI engineering agent in Arize AX, using one structured memory file per user and space. Read How we built long-term memory for Alyx with one 8,000-character file for the architecture, update pipeline, compaction logic, and evals behind it.
What is long-term memory for AI agents?
Long-term memory is information an agent keeps between sessions and uses to change what it does later. Short-term memory is the current context window: the messages, tool calls, and results in the conversation happening now. Long-term memory is whatever is still there when the user opens a new chat next week.
The idea is older than LLMs. Endel Tulving proposed the split between episodic memory (specific experiences) and semantic memory (general knowledge) in 1972, and it’s still a core concept in memory research (Renoult and Rugg, Neuropsychologia, 2020). In 2023, the CoALA paper applied that vocabulary to language agents, describing working, episodic, semantic, and procedural memory for LLM-based systems (Sumers et al., arXiv:2309.02427). The same year, MemGPT borrowed from operating systems and treated the context window like RAM, paging information in and out of external storage (Packer et al., arXiv:2310.08560).
Products took longer. OpenAI started testing memory in ChatGPT in February 2024 (TechCrunch) and extended it to reference past chats in April 2025 (TechCrunch). Anthropic launched memory for Claude Team and Enterprise users in September 2025 (Claude blog) and opened it to free users in March 2026 (Claude release notes). What’s new isn’t the concept. It’s that the teams serving memory to the most users have converged on something far simpler than the research diagrams.
Short-term memory vs. long-term memory vs. RAG
Memory gets confusing because several different kinds of context can end up in the same model prompt.
The simplest distinction is to ask where the information came from and how long it needs to survive.
| Context type | What it does | Example |
|---|---|---|
| Short-term or working memory | Keeps context for the current task or session | Recent messages, tool calls, intermediate results |
| Long-term memory | Carries learned information across sessions | Preferences, goals, corrections, previous decisions |
| Retrieval-augmented generation (RAG) | Retrieves external knowledge relevant to the current request | Documentation, policies, source code, product data |
The boundaries can overlap. A long-term memory system may use a vector database for retrieval, while a RAG system may also draw on information from earlier interactions.
That makes it useful to distinguish what job the context is performing. RAG usually helps an agent find something it could look up again. Memory helps it retain something it learned.

For example, an agent can fetch the current configuration of a dataset through a tool and more than likely doesn’t need to memorize that configuration. But if the user explains why that dataset exists or says that every dataset they create should contain exactly 70 rows, that’s harder to reconstruct later and much more useful as memory.
Build better agents with Arize
Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.
Prefer open source? Try Arize Phoenix for self-hosted, open source agent observability.
What does an agent memory system actually have to do?
Whatever you use to store memory, the system still has to solve the same four problems:
- Keep track of the current interaction. The agent needs enough short-term context to complete the task in front of it.
- Decide what should persist. Most conversations contain far more information than the agent should remember permanently.
- Make useful memories available later. That might mean loading a profile directly into context or retrieving a subset from a larger store.
- Update and forget. New information can contradict old information, and memories that were once useful can become stale.
That makes memory a loop more than a database. The agent interacts with the user, learns something, decides whether it should persist, and may use or revise that information in a later session.

You can see that separation in production memory systems:
- Amazon Bedrock AgentCore keeps raw session events as short-term memory, then asynchronously extracts and consolidates long-term records that can be retrieved semantically.
- Redis Agent Memory follows a similar pattern: session events are stored first, while summarization and long-term extraction happen in the background.
How most of the industry talks about building agent memory systems
Most published guidance treats memory as a retrieval problem. You extract memories from conversations, index them in a vector store, a knowledge graph, or both, and search for relevant ones at query time.
The guidance is consistent across sources:
- IBM splits long-term memory into episodic, semantic, and procedural types, stored in databases, knowledge graphs, or vector embeddings (IBM Think).
- Mem0 describes a hybrid store that blends vector similarity with graph-traversal confidence into one relevance score (Mem0).
- Redis recommends a read, reason, act, write loop built on embeddings, HNSW indexes, and hybrid retrieval, which it calls the strongest default (Redis, April 2026).
- Machine Learning Mastery maps episodic memory to vector databases, semantic memory to knowledge graphs, and procedural memory to fine-tuning or workflows (MLM, December 2025).
- AWS AgentCore extracts memories asynchronously with an LLM, then consolidates them with a second LLM step. AWS says the pipeline finishes in 20 to 40 seconds for a standard conversation (AWS Machine Learning Blog, October 2025).
- Zep’s Graphiti builds a temporal knowledge graph in which every fact carries a validity window (Graphiti on GitHub).
Episodic, semantic, and procedural memory
Most of these systems borrow the cognitive-science taxonomy:
| Memory type | What it holds | Alyx example | Typical store |
|---|---|---|---|
| Episodic | Specific past events | “Last Tuesday, prompt v3 scored 71% on faithfulness” | Vector database of past interactions |
| Semantic | Facts and generalizations | “The golden regression set is rag-query-dataset” | Knowledge graph or vector database |
| Procedural | How to do things | “Confirm before running an experiment“ | Prompt updates, fine-tuning, workflows |
Each type usually gets its own extraction job, its own store, and its own retrieval path.
Knowledge graphs and temporal facts
Knowledge-graph memory goes further by storing relationships. In Graphiti, every ingested message becomes an episode. An LLM extracts entities, such as people, projects, and datasets, and facts that connect them, like “Alice works at Acme Corp.” Each fact tracks two timelines: when it was true in the world (valid_at and invalid_at) and when the system learned or retired it (created_at and expired_at). A newer, contradicting fact invalidates the old one instead of deleting it, so the graph can answer “what did we know about Alice’s employer as of last Tuesday?” (Rasmussen et al., arXiv:2501.13956).
Retrieval combines three searches: cosine similarity over embeddings, BM25 keyword search, and breadth-first traversal outward from seed nodes. A reranker merges the results. The options are reciprocal rank fusion, maximal marginal relevance, graph distance from a center node, episode-mention counts, or a cross-encoder.
This is well-designed software, and it’s impressive architecturally. Temporal invalidation cleanly solves a problem flat memory stores ignore: facts go stale. When it’s tuned, it performs. Zep’s paper reports accuracy gains of up to 18.5% on LongMemEval while shrinking the context from about 115,000 tokens to about 1,600 (Rasmussen et al.).
Choosing a memory architecture
The retrieval-heavy approach above is one way to build memory, but it isn’t the only one.
In practice, most architectures fall into four broad patterns:
| Architecture | How it works | Where it fits |
|---|---|---|
| Profile or memory file | Keep one compact representation of what the agent should know and load it each time | Preferences, goals, conventions, bounded user or project state |
| Searchable collection | Store memories independently and retrieve a subset for each request | Large or open-ended memory |
| Knowledge graph | Represent entities, relationships, provenance, and sometimes time | Relational knowledge and historical questions |
| Hybrid | Always load a small core and retrieve the long tail | Systems that need predictable high-priority context plus larger history |
The main tradeoff is between capacity and retrieval.
A collection or graph can contain far more information than you would ever put into a prompt. But the right memory only helps if the retrieval system selects it.
A profile has the opposite constraint. If the information fits, the entire thing can be placed in context. There’s no retrieval cutoff, but the profile has to stay small enough to load every time.
That makes one question critical: Could the information that really matters fit comfortably in a few thousand tokens?
If the answer is yes, you should seriously consider a profile before building a retrieval system.

AI memory tools and platforms
Those architecture choices show up differently across current memory tools. There isn’t one standard stack, and I wouldn’t choose between them from a feature checklist. Start with the shape of the memory problem you need to solve.
| Tool | Approach | Where it fits |
|---|---|---|
| LangMem | Profiles and searchable collections, with foreground or background memory management | Useful when you want primitives for building your own memory behavior, particularly in the LangGraph ecosystem |
| Mem0 | Vector-based memory with optional graph relationships | Useful when semantic retrieval is the main access pattern but relationships can add context |
| Graphiti / Zep | Temporal knowledge graphs with semantic, keyword, and graph retrieval | Useful when relationships, provenance, and changing facts are first-class requirements |
| Amazon Bedrock AgentCore Memory | Managed short-term events plus asynchronous extraction, consolidation, and semantic retrieval | Useful when you want a managed memory layer inside an AWS agent stack |
| Redis Agent Memory | Session memory plus background extraction into searchable long-term memory | Useful when you need semantic, keyword, or hybrid retrieval with configurable retention |
LangMem makes one distinction I find particularly useful: profiles represent bounded current state in one document, while collections let memory grow across interactions and retrieve only what is relevant later. Mem0 can layer graph relationships alongside vector-based retrieval, while Graphiti makes time and provenance part of the graph itself. AgentCore and Redis both provide more managed versions of the extraction-and-retrieval loop described above.
The architecture matters more than the product. Decide what the agent needs to remember and how that information needs to be accessed before choosing the infrastructure.
What should an agent actually remember?
More memory isn’t automatically better.
Every memory takes up context somewhere and can quickly become stale. Moreover, if it’s wrong, the agent may carry that mistake into a completely unrelated future task.
A good default is to store information that is:
- Durable. It will probably remain true long enough to matter.
- Difficult to reconstruct. The agent cannot simply fetch it again from an authoritative tool or API.
- Useful across sessions. It is likely to change a future response or decision.
For a scoped agent, that often means a fairly small set of information:
| Category | What belongs there | Example |
|---|---|---|
| User profile | Durable preferences, expertise, terminology | “Prefers raw span filters instead of summaries” |
| Goals | Objectives, progress, blockers, major decisions | “Reach 90% retrieval accuracy before launch” |
| Projects | Durable facts a fresh tool query will not reveal | “Ground truth lives in metadata.label” |
| Assets | Intent and history behind prompts, datasets, evals, or experiments | “v4 removed chain-of-thought after latency regressed” |
| Conventions | Repeated workflow preferences | “Do not rename prompts when saving a new version” |
| Corrections | Mistakes the agent should not repeat | “Confirm before running an experiment” |
One distinction that’s particularly useful is state vs. history.
If an agent can fetch the current value again, memorizing it may create a stale copy. A historical checkpoint is different. “Prompt v2 scored 0.648 against a 0.90 launch target” records progress and explains why later decisions were made.
The hard part is maintaining an agent memory system
Choosing a store is only part of the implementation.
A production system still has to decide when to write memory, how to update it, what to forget, and how to stop bad memory from propagating into later sessions.
Write memory in the foreground or the background
You can let the main agent manage memory directly while it works. This makes sense when remembering something is an explicit part of the task.
The alternative is to process the interaction after the foreground task finishes and update memory asynchronously.
In our work at Arize, we generally prefer the second model when memory is incidental to the user’s request. It keeps memory management from competing with the actual task for attention, latency, and tool calls.
It also means the memory pipeline can inspect more than the final answer.
Agents can also learn things through tools, and that’s important to keep in mind. A trace query might reveal where ground truth is stored. An experiment might show that one prompt improved quality but increased latency. If memory extraction only reads the final response, it can lose the exact discovery that should persist.
Make updates as small as possible
Repeatedly asking an LLM to rewrite an entire memory document creates drift.
A correct detail can disappear. Wording can change for no reason. One “cleanup” pass can quietly alter a fact that had been correct for months.
In our AI engineering agent Alyx, we moved toward surgical operations instead: add text at a known location, replace an exact string, or delete an exact string. Everything else stays untouched.
The broader principle applies regardless of storage layer: if you care about understanding how memory changed, explicit inserts, updates, invalidations, and deletes are easier to reason about than unconstrained regeneration.
Every memory system eventually needs to forget
This may initially sound counterintuitive, but forgetting is part of keeping memory useful.
Profiles need compaction, while collections need some combination of consolidation, deduplication, expiration, decay, or deletion. Otherwise, useful context gets buried under information that is stale, redundant, or no longer relevant. Old preferences, superseded decisions, and incorrect conclusions can also keep influencing future tasks long after they should.
Different memories have different lifespans. A current goal may matter for weeks, a debugging hypothesis for an hour, and a correction like “never modify production without confirmation” much longer. Treating them all as equally durable makes memory noisier and less reliable.
Forgetting should therefore be built into the system from the start. A compaction policy should preserve high-priority information such as user-provided facts, active goals, preferences, and corrections while removing completed work, redundant details, and information the agent can easily reconstruct. If compaction fails validation, keeping the previous good state is usually safer than overwriting it.
Facts change
Suppose the system remembers: Alice works at Acme.
Six months later, Alice works somewhere else.
If you only care about the current value, replacing the old fact may be enough. If the application needs to know where Alice worked last year, then you need some representation of time and provenance.
This is one of the cases where temporal graph memory earns its complexity.
Treat persistent memory as untrusted input
Memory eventually gets put back in front of the model.
If malicious text gets written once and is then reinserted into context for every future session, a temporary prompt-injection problem can become a persistent one.
Memory systems therefore need the same kinds of boundaries as any other persistent application state: isolation between users and tenants, controls around what can be written, and validation before untrusted text becomes durable context.
You also need to think about concurrency. Two sessions can read the same memory state, update it independently, and overwrite one another. Serialization, version checks, transactions, or locks are boring compared with vector search, but they matter once memory is real production state.
How do you evaluate AI agent memory systems?
Agent evaluation in memory systems has to test more than whether the system retrieved the right fact.
A fact can be written correctly and never influence the agent. It can also be retrieved perfectly and make the answer worse because it is stale, irrelevant, or applied too broadly.
From experience, we recommend testing agent memory at three levels.
1. Did the system write the right thing?
Inspect memory after interactions that should create or update it.
This catches missing facts, over-extraction, bad updates, and information that disappears during compaction.
Save snapshots as the scenario progresses. If a fact is missing at the end, the useful question is not just whether it disappeared. You want to know which turn removed it.
2. Does the agent use memory when it should?
Build multi-session scenarios where the correct response depends on information learned earlier.
For example, a user specifies a preference in session one and the conversation then ends.
In session two, the user makes a related request without repeating the preference.
The agent should apply it.
Run the test through the real write and read paths. Manually inserting the expected memory into the prompt proves the model can use the fact, but not whether your memory system actually works.
3. Does the agent ignore memory when it should?
This is the other half of the problem.
If the user once asked for a 70-row dataset and later explicitly asks for a small dataset, the new instruction should win.
If project A had a hallucination problem, the agent should not automatically diagnose project B the same way just because that history is nearby.
We call these memory-corruptible cases, or tasks where remembered information could plausibly influence the answer but should not.
A useful memory system has to pass both sides. It should make the agent better when history matters and stay out of the way when it doesn’t.

What memory is already built into agent frameworks?
If you’re already building with an agent framework, you may not need separate memory infrastructure on day one. Most major frameworks now provide some way to persist state or put saved information back into an agent’s context.
The implementations differ.
- LangGraph’s ecosystem supports state that persists within a thread and long-term memory that persists across threads. Deep Agents, which builds on LangGraph, can use its Memory Store to save and retrieve information from previous conversations.
- AutoGen exposes a Memory interface around adding, querying, and injecting stored information into model context. Its built-in examples range from a simple chronological list to persistent ChromaDB and Redis stores, and it also includes a Mem0 integration.
- CrewAI now has a unified memory system that can extract facts from completed tasks, organize them into scopes, and retrieve relevant memories before later tasks. Recall can combine semantic similarity, recency, and importance rather than relying on similarity alone.
The distinction between framework memory and a standalone memory layer is therefore less about whether the framework can remember anything and more about whether its memory model matches the lifecycle you need.
If one framework owns the agent and its persistence model fits the application, built-in memory may be enough. A separate layer becomes more useful when memory needs to be shared across different agents or frameworks, governed independently, or stored and retrieved in ways the framework does not provide.
When should you add a standalone memory layer?
First and foremost, you may not need a standalone memory layer.
If your framework already gives you the persistence and retrieval behavior the agent needs, adding another memory system creates another component to operate and evaluate.
A standalone layer starts to make more sense when memory needs to outlive a particular agent implementation, be shared across agents or frameworks, or support capabilities such as large-scale semantic retrieval, temporal relationships, independent retention policies, or custom access controls.
The same rule applies here as everywhere else in the stack: start with the smallest system that reliably gives the agent the context it needs, then add infrastructure when a concrete requirement justifies it.
Does a bigger context window replace agent memory?
A larger context window helps, but it does not remove the need for memory.
More context gives the model room for a longer conversation history. It does not decide which information should persist across sessions, how long it should remain relevant, or what happens when a preference, goal, or fact changes.
You could keep sending months of interaction history with every request, but eventually the same problems return: cost grows, stale context accumulates, and useful information becomes harder to identify.
Context windows give the model capacity. Memory decides what is worth carrying forward.
When should you use a knowledge graph for AI memory?
Knowledge graphs are most useful when the relationships between memories matter as much as the memories themselves.
That usually means applications where facts change over time, historical queries are important, or the agent needs to reason across people, projects, assets, dependencies, or organizations.
If the main requirement is remembering a user’s preferences, goals, or working conventions, a graph may add more complexity than value. But if you need to answer questions like “who owned this account when the incident happened?” or “how did this relationship change over time?”, graph structure and temporal validity become much more useful.
Can memory make an agent worse?
Memory can absolutely make an agent worse if the underlying system isn’t well designed.
A memory system can preserve incorrect information, keep stale assumptions around too long, or carry context from one task into another where it no longer applies. In the worst case, one bad conclusion can become something the agent keeps reusing across future sessions.
That is why memory should be evaluated as part of the agent’s behavior, not just as a storage or retrieval layer.
Test whether the agent uses memory when it should. Just as importantly, test whether it ignores memory when the current task should take precedence.
A practical default for building agent memory
Start with the information your agent actually needs to carry across sessions. In most cases, that set is smaller than you think.
If the important context fits comfortably in a compact profile, I would start there before adding retrieval. Prioritize information that is durable, difficult to reconstruct from the system itself, and likely to change what the agent does later.
Memory formation also does not need to sit in the foreground path unless the task requires it. Moving extraction and updates into the background keeps the agent focused on the user’s request and makes the memory pipeline easier to inspect independently. From there, add an explicit forgetting strategy, treat persistent memory as untrusted input, and test the whole system with multi-session evals.
Add retrieval when the useful memory no longer fits reliably in context. Add graph structure when relationships, provenance, or temporal history become requirements.
The architecture should get more complicated when the problem demands it, not before.
Related reading
- How to evaluate AI agents in production
- Agent observability
- Session-level evaluations
- How to find and debug agent failures your evals are missing
- LLM evaluation
- Agent evals from production traces
- Memory is still a missing primitive: What memory systems are actually shipping
- What is an agent observability platform?
Build and evaluate agents with Arize
Memory makes an agent stateful across sessions. That means failures can come from more than the model response itself: what the system remembered, what it retrieved, what it forgot, and which old context affected the next decision all become part of the debugging surface.
For the implementation behind these ideas, read How we built long-term memory for Alyx with one 8,000-character file.