Over the last few months, I built long-term memory for Alyx, the AI engineering agent inside Arize AX. I started with a fairly sophisticated architecture (and coincidentally where most of the internet told me to start): vector retrieval, separate episodic and semantic stores, and a temporal knowledge graph.
What we ended up shipping is much smaller. Each user gets one structured text file for each Arize space they work in, capped at 8,000 characters, and Alyx reads the whole thing on every request.
I spent weeks building the more complicated version before arriving there, though. The interesting part was figuring out what Alyx actually needed to remember, how to keep that memory useful as it changed over time, how to keep the pipeline cost-efficient, and how to test whether memory was helping rather than quietly making the agent worse.
That simpler design also puts Alyx close to the approach OpenAI and Anthropic have taken with ChatGPT and Claude, and after building both versions, I think it’s the right default for most scoped agents.
This post explains why that design fit Alyx, how we update and compact the memory file, and how we test whether memory improves the agent’s work. It also covers the failures we found, including the preferences that disappeared and the mistakes the agent wrote back into its memory.
Key takeaways
- Long-term memory carries useful context across sessions, including user preferences, project conventions, and the reasoning behind earlier decisions.
- Retrieval and knowledge-graph memory are powerful but costly to run. Every write fans out into many LLM calls, and retrieval quality becomes its own tuning project.
- Alyx now uses one structured memory file per user and space as its core long-term memory system. This removes memory retrieval from the read path while leaving the challenge of deciding what to preserve and when to apply it.
- Background jobs summarize completed turns and propose targeted edits. Compaction keeps the file within its size limit, but a rejected update can leave recent information unsaved. ChatGPT and Claude both put a compact memory profile straight into the model’s context and keep search for the long tail.
- For a scoped agent, one always-loaded, well-structured file can be enough. In our final evals, Alyx answered 16 of 18 memory-dependent questions correctly, both across three separate scenarios and in a mixed run that interleaves them.
- We also tested memory from two directions: questions that need memory, and questions that memory could make worse. Alyx passed every question in the second group.
Need a primer on long-term memory in agents?
Long-term memory is context an agent keeps across sessions and uses to inform what it does later, including durable user preferences, goals, conventions, and decisions. This post focuses specifically on how we implemented long-term memory in Alyx. For the broader architecture, terminology, and approaches to agent memory, read our guide to long-term AI memory.
Build better agents with Arize
Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today.
Prefer open source? Try Arize Phoenix for self-hosted, open source agent observability.
Why memory is a different kind of context
Agents already get several kinds of context. The system prompt says who they are. Tool definitions say what they can do. Retrieved documents supply facts, and the session history records what has happened so far. Long-term memory is its own category: what the agent knows about this user that none of the others can supply.
The problem we wanted to solve with memory in our agent Alyx
Without long-term memory, the session in Alyx has to carry everything, and sessions don’t hold up well:
- They end. A new chat starts from zero, so the user explains their project, their preferences, and last week’s decision all over again.
- They degrade as they grow. Chroma tested 18 models and found that “performance grows increasingly unreliable as input length grows,” even on simple tasks (Chroma, “Context Rot,” July 2025). Earlier research showed models are worst at using information buried in the middle of a long input (Liu et al., “Lost in the Middle,” 2023).
- They get compacted. Long sessions eventually get summarized to fit. A summary written for the current task drops details that matter for the next one.
Alyx has this problem worse than most chat apps. Users move between trace views, eval pages, the prompt playground, and dashboards, and each surface can start its own thread.
If Alyx learns while investigating traces that ground truth lives in metadata.label, that fact should still be available when the same user opens the eval builder in a new session. The same is true for preferences like “datasets should have exactly 70 rows,” project conventions, previous decisions, and corrections to mistakes Alyx made.
Keeping all of that in conversation history doesn’t solve the problem. Sessions end, long histories become harder for models to use reliably, and conversations eventually get compacted. We needed a separate layer for the small amount of context that should still matter when the next session starts.
Our first implementation used retrieval and a knowledge graph
My first Alyx prototype treated memory as a retrieval problem. It modeled users, projects, assets, goals, and product areas such as graph entities, then retrieved relevant memories for each request.
This paradigm is consistent and has shown success for many reputable sources:
- IBM splits long-term memory into episodic, semantic, and procedural types, stored in databases, knowledge graphs, or vector embeddings (IBM Think).
- Mem0 describes a hybrid store that blends vector similarity with graph-traversal confidence into one relevance score (Mem0).
- Redis recommends a read, reason, act, write loop built on embeddings, HNSW indexes, and hybrid retrieval, which it calls the strongest default (Redis, April 2026).
- Machine Learning Mastery maps episodic memory to vector databases, semantic memory to knowledge graphs, and procedural memory to fine-tuning or workflows (MLM, December 2025).
- AWS AgentCore extracts memories asynchronously with an LLM, then consolidates them with a second LLM step. AWS says the pipeline finishes in 20 to 40 seconds for a standard conversation (AWS Machine Learning Blog, October 2025).
- Zep’s Graphiti builds a temporal knowledge graph in which every fact carries a validity window (Graphiti on GitHub).
I spent weeks inside Graphiti building that version. The system was capable, particularly around relationships and facts that change over time, but the implementation exposed three problems for Alyx.
Where retrieval and graph memory got hard for us
1. Broad preferences were easy to miss. A retrieval query built from the user’s latest message may not surface the conventions that should shape the response. “Make a dataset of agent-prod traces” has little semantic overlap with “this user wants datasets with exactly 70 rows” or “the hallucination eval only covers the summarize_book span.” For Alyx, loading these conventions on every request avoided having to retrieve them separately.
2. Knowledge graphs are expensive to maintain. Every write is LLM work. Reading Graphiti’s add_episode path while I prototyped, I counted one LLM call to extract entities, one to resolve ambiguous entities against the existing graph, and one to extract facts. Then it fans out.
There’s one call per fact to check for duplicates and contradictions, whenever similar facts already exist. There are per-fact calls for timestamps when dates are missing, per-fact and per-entity calls for attributes when you define custom types, and batched calls for entity summaries.
For a turn with 5 entities and 6 facts, that’s roughly 10 LLM calls without custom types and more than 20 with them.

That cost repeats on every turn for every user. It’s also latency you have to hide and rate limits you have to manage. Graphiti’s README explains that its default concurrency limit exists to help prevent HTTP 429 rate-limit errors (Graphiti on GitHub).
3. It becomes a difficult data science problem. Once memory is retrieval, the debugging questions change. Instead of “did Alyx remember the right thing?” you’re asking about recall at k, embedding models, and whether a similarity threshold merges two different datasets that happen to have similar names.
Answering those questions takes a benchmark, and memory benchmarks are shaky. The standard ones, LoCoMo and LongMemEval, test whether a system can recall facts from long chat histories (Maharana et al., arXiv:2402.17753; Wu et al., arXiv:2410.10813). Even on those, vendors can’t agree on the numbers. Zep’s corrected LoCoMo score for itself is 75.14% (Zep). Mem0’s paper reported 65.99% for Zep (Chhikara et al., arXiv:2504.19413), and a Mem0 re-run posted on GitHub put it at 58.44% (getzep/zep-papers #5). Meanwhile, Letta showed that a GPT-4o mini agent that stores conversation history in plain files and searches it with grep-style tools scored 74.0% (Letta, August 2025). And in Mem0’s own paper, putting the full conversation in context beat every memory system on accuracy.
The tradeoff was cost: Mem0 reports 91% lower p95 latency and more than 90% lower token cost than full context (Chhikara et al.).

Even a clean score wouldn’t tell you whether an agent does better work in your product. For Alyx, a fair retrieval benchmark would mean labeling which memories should come back for every question. That’s a labeling project before you’ve written any memory code.
Then there are the knobs. Rerankers alone give you five options with different tradeoffs. RRF is cheap. MMR trades relevance for diversity. Node distance needs a center node. Episode mentions favors facts that come up often. A cross-encoder is the most accurate, but it’s another model call. On top of that you have top-k, traversal depth, entity types, and edge types. Every knob changes what the agent sees, and you can’t measure the effect without an eval, which brings you right back to the benchmark problem.
The long-term agent memory architecture we shipped
The alternative is a profile document: one text file holding the user’s goals, preferences, conventions, and current work, loaded into context on every request. Nothing is searched at read time.
That’s also roughly where the most popular AI tools have landed:
- ChatGPT. Reverse-engineering in 2025 found that ChatGPT’s memory layers aren’t retrieved at all. Saved memories, summaries of recent conversations, and usage metadata are sent with every message. Shlok Khemani summed it up: “OpenAI just includes everything with every message” (Khemani, September 2025; see also Simon Willison, May 2025). In 2026, OpenAI reportedly added a periodic background pass that revisits recent conversations and updates the user’s profile (Khemani, June 2026).
- Claude. At launch, Claude’s memory was an editable summary with a separate memory for each project (Claude blog), and the synthesis refreshed every 24 hours (Claude Help Center). In 2026 it became categorized entries that Claude reads and updates during conversations (Claude release notes).
- Claude Code.
CLAUDE.mdfiles load in full at the start of every session. Anthropic recommends keeping each one under 200 lines, because longer files use more context and reduce adherence (Claude Code docs). - Letta and LangMem. Letta, the company behind MemGPT, keeps core memory blocks in the context window at all times. Its docs say they’re “always visible – no retrieval needed” (Letta docs). LangMem calls the same pattern a profile: a single document that’s updated in place (LangMem).
For Alyx, this also changes the central memory problem. Instead of asking whether the right fact can be retrieved for a particular query, we can focus on whether we saved the right information in the first place, whether it survived updates and compaction, and whether Alyx actually uses it when it matters.
Why we chose an always-loaded file for Alyx
- Nothing to miss. The whole file is in context, so no retrieval step can go wrong. The model decides what’s relevant, and that’s something LLMs are good at.
- It’s cheap. 8,000 characters is roughly 2,000 tokens. Reading it means fetching one database row. Writing it takes two background LLM calls per turn plus a small-model safety check. In an earlier version of our combined eval run, that background step took a median of about 13 seconds on turns that didn’t need compaction and about 67 seconds on turns that did, all of it after the user already had their answer.
- It’s easy to debug. When memory is wrong, you open the file and read it. There are no retrieval scores to interpret.
- It’s often enough. Whether a file is enough depends on scope. A general-purpose assistant might need to remember anything. A scoped agent doesn’t, and that changes the math.
Making one file enough for Alyx: what we actually store
The memory file isn’t a blank scratchpad. Two design choices make 8,000 characters go a long way.
A fixed structure built around what Alyx does. Every memory file starts from the same skeleton:
## User Profile
## Goals
## Projects
## Assets
## Conventions
## Interaction corrections
| Section | What goes in it | Example |
|---|---|---|
| User Profile | Role, technical depth, answer-style preferences, the app’s domain, personal aliases | “Wants raw span filters, not summaries”; “‘the judge’ = the hallucination eval” |
| Goals | The active objective, its status and blocker, and decisions with rationale, edited in place | “Get retrieval accuracy to 90% before launch; at 65%, blocked on the vector store” |
| Projects | Facts about a traced app that a fresh query won’t reveal | “Ground truth lives in metadata.label“; “latency spikes trace to the reranker” |
| Assets | The intent behind evals, datasets, prompts, and experiments | “v3 added chain-of-thought; v4 removed it after it hurt latency” |
| Conventions | How this user works across all their projects | “Don’t rename prompts when saving a new version” |
| Interaction corrections | Past mistakes and the fix | “Confirm before running experiments” |
The updater follows one selection rule. An entry has to be durable (still true next month), non-re-fetchable (a tool call can’t reconstruct it), and useful across sessions (it shapes many future requests, not just one). Configs, IDs, URLs, filter strings, and live metrics stay out because Alyx can look them up. The exception is results that measure progress toward a goal, like “v2 scored 0.648 against a 0.90 target.” Those are history, not live readings.
When the user states a fact directly, the line gets a [user] prefix. Those facts are the hardest to rediscover, so compaction protects them.
Here’s a lightly trimmed excerpt from a memory file Alyx wrote during one of our eval scenarios, a legal RAG app. It comes from an earlier version of the updater prompt, which gave each goal and project its own ## header. The current prompt nests them as ### subsections under the fixed categories.
## Goal: Get the Legal RAG application production-ready
- Active objective: Raise retrieval accuracy/score to the 90% production target
as the main path to improving RAG answer quality and readiness.
- Current status/blocker: Not ready for production yet. V2 is the best
query-construction prompt so far, but retrieval quality remains below target;
the main blocker is retrieval/knowledge coverage rather than generation,
latency, or cost.
...
- Decisions/rationale:
...
- Use v2 over v3 because v2 wins consistently by category, not just overall.
- Prompt tuning alone is unlikely to close the remaining gap to the 90% target.
...
## Project: Legal RAG application
...
- Key recurring root cause: Poor answer quality usually stems from retrieval
misses/weak semantic matching, not LLM behavior.
...
- Architecture quirk: Trace schema appears to include reranker-related fields,
but reranking seems unused/null.
We intentionally maintained a narrow scope. Alyx works inside Arize AX on the agents and apps you trace there. That means it doesn’t need your travel plans or your dog’s name. Memory is stored per user and per space, so project and dataset details never cross a space boundary.
Compaction
Ingestion never compresses, so a new turn is allowed to push the file past its 8,000 target. When the file goes over 7,000 characters, a separate compaction call rewrites it toward 6,000. Compaction follows a strict priority order. It keeps the User Profile and active Goals intact first, then [user] lines, then the historical checkpoints that tell a goal’s story. It merges redundant bullets first, tightens wording second, and cuts only as a last resort.
The server rejects anything over 8,000 characters. If compaction can’t get under the cap, the pipeline doesn’t write, and the last good file stays in place.
This isn’t hypothetical, and compaction was the weakest part of the pipeline in our evals. In an earlier version of our combined eval run (described below), on a build with an 8,192-character cap, compaction ran on 22 of 48 turns and never reached its target. On six of those turns the file stayed over the cap, so those turns’ new edits were dropped. In one case, three attempts in a row shrank the file from 9,009 characters to 9,005.
Each time, the existing file survived. Setting the trigger 1,000 characters below the cap means a compaction that falls short can still produce a file the server accepts, as long as the turn didn’t push the file far past the trigger.
How Alyx updates memory after every turn
Our first version gave Alyx a tool, edit_alyx_memory, and told it to save anything durable before finishing a turn. That meant the agent had to decide mid-task what was worth remembering, and memory writes competed with the user’s actual request.
Because of that, we moved memory out of the foreground entirely. It now runs as two background LLM jobs, which we call sidecars, after each turn:
- Turn summarizer. When a turn completes (not when it pauses for user approval), a background task reloads the saved turn and asks an LLM to summarize it. The turn includes the user’s messages, Alyx’s replies, and every tool call and result. We include tool calls on purpose. Many of the best memories are discoveries made along the way, like a column that turned out to be empty or the variable mappings Alyx chose for an eval. The final reply rarely mentions those.
- Memory updater. A second LLM call reads the current memory file and the summary. It returns a short list of edit operations, not a rewritten file. An
addinserts text after an anchor. Areplaceswaps one exact string for another, and replacing with an empty string deletes. Each anchor must match exactly once, or the operation is rejected and the model retries with the error as feedback, up to three attempts. An empty list means the turn produced nothing worth keeping, which is a normal outcome. Because untouched text stays byte-for-byte identical, the file doesn’t drift between compactions.

A few guardrails sit around those two calls:
- Safety gate. Memory is injected as a system message in every future session, so a poisoned entry would be a persistent prompt injection. A small model classifies every new piece of text as allow or block before it’s saved. Compaction output gets checked in full. If the check can’t run, the write is skipped.
- Serialization. Updates for the same user and space run one at a time, enforced by a local lock plus a Postgres advisory lock, so two replicas can’t interleave read-modify-write cycles. The advisory locks get their own small connection pool so memory jobs can’t starve foreground requests.
- Best effort. Any failure in the pipeline is logged and swallowed. Memory writes never delay or break the user’s response.
On the read side, Alyx loads the latest file once per request and injects it as a system message that’s never saved to chat history. The injected block comes with three instructions: answer from memory first, check memory for an existing asset before creating a duplicate, and treat [user] lines as high-confidence facts.
The Graphiti prototype vs. the memory system we shipped
Since we originally built Alyx’s long-term memory using Graphiti, we were able to compare it with the solution we ended up adopting.
Two differences stood out:
- The graph-based system could contain the right memory without actually putting it in front of the agent. This was because every request depended on retrieval and reranking. With a file-based approach, the entire memory is always in context.
- The simpler architecture removed a surprising amount of machinery from both the read and write paths. This made the system much easier to debug and evaluate.
Here’s how the two approaches compare:
| Graph memory (Graphiti-style) | Alyx memory file | |
|---|---|---|
| LLM calls per write | 2-3 fixed calls, plus calls per fact and per entity | 2 (retried on bad anchors), a small-model safety check, and up to 3 compaction calls above 7,000 characters |
| Read path | Hybrid search and reranking on every query | Read one row |
| What the agent sees | Top-k retrieved facts | The whole file |
| Debugging | Inspect the graph and retrieval scores | Read the file |
| Typical failure | The right fact exists but isn’t retrieved | The right fact is in context but isn’t applied |
How we evaluated Alyx’s memory
A smaller system is easier to get right, but it still needs agent evals.
We built a harness that runs scripted, multi-session scenarios through the real Alyx agent with memory turned on. Each scenario starts with seeding sessions, where the simulated user works normally. Later sessions ask the eval questions. Conversation history resets between sessions, and only the memory file carries over. Every turn runs the full summarize-and-update pipeline, and the harness saves a copy of the memory file after each turn.
We ran three scenarios on their own:
- a book-summary agent with separate local and production projects
- a legal RAG app
- Alyx’s own production traces
The fourth combined run is a separate run that mixes all three cases. Instead of running one scenario after another, it interleaves their sessions, so the simulated user jumps between projects the way someone working on several things at once would. That makes it the hardest test in the set. With three projects’ worth of context competing for the same 8,000 characters, it puts the most pressure on compaction and on the updater’s selectivity, meaning its judgment about what’s worth keeping. It runs 48 turns across 15 sessions.
Every turn has a ground-truth expectation for what memory should record from it, and every eval question has a written pass/fail criterion. Answers were graded by reading the response and its trace, not by an automated judge, from a single run per scenario. Treat the numbers as a regression check, not a benchmark. If you want to automate this kind of grading, an LLM-as-a-judge evaluator with written criteria is the natural next step.
The results below come from our final runs, after the prompt and pipeline changes described in this post. Many of those changes came out of earlier runs, and I’ll point out a few of their failures along the way.

Evaluating memory writes
The first thing to check is what gets written. A write counts as successful when the memory file after a turn matches that turn’s ground truth. In the final runs, writes matched on 22 of 23 turns for the book-summary agent, 7 of 10 for the legal RAG app, 13 of 13 for Alyx’s production traces, and 44 of 48 in the combined run. Most of our prompt changes came from reading these files turn by turn.
An early replay of 70 real Alyx requests across 39 sessions used the original tool-based design. It produced a file full of trace IDs, encoded asset IDs, filter strings, and counts that were stale by the next session. That’s where the “don’t store what a tool can fetch” rule came from.
Reading the file turn by turn also catches silent losses. In an earlier run of the book-summary scenario, the user said in the first session that datasets should have exactly 70 rows. Memory recorded it, but only as a note on one dataset. Then the update after the first turn of session two deleted that line with a targeted edit, and it never came back. An end-of-run check would show the fact missing, but not when it disappeared or which edit removed it.
The [user] prefix is our response: it marks facts the user stated directly as worth keeping, compaction ranks them just below the profile and active goals, and Alyx treats them as high-confidence. In a later combined run, where the updater used the prefix, the 70-row preference was still in memory at the end, tagged [user] on that dataset.
Memory-dependent questions
These questions can only be answered well if memory worked. Some examples, quoted as the simulated user typed them:
- “In my agent-prod traces how does the hallucination compare to agent-local?” A good answer limits the analysis to the
summarize_bookspan. The user stated that preference in a different session, about a different project. - “Do we have any evals for latency?” A good answer identifies
priyan-evaluator-testas the latency evaluator, even though the name doesn’t say so. - “Are we ready to push to prod?” A good answer says no, because the best prompt scored 0.648 against a 0.90 target. It shouldn’t run a new experiment to find out.
- “Whats the PR number for merging the alyx agent and playground agent? I forgot” A good answer recalls the number the user mentioned several sessions earlier.
In the final runs, Alyx answered all 7 memory-dependent questions correctly for the book-summary agent and all 6 for its own production traces. The legal RAG app was the weak spot at 3 of 5, and it also had the lowest write score, 7 of 10. In the combined run, with three projects competing for the same file, Alyx still got 16 of 18.
Memory-corruptible questions
The other direction matters just as much. These are questions where a sloppy memory makes the answer worse:
- “Hallucination rate of copilot-prod?” Memory is full of hallucination context from the book-summary projects. None of it should leak into a different project.
- “Make a small dataset of agent-local traces.” The scenario stored a 70-row dataset preference earlier. The user said “small,” so Alyx should override the stored default.
- “How should we improve hallucination rate of copilot-prod?” Upgrading the model helped the book-summary agent, but it’s the wrong advice here, because copilot-prod traces use several different LLMs.
In the final runs, Alyx passed every memory-corruptible question: 3 of 3 in each of the single scenarios and 9 of 9 in the combined run, where the other projects’ context sat in the same file the whole time.
Earlier runs also failed in instructive ways:
- The fact was in memory, but Alyx didn’t use it. Memory recorded that hallucination should be measured only on
summarize_bookspans. When asked to compare projects, Alyx aggregated every span anyway. An always-loaded file removes retrieval misses, but the model still has to apply what it reads. - Then memory learned the wrong lesson. A few turns later, the updater wrote the flawed approach back into memory as a convention: compare the two projects on the shared
attributes.hallucinatedfield. A memory that records the agent’s own mistakes as rules could make later answers worse, even though no later question in that run happened to test it. It’s also why we read the file itself, not only the answers. - The fact was never stored. Asked how top users changed over time, Alyx re-queried both time windows instead of recalling earlier results. The selection rule deliberately skips most live metrics, and this sits right on that line: a result the user will want to compare against later.
Where this design stops working (or what if a single file isn’t enough?)
For Alyx, one file per user and space has held up because the agent has a narrow scope and tools that can fetch most live state from the platform. Memory only has to hold the durable context those tools can’t reconstruct.
That assumption won’t hold for every agent. A coding agent working across a large repository, a team-wide knowledge system, or a general-purpose assistant may accumulate more useful context than you can reasonably keep loaded on every request. At that point, I would keep an always-loaded core for high-priority context and add retrieval for the long tail.
Retrieval solves a demonstrated capacity problem instead of becoming the starting architecture by default.
lore is a good example of the next step. It’s a persistent memory layer for coding-agent harnesses like Claude Code, OpenCode, and Codex CLI. Knowledge lives in plain markdown entries with provenance, tagged by scale (from implementation details up to architecture) and by category (gotchas, conventions, preferences). At the start of a session, lore loads an index plus high-priority files within a budget. It loads domain files on demand and searches everything else with SQLite full-text search and BM25 ranking. (Disclosure: lore is built by my Arize colleague Dustin Ngo.)
Lore is still files you can read, with lexical search layered on top. That’s the same shape as everything in this post: an always-loaded core, with retrieval saved for the long tail.
What we’d keep if we built it again
The system we shipped for Alyx ended up much smaller than the one I started building, but getting there required being strict about what memory is actually for.
Retrieval and knowledge-graph memory are powerful, and for some problems they’re the right tool. For a scoped agent, they’re usually more machinery than the problem needs. Here’s what worked for Alyx:
- One structured file per user and space, always loaded, with no search at read time.
- A selection rule that keeps only what’s durable, can’t be fetched with a tool, and matters across sessions.
- Background sidecars that summarize each turn and apply small, exact edits, with compaction and a safety gate around them.
- Evals that test both whether memory helps and whether it hurts.
The failures were as useful as the successes. We found preferences that disappeared during compaction, facts that survived but weren’t applied, and cases where Alyx wrote its own bad reasoning back into memory as a convention. Those are the kinds of bugs you only find when you inspect the memory state alongside the agent’s answers.
Start with the file. Add retrieval when an eval shows you the file isn’t enough. If you want to see how Alyx uses memory on your own traces, start with the Alyx docs.
Frequently asked questions
Is a memory file just a longer system prompt?
When should I use a knowledge graph for agent memory?
How do you keep memory from becoming a prompt-injection vector?
Priyan Jindal works on Alyx at Arize AI and built its long-term memory system.