# 10. Memory Built for Agents

*The question: What changes when the main reader of our customer data is an AI agent, not a person at a dashboard?*

> **Questions this chapter answers**
> - Our CRM works fine for our people. Why does it fail our agents?
> - We already have RAG. Is that memory?
> - What should customer memory actually look like, and what goes where?
> - Context windows keep growing. Why not give the agent everything?
> - How do several agents share what they learn about the same customer?

## The short answer

For forty years, customer data was designed for one kind of reader: a person looking at a screen, who already knew most of the story and needed a reminder. That reader could interpret a status field because they were in the meeting where it was set. They could skim 200 activity entries because they knew which three mattered.

The main reader is now an agent. It was not in the meeting, it reads thousands of records a day, it pays for every token, and it cannot tell which field is authoritative or which note is stale. Hand it a dashboard-era record and it reads too much, reads the wrong slice, or acts on a field whose meaning it guessed.

The answer is not a bigger context window or a better search index. It is a different model of the customer. I call the pattern **Unified Record Memory**; in this handbook, where the record is almost always a customer, it appears as **Unified Customer Memory**. It is one record per entity that unifies everything an agent can know about it, kept in four forms that answer four different questions:

- **Typed properties** answer *which facts qualify, count, and govern?* (stage, renewal date, consent status).
- **Evidence-level memories** answer *what is relevant here?* (small, sourced observations: "the ops lead said route sync failed twice in June").
- **Long-form synthesis** answers *what is understood?* (an account brief an agent can read in one pass).
- **The graph** answers *what is connected?* (who reports to whom, which site belongs to which account).

All four are written in one save and read back in one call under a token budget. Every agent reads the same record before it acts and writes back what it learned.

For a leader, three things follow. This is where agent quality is decided: the same model writes a generic email or a specific, true one depending on what memory hands it. It is where cost is decided, because agents are billed by what they read. And it does not fix bad identity or bad objectives; it makes those mistakes faster.

## The reader changed

Open a Larkspur Systems account record the way its account managers have for ten years. (Larkspur is the fictional company that runs through this handbook: field-service and fleet software, about 40,000 customer accounts and 250,000 contacts.) Harlow Fleet Services, 400 technician seats. Tier: Gold. Health: Yellow. Renewal: March 31. Below that, 212 activity entries, most of them logged emails.

Priya, the account manager, reads it in thirty seconds. She knows Yellow means the route-sync failures in June, because she set it after the call. She knows the renewal date moved once, and why. She skips the activity log, because she was there for most of it.

Now watch the other reader. This quarter Larkspur runs a renewal agent, a support copilot, and an outbound pilot. Between them they read Harlow's record dozens of times a week, and across the book of business, several hundred account records a minute at peak. None of them was on the call. The renewal agent sees Yellow and does not know who set it, when, why, or whether the problem was fixed. It sees 212 activity entries and either reads all of them, paying for that on every call, or reads the five most similar to its question and misses the one that matters. It sees a renewal date and cannot tell whether it came from the contract, the billing system, or a rep's guess.

The CRM did not get worse. The reader changed.

## Dashboard-era data fails agents in four specific ways

This is not a data quality problem; Larkspur's data is as clean as any mid-market company's. The data was designed for a reader who carried the missing context in their head. Take the head away and four gaps appear.

**Fields without meaning.** "Health: Yellow" is defined by a team's habits. Two agents asked to update it will write two different concepts into the same field. A person resolves the ambiguity by asking a colleague. An agent resolves it by guessing, fluently.

**No provenance.** Without source, observation time, and confidence written down, an agent cannot tell a contract date from a rep's estimate. (Chapters 7 and 11 build the provenance envelope.)

**No narrative.** The activity log records that 212 things happened, not what they mean. The two-minute story Priya would tell a colleague exists nowhere the agent can read it.

**No budget.** A dashboard costs nothing to glance at. An agent pays for every token, and its accuracy degrades as the input grows. Chroma's 2025 study of 18 models found that "performance grows increasingly unreliable as input length grows," even on simple tasks [1]. Earlier, Liu and colleagues showed that models use information in the middle of long inputs worst of all [2]. Anthropic's engineering guidance puts it plainly: context "must be treated as a finite resource with diminishing marginal returns" [3].

That last gap changes the objective of data modeling. Relational design optimized for storage integrity: store each fact once, return exactly the rows that match. Memory for an agent optimizes for **the most task-relevant signal per token**. The discipline of modeling entities, relationships, and constraints transfers almost entirely. The objective does not.

## A prompt is not a memory

The first response most teams try is a bigger prompt: paste the account record, the last twenty emails, and the open tickets, and let the model sort it out. It works in a demo, because a demo asks one question about one customer once.

A prompt is temporary material assembled for one call. Durable memory is organizational state: it persists, can be retrieved selectively, carries source and time, and can be corrected, expired, or deleted. When the support copilot learns on Tuesday that Harlow's route-sync problem is fixed, the prompt the renewal agent assembled on Monday does not change. Nothing in a prompt compounds.

A prompt helps one call. A memory changes the next one.

## Model entities, not dossiers

The first design choice is what the system remembers. The tempting answer is a profile per customer: one big document with everything in it. That becomes brittle quickly, because the things we call "customer" are several different things.

Use a small set of durable entities instead, and let roles and relationships connect them:

| Entity | Larkspur example | Why it is separate |
|---|---|---|
| Person | Dana, Harlow's fleet operations lead; Priya, Larkspur's account manager | A person persists while roles change |
| Organization | Harlow Fleet Services; a regional reseller | Contracts often belong to organizations, not people |
| Role | "Operations lead at Harlow, 2023 to present" | Connects a person to an organization for a period and a purpose |
| Account | Harlow's subscription, 400 technician seats | Billing, usage, and entitlements attach here |
| Deal or case | The 2027 renewal; the June route-sync ticket | Work has its own status, participants, and evidence |
| Contract | The master agreement and its renewal clause | Commercial terms, and what is actually promised, live here |
| Event | A call, a ticket, a week of product usage | Time and sequence change what information means |

"Customer," "champion," "admin," and "former contact" become typed, dated relationships rather than containers. Dana **works at** Harlow; Harlow **owns** an account; Dana **is the admin of** it. When Dana changes jobs, the person persists, the role gets an end date, and the account needs a new admin. In a dossier model, Dana keeps receiving messages written for a job she no longer has.

Every one of these entities needs a resolved identity first (Chapter 9). Memory built on a wrong merge is a very efficient way to tell one customer another customer's renewal terms.

## One record, four ways of knowing

Here is the sentence this chapter turns on. **Memory for agents is not a place to store what you know about a customer; it is the working model of the customer that every agent reads before it acts and writes back to after.**

That model needs more than one shape, because knowledge about a customer comes in more than one shape. A single sales call about Harlow contains all four:

**Typed properties: what qualifies, counts, and governs.** Dana's role is "Fleet Operations Lead." Renewal stage is "Evaluating." Competitors under evaluation: 2. These values have one current state, and other systems need to filter, count, and trigger on them. `stage = 'evaluating' AND renewal_date < '2027-04-01'` must return the right accounts with no model in the loop.

**Evidence-level memories: what is relevant.** "On the September 12 call, Dana said the June route-sync failures cost her dispatchers two days of manual scheduling." That is a self-contained, dated, sourced observation. No schema anticipated it, and it is exactly what a renewal conversation should acknowledge.

**Long-form synthesis: what is understood.** The account brief: history, open risks, stakeholders, what we have promised, what is unresolved. An agent reads it in one pass instead of reconstructing the story from 212 fragments. It is a generated view, not self-authenticating truth: the underlying evidence and the system of record still win on conflict.

**The graph: what is connected.** Dana reports to Harlow's COO, who signed the original contract. Harlow shares a reseller with four other Larkspur accounts. The route-sync ticket is linked to a known defect affecting 60 other accounts. Edges should be typed, directional, dated, and sourced: `reports_to` is not `approves`, and `previously_admin` is not `admin`.

The division of labor is simple: **evidence preserves, properties operationalize, synthesis explains, the graph connects.**

![One piece of raw content, such as a call, email, or ticket, goes through a single save that embeds it for recall, extracts typed properties and atomic memories, and applies write-time gates (resolve pronouns, anchor time, deduplicate). The save writes onto one record per entity, kept in four forms: typed properties answer what qualifies, counts, and governs; evidence memories answer what is relevant here; long-form synthesis answers what is understood; the graph answers what is connected. Evidence preserves, properties operationalize, synthesis explains, and the graph connects, and every agent reads the record in one budgeted call and writes back what it learned.](/images/handbook/ch09-one-save-four-forms.svg)

*Figure 10.1. One save, four forms on one record, each answering a different question.*

This is related to the research taxonomies but not the same. CoALA splits an agent's long-term memory into episodic, semantic, and procedural stores [4]; MemGPT, which became Letta, pages information between context and an external store the way an operating system pages memory [5]. Both concern an agent's memory of its own experience. Unified Customer Memory is many agents' shared memory of a customer: a system of record read by models, not a model that keeps notes.

## If you would put it in a WHERE clause, it is a property

The hardest design question in memory is not which database to use. It is deciding, per piece of knowledge, what to crystallize into a typed property and what to leave as a memory. The rule I use:

> **The Boundary Rule.** If you would put it in a WHERE clause or a chart, it is a property. If you would put it in a paragraph, it is a memory.

Crystallize when the fact is filterable or countable, when it governs a workflow (consent, stage, tier), when it is stable and enumerable, or when another system needs it synced. Leave it as a memory when it is nuanced, explanatory, one-off, or "the why."

Both sides fail when the rule is ignored. Keep everything as prose and "which Gold accounts renewing in Q1 have an open severity-one ticket?" means reading every record with a model, at a model's cost and error rate. Force everything into fields and you get a property called `insights` that agents fill with paragraphs no workflow can test: the dossier again, wearing a schema.

![The Boundary Rule as two columns. Knowledge you would put in a WHERE clause or a chart becomes a typed property: it is filterable or countable, governs a workflow such as consent, stage, or tier, is stable and enumerable, or must sync to another system, for example stage = 'evaluating' AND renewal_date < '2027-04-01'. Knowledge you would put in a paragraph stays an evidence memory: nuanced, explanatory, one-off, or the why, for example Dana's comment on the September 12 call that June route-sync failures cost her dispatchers two days of manual scheduling. Ignoring the rule fails both ways: everything as prose means every filter reads every record with a model, and everything as fields produces an insights property full of paragraphs no workflow can test.](/images/handbook/ch09-boundary-rule.svg)

*Figure 10.2. The Boundary Rule: filter or chart it, make it a property; explain it, keep it a memory.*

Two design details carry most of the weight.

**The property description is the extractor's instruction.** When a model fills a property from a transcript, the description you wrote is the prompt it follows. A good description has four parts: what to capture, where to find it (signatures, call transcripts, order forms), what it is not, and one to three examples. "Renewal risk" with no description gets you a model's mood. "Renewal risk: the customer's stated or strongly implied intent not to renew, from calls, emails, or tickets; not general dissatisfaction with a feature; example: 'we are putting the renewal out to tender'" gets you a field two agents fill the same way.

**Each property declares how it changes and who may write it.** Some values replace (current role); some accumulate (pain points raised). Some may be written by a model (buying stage); some belong to a system of record and must never be overwritten by extraction (contract dates, consent). A model that "corrects" a contract date from a rep's offhand comment is a governance failure that looks like helpfulness.

## One save, several representations

If four forms have to be maintained by four pipelines, they drift apart. The fix is a **unified write**: one save of raw content produces several representations at once. The transcript is stored and embedded for semantic recall. In the same pass, a model extracts typed property values against the schema and atomic memories that no schema anticipated. Documents saved as documents can optionally feed the same extraction.

In the system I built, a save runs one model call per collection of properties plus one for free-text memories: three collections, about four calls per save. That is the price of thorough extraction.

In our own measurements the two extraction modes cover different ground. Across 20 annotated samples, 34 percent of the valuable information was captured by both typed and free-text extraction, 38 percent only by free-text memories, 12 percent only by typed properties, and 16 percent by neither, for combined recall of 82.8 percent [6]. Treat those as internal numbers from a small sample, not a benchmark. The shape of the result is the point: the schema catches what you anticipated, the free text catches what you did not, and neither alone is enough.

My paper reports a much higher figure, 99.6 percent fact recall with complementary dual-modality coverage [20], and the two numbers are not in conflict because they measure different things. The 82.8 percent is the share of the valuable information in 20 annotated samples that either extraction mode captured in the internal experiment above. The 99.6 percent is the paper's fact recall in controlled experiments (N=250, five content types), with typed and free-text extraction combined. Read the first as a small-sample estimate and the second as a controlled-experiment result; neither is a benchmark of your data.

Write time is also where quality is cheapest to buy. Before a memory enters the store it should be self-contained, have its pronouns resolved ("she" becomes "Dana, Harlow's fleet operations lead"), and have relative time anchored ("recently" becomes "September 2026"). It should also be checked against what is already stored. When we ran five overlapping sources about the same entity through a shared store, 83 percent of the candidate memories duplicated something already there [7]. Without deduplication at write time, the store fills with echoes and each echo makes a fact look more important than it is.

Write-time quality gates are a one-time cost. Read-time noise is a tax you pay on every call, forever.

## Documents are memories too

Short-form memory (atomic facts and properties) is precise and cheap to retrieve. Some knowledge is bigger than a fact and smaller than a database: an account brief, a renewal plan, a record of what was promised. Agents want to write these as sectioned notes and revise one section at a time, the way coding agents keep a notes file.

Long-form memory is not a separate store. It is another form on the same record, with the same identity, permissions, and retrieval call. Its value is compression. In one of our experiments, 1.5 million tokens of source material was compacted into about 27,000 tokens of long-form notes, and the agent still had what it needed [8]. That is one internal run, not a general ratio, but it shows why synthesis earns its place: it is the only form an agent can read whole.

Keep synthesis subordinate to evidence. A beautifully written brief that says "renewal is safe" is an opinion with good formatting. It should cite what it rests on and be regenerated when that changes. (Chapter 11 calls that ongoing work dreaming.)

## Retrieval is a conversation, not a query

Once memory is shaped this way, reading it changes too.

The usual approach hands knowledge to an agent by similarity: chunk everything, embed it, and return whatever looks most like the question. That finds text that resembles the query. It does not assemble what the task needs. The fact that matters often shares no words with the question: "should we offer Harlow a multi-year renewal?" needs the open severity-one ticket, the stakeholder change, and the contract's price-protection clause, none of which look like the word "renewal." Similarity is not relevance.

Retrieval for agents has three properties the chat era did not need.

**One call, several sources, one identity.** A request for Harlow should return typed properties first (they are authoritative), then the most relevant memories, document sections, and graph neighbors, all resolved to the same entity, with citations to specific facts. An agent that calls four stores and joins the results itself rewrites that plumbing in every workflow and, sooner or later, joins two slightly different Harlows.

**A token budget, not a result count.** More is not better. In our end-to-end evaluation of personalized sales emails (10 prospects, 3 runs each, scored on a 100-point rubric), quality rose from 79.5 with no memory to about 86 with governed memory, and most of that gain came from roughly the first seven well-chosen memories per entity; adding more did not help [9]. Small sample, one task, internal rubric. The paper's larger controlled experiments found the same shape: output quality saturated at approximately seven governed memories per entity [20]. The durable lesson matches the context research above: the bottleneck is curation, not capacity. Store generously; inject sparingly.

![A budgeted read is one call scoped to one entity that returns typed properties first because they are authoritative, then the most relevant memories, relevant document sections, and linked graph neighbors, stopping at the token budget while everything else stays stored; the next step in the session returns new items, not repeats. Below, drawn to scale, the chapter's cost arithmetic for 40,000 accounts read 5 to 20 times a month: the full history at 30,000 to 80,000 tokens per read is 6 to 64 billion input tokens a month, while a curated payload of 3,000 to 6,000 tokens is 0.6 to 4.8 billion.](/images/handbook/ch09-budgeted-read.svg)

*Figure 10.3. Store generously, inject sparingly: a budgeted read costs about an order of magnitude less than reading the full history.*

**Continuation across steps.** An autonomous agent does not retrieve once. The first read about Harlow reveals that the admin changed; that makes the contract relevant; the contract reveals a notice period; the plan changes. Retrieval becomes an action reasoning selects, so the memory layer should remember what it already delivered about this record and return the next most useful items, not the same ones again. In one research agent we instrumented, 23 percent of output tokens went to the agent explaining to itself why it had already seen repeated content [10]. The fix was not a better ranker. It was a session that knew what had been shown. In my paper's controlled experiments, progressive delivery across steps cut delivered tokens by 50 percent [20].

The retrieval I built exposes this as modes, and the rule is the cheapest mode that answers: a record scan with no model call, a deterministic property filter, a semantic brief that cites the memories it used, or a session that continues where the last read stopped. "Which Gold accounts renew in Q1?" should never cost a model call.

## Four forms do not need four databases

None of this needs four databases. The self-hosted memory system I built keeps all four forms, plus the governance documents, in one Postgres database with the pgvector extension. Every prior property value is kept with the dates it was valid, so a superseded renewal date is retrievable, not lost (Chapter 11). One database means one backup, one permission model, and one place where "Harlow" is resolved, which is most of what makes a single budgeted read possible.

The schema installs as a bundle: entity types, property collections with descriptions and update semantics (replaceable or append-only), relation types, rules that turn properties into edges (a saved competitor list becomes `competes_with` links), and seed guidelines. The Boundary Rule becomes a reviewable artifact instead of a habit. Reinstalling is harmless, but in my system graph rules cannot be edited after install, so test the bundle in a throwaway namespace first.

Agents reach it through the Model Context Protocol. A session opens with an orientation call that returns the schema, the index of governance documents (names, never contents), and only the tools this agent's key may call. A read-only agent's write tools are absent, not merely refused, and autonomous agents are never offered delete. Each turn then runs one loop: recall, retrieve the rules, act, write back.

A governed personalization engine I built on memory like this used a short list: exact properties for campaign IDs, scores, and status flags; free-text memories for research; batch ingestion; semantic recall; a compiled digest per account; property history for audit; deterministic filters for QA and exports; and bulk updates to write scores back. The hard part is keeping it on one record.

## Memory is one of three questions

Memory is often asked to do jobs that belong elsewhere. Three questions keep them apart. *What is relevant?* is retrieval over a corpus (product docs, past proposals). *What do we know about this customer?* is memory: entity-scoped, persistent, written by many agents. *What are the rules?* is governance: pricing, brand, consent, contact limits, applied whether or not the agent thought to ask.

The common mistake is storing rules inside customer memory; when Larkspur's discount policy changes, it should change in one place, not in 40,000 records. The opposite mistake is serving rules by similarity search, where a draft policy outranks the approved one because it uses the right words. Retrieval ranks by similarity; governance has to rank by authority (Chapter 18). In my paper, tiered governance routing reached 92 percent routing precision [20]. That is a different measure again: of the governance documents routed to a task, the share that were the relevant ones. It says nothing about recalling customer facts.

## What this does not do

An argument for memory that lists only its strengths should not be trusted, so here are the limits, my own numbers included.

**Memory does not fix identity or objectives.** If Chapter 9's identity resolution merged two Danas, unified memory will faithfully give both of them one history. If the decision layer is optimizing for sends instead of outcomes (Chapter 14), better memory produces better-targeted versions of the wrong message.

**The benchmarks do not measure what you need.** In 2025, memory vendors fought publicly over the LoCoMo benchmark [11]. Mem0's paper claimed a 26 percent relative improvement over OpenAI's memory, with large latency and token savings [12]. Zep replied that a simple full-context baseline, just feeding the whole conversation to the model, scored about 73 percent against Mem0's best of about 68, that LoCoMo "lacks questions designed to test knowledge updates," and that Zep's own corrected score was 75.14 [13]. Mem0's CTO in turn alleged an error in Zep's earlier headline figure [14]. Then Letta ran a plain agent with file tools and no specialized memory system and reported 74.0 percent with a small model [15]. All of these are vendor claims, and I report them as claims. The architecture I published scores in the same band: 74.8 percent overall on LoCoMo [20], which is part of why I stopped leading with the number. When a filesystem, a full-context baseline, and several dedicated products land within a few points of each other, the benchmark is mostly measuring context management on one conversation. It is not measuring what breaks in customer memory: many writers, facts that change, and questions that filter across thousands of accounts. LongMemEval is closer, because it tests knowledge updates and abstention, and it found that commercial assistants and long-context models showed a 30 percent accuracy drop on sustained interactions [16]. Build your evaluation from your own freshness and contradiction cases.

**The strongest opposing view deserves an honest answer.** Context windows reached a million tokens in 2024 [17] and keep growing, and Letta's result suggests a filesystem is all you need; so skip the architecture, keep good files, and let the model read. I agree with Letta's conclusion that "memory is more about how agents manage context than the exact retrieval mechanism used" [15]. For one agent on one conversation or one codebase, files and a long window are often right, and I would not overbuild. The argument breaks on the customer. "Which Gold accounts renewing next quarter mentioned a competitor?" is a filter across 40,000 records, not a reading task. A file has no field-level permissions, no system-of-record protection, and no wall between one customer and the next. And a long window does not decide which of two contradictory facts is current; it hands the contradiction to the model on every call, in a context the research [1][2] says it reads less reliably as it grows.

Cost makes the same point with arithmetic. Suppose each of Larkspur's 40,000 accounts is read by agents 5 to 20 times a month. Reading the full history at 30,000 to 80,000 tokens per read is 6 to 64 billion input tokens a month. A curated payload of 3,000 to 6,000 tokens is 0.6 to 4.8 billion. The assumptions carry the weight (read frequency and history length vary widely), and prompt caching changes the price of repeated reads but not the attention cost of a bloated context. The order-of-magnitude gap is the durable part.

**In-model memory may change the picture.** Research such as Titans explores models that learn to memorize at test time [18]. That is a lab direction, not a production option for shared, governed customer records, and I label it a forecast. Even if it matures, an organization will still need a record it can inspect, correct, permission, and delete.

## At scale

At one agent and a hundred customers, almost any memory approach works. At Larkspur's scale, three things change.

**Writers multiply.** Enrichment, support, sales, telemetry, and humans all write about the same account. Without deduplication and precedence rules, the record becomes an argument. Every property needs an owner and a rule for who wins on conflict. The customer is a writer too: a "not interested" pressed in one surface must land here so every surface honors it, or the control reads as broken (Chapter 3). Synthetic research respondents never write here at all.

**Isolation becomes architectural.** Similar accounts in the same industry have similar embeddings. Keeping Harlow's memories out of another customer's retrieval must be enforced by scoping every read and write to a resolved entity identifier, not by semantic distance. In my paper's controlled test, entity scoping leaked nothing across 500 adversarial queries [20]: one test, not a guarantee.

**Schemas become living documents.** Properties nobody fills, or that two agents fill differently, show up only across thousands of records. Measure extraction per property and revise descriptions the way you revise code.

A memory precise enough to serve a customer this well is precise enough to pressure them, which is why governance and privacy sit on this layer rather than beside it (Chapters 18 and 19). Memory is, in the end, a polite word for a model of someone.

## Failure story: The $450K email

The deal was worth $450,000. (The scenario is a composite, told here at Larkspur; the pattern is one I have seen repeatedly.)

In month one, Larkspur's enrichment agent found that a customer's COO was evaluating two competing platforms and had complained publicly about route-sync reliability. It stored this in its own workspace. In month two, the outbound agent sent the account a segment-level email about "modernizing your fleet operations." In month three, the support copilot resolved a severity-one route-sync ticket for the same account and closed it. In month four, the scoring model saw no recent engagement and lowered the account's priority. In month five, a renewal agent pitched the fixed route-sync feature as new. In month six, the customer chose a competitor, and nobody at Larkspur could say why.

Four agents. One account. Zero shared context. Each learned something true; none of it reached the others, because each workflow's memory belonged to the workflow, not to the customer. The difference between the email the COO deleted and the one she would have answered was not model quality. It was whether the agent could read what the organization already knew. The anti-pattern: **Siloed Agent Memory**.

## Patterns

**Memory Layers.**
*Problem:* one memory shape cannot serve filtering, recall, explanation, and connection. *Forces:* agents need exact answers and nuance; tokens are expensive. *Solution:* keep four forms on one record: typed properties, evidence-level memories, long-form synthesis, and a graph, each answering its own question. *Tradeoffs:* more design work up front; some intentional redundancy between forms (overlap is cheaper than loss).

**Boundary Rule.**
*Problem:* teams either trap facts in prose or force nuance into fields. *Forces:* filterability versus expressiveness; schema maintenance cost. *Solution:* WHERE clause or chart, property; paragraph, memory. Every property carries a four-part description, update semantics, and write permission. *Tradeoffs:* schemas need owners and revision; some facts will live in both forms for a while.

**Unified Write, Budgeted Read.**
*Problem:* separate pipelines per form drift apart; agents over-read. *Forces:* extraction cost, consistency, context limits. *Solution:* one save produces all representations with quality gates and deduplication; one read returns properties first, then curated memories, documents, and graph neighbors under a token budget, with session continuation. *Tradeoffs:* extraction cost moves to write time (usually a good trade, since each record is written once and read many times).

## Leader questions

1. When two of our agents act on the same customer, can the second one see what the first one learned? Show me.
2. For our five most important customer fields, who owns each one, where does its value come from, and can a model overwrite it?
3. How many tokens does an agent read to answer one question about one account, and how do we know that number is not mostly noise?
4. Where do our rules live: in one governed place, or copied into prompts and customer records?
5. How do we test whether memory stays correct when facts change, not just whether it recalls them?

## Build checklist

- [ ] Every entity has a resolved identity before anything becomes memory (Chapter 9).
- [ ] Entities are modeled separately from roles; roles and relationships are typed and dated.
- [ ] Each property has a type, a four-part description, update semantics (replace or append), and a write permission (model-writable or system-of-record only).
- [ ] One save produces properties, atomic memories, and (optionally) document content in a single pass.
- [ ] Write-time gates: coreference resolved, time anchored, self-contained, deduplicated against the store.
- [ ] Every memory carries source, observation time, and confidence.
- [ ] One retrieval call returns properties first, then memories, documents, and graph neighbors, scoped to one entity, under a stated token budget.
- [ ] Multi-step agents continue a retrieval session instead of re-reading the same items.
- [ ] Organizational rules live in governance, not in customer records.
- [ ] Reads and writes are isolated by entity identifier, not by embedding similarity.

## Metrics to watch

- **Tokens read per decision**, and the share that were cited or used.
- **Cross-agent reuse:** the share of decisions that used a memory written by a different workflow.
- **Property fill rate and agreement** per property (null rates, disagreement between extractors or with human labels).
- **Duplicate rate at write time**, trending over time.
- **Freshness accuracy** on your own test set of changed facts (role changes, fixed issues, moved dates).

## Reader Q&A

**We already have RAG over our CRM notes. Isn't that memory?**
It is retrieval. It makes the agent well read. It does not maintain a customer's current state, decide which fact is current, or give other systems typed values. Once you add extraction, write-back, deduplication, and supersession, you have built memory that uses retrieval as one component.

**Do we need a graph database to start?**
No. Start with typed, dated, sourced relationships on the record. Move to a dedicated graph when traversal questions become frequent. GraphRAG shows graphs help most on questions spanning a whole corpus [19]; most early customer questions are about one account and its neighbors.

**How many properties should we define?**
Fewer and better than you think: the fields that drive your first three decisions (Chapter 6). Free-text memories will catch what you did not anticipate.

**Is this just a CDP with a vector index?**
A CDP unifies identity and events for activation. Unified Customer Memory sits on that identity and adds extracted evidence, typed properties with extraction instructions, synthesis, and budgeted retrieval. Many companies will keep both.

## For your AI

```yaml
chapter: 10
title: "Memory Built for Agents"
concepts:
  - name: Unified Customer Memory
    alias: Unified Record Memory
    definition: "One record per entity that unifies everything agents can know about it, kept as typed properties, evidence-level memories, long-form synthesis, and graph relations, written in one save and read in one budgeted call by every agent."
  - name: The reader changed
    definition: "Customer data was designed for humans who carry missing context in their heads; agents do not, and pay per token, so data must carry meaning, provenance, narrative, and a budget."
  - name: Memory Layers
    definition: "Properties answer what qualifies and governs; memories answer what is relevant; synthesis answers what is understood; the graph answers what is connected."
  - name: Boundary Rule
    definition: "If you would put it in a WHERE clause or a chart, it is a property; if you would put it in a paragraph, it is a memory."
  - name: Unified write
    definition: "One save of raw content produces embeddings, typed property values, and atomic memories in a single extraction pass, with quality gates and deduplication."
  - name: Property description as extractor instruction
    definition: "Each property's description (what to capture, where to find it, what it is not, examples) is the prompt the extractor follows."
  - name: Retrieval as conversation
    definition: "Multi-step agents continue a retrieval session anchored to an entity, receiving new items rather than repeats, under a token budget."
  - name: Four forms, one database
    definition: "All four forms, plus governance documents and property history, can live in one Postgres database with pgvector, giving one backup, one permission model, and one identity resolution."
  - name: Three questions
    definition: "What is relevant (retrieval), what do we know about this entity (memory), what are the rules (governance) are separate problems with separate infrastructure."
decision_rules:
  - if: "a fact will be filtered, counted, charted, synced, or used to gate a workflow"
    then: "model it as a typed property with a description, update semantics, and write permission"
  - if: "a fact is nuanced, explanatory, or one-off"
    then: "store it as an evidence-level memory with source, time, and confidence"
  - if: "a field is owned by a system of record (contract dates, consent)"
    then: "mark it not model-writable; extraction may propose but never overwrite"
  - if: "an organizational rule would be copied into many customer records"
    then: "store it once in governance and route it to agents by authority, not similarity"
  - if: "an agent retrieves more than once about the same entity in one task"
    then: "continue the retrieval session so it receives new items, not repeats"
  - if: "a question can be answered by filtering typed properties"
    then: "use a deterministic filter read with no model call; reserve model synthesis for questions that need it"
  - if: "memory quality is being evaluated"
    then: "test on your own changed-fact and contradiction cases, not only public recall benchmarks"
assessment_questions:
  - "Which agents or workflows act on the same customers today, and can each see what the others learned?"
  - "Which customer fields drive your first decisions, who owns each, and which may a model overwrite?"
  - "How many tokens does an agent read per decision, and what share is used?"
  - "Where are organizational rules stored, and how many copies exist?"
  - "Is identity resolved before content becomes memory?"
patterns: [Memory Layers, Boundary Rule, Unified Write Budgeted Read]
anti_patterns: [Siloed Agent Memory, The Dossier, Modality Monoculture, Prompt as Memory, Rules in Records]
maturity_dimension: identity_and_memory
```

## References

1. Hong, Troynikov, Huber. "Context Rot: How Increasing Input Tokens Impacts LLM Performance." Chroma, 2025-07-14. https://www.trychroma.com/research/context-rot
2. Liu et al. "Lost in the Middle: How Language Models Use Long Contexts." TACL 2024; arXiv:2307.03172. https://arxiv.org/abs/2307.03172
3. Anthropic. "Effective context engineering for AI agents." 2025-09-29. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
4. Sumers, Yao, Narasimhan, Griffiths. "Cognitive Architectures for Language Agents" (CoALA). TMLR 2024; arXiv:2309.02427. https://arxiv.org/abs/2309.02427
5. Packer et al. "MemGPT: Towards LLMs as Operating Systems." arXiv:2310.08560, 2023. https://arxiv.org/abs/2310.08560
6. Taheri, H. "Dual Memory: Why You Need Both Free-Text Facts and Typed Properties." hamedtaheri.com, 2026-03-14. Internal experiment, 20 samples. https://hamedtaheri.com/articles/dual-memory-free-text-and-typed-properties/
7. Same source as 6 (five-source deduplication experiment). Internal.
8. Taheri, H. "Unified Record Memory, part 3" (LinkedIn, 2026-07). Internal experiment, single run.
9. Taheri, H. "Seven Memories Per Entity Is All You Need." hamedtaheri.com, 2026-03-14. Internal evaluation, 10 prospects x 3 runs. https://hamedtaheri.com/articles/seven-memories-per-entity/
10. Taheri, H. "Retrieval Is a Conversation, Not a Query." hamedtaheri.com, 2026-06-03. Internal instrumentation. https://hamedtaheri.com/articles/retrieval-is-a-conversation/
11. Maharana et al. "Evaluating Very Long-Term Conversational Memory of LLM Agents" (LoCoMo). ACL 2024; arXiv:2402.17753. https://arxiv.org/abs/2402.17753
12. Chhikara et al. "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory." arXiv:2504.19413, 2025. Vendor claims. https://arxiv.org/abs/2504.19413
13. Zep. "Lies, Damn Lies, & Statistics: Is Mem0 Really SOTA in Agent Memory?" 2025-05-06. Vendor claims. https://blog.getzep.com/lies-damn-lies-statistics-is-mem0-really-sota-in-agent-memory/
14. Mem0 CTO, GitHub issue getzep/zep-papers #5, 2025-05-08. Contested claim. https://github.com/getzep/zep-papers/issues/5
15. Letta. "Benchmarking AI Agent Memory: Is a Filesystem All You Need?" 2025-08-12. Vendor claims. https://www.letta.com/blog/benchmarking-ai-agent-memory
16. Wu et al. "LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory." ICLR 2025; arXiv:2410.10813. https://arxiv.org/abs/2410.10813
17. Google. "Our next-generation model: Gemini 1.5." 2024-02-15 (1 million token context in limited preview). https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/
18. Behrouz, Zhong, Mirrokni. "Titans: Learning to Memorize at Test Time." arXiv:2501.00663, 2025. https://arxiv.org/abs/2501.00663
19. Edge et al. "From Local to Global: A Graph RAG Approach to Query-Focused Summarization." arXiv:2404.16130, 2024. https://arxiv.org/abs/2404.16130
20. Taheri, H. "Governed Memory: A Production Architecture for Multi-Agent Workflows." arXiv:2603.17787, 2026-03-18. Author's own controlled experiments (N=250, five content types) and LoCoMo result. https://arxiv.org/abs/2603.17787
