Questions this chapter answers
- Our AI remembers our customers. How do we stop it remembering things that are no longer true?
- When two systems disagree about a customer, which one wins?
- How do we know where a fact came from, and whether we may still use it?
- What should expire, and what should never have been stored?
- Everyone is talking about "dreaming." What is it, and does it apply to customer data?
The short answer
Memory does not decay on its own. It accumulates by default and is corrected only on purpose, so a memory nobody maintains does not fade into irrelevance. It gets larger, more confident, and more wrong at the same time.
Keeping memory true takes four disciplines, and none of them is a model upgrade. Every property gets a freshness policy: how long a value can be trusted before it must be re-verified or stop being used. Every fact carries a provenance envelope (introduced in Chapter 7, extended here): where it came from, when it was observed, when it was valid, which extractor produced it, and whether it is a fact or an inference. Contradictions are resolved by precedence, not by whichever write arrived last. And the memory is curated continuously in the background: duplicates merged, contradictions surfaced, summaries rewritten, stale facts expired.
The field now calls that fourth discipline "dreaming." OpenAI, Anthropic and Letta all use the word, building on research into sleep-time compute, and they apply it to an agent's or an assistant's own memory. This chapter applies it to the customer: shared memory that many agents and teams read, sitting on resolved identity, holding typed facts with their sources, and changed only through reviewable, non-destructive steps.
For a leader, the test is simple. Pick one customer fact your AI used this week and ask: where did it come from, when was it last true, and what would have told us if it stopped being true? If nobody can answer the third, the memory is not being kept. It is being hoarded.
The customer who changed jobs and the memory that did not
Larkspur Systems (a fictional company, a composite of many real ones) sells field-service and fleet software. For three years its main contact at a regional HVAC contractor was Elena, the fleet operations manager. Her calls with Larkspur were about technician scheduling, route sync, and the dispatch board.
Last spring Elena was promoted to run procurement, and the Larkspur renewal became hers to sign. An enrichment feed noticed the new title in May, and the CRM's job-title field changed.
Nothing else did. The account summary Larkspur's agents read before every touch still described Elena as "the operations lead who cares about dispatch efficiency." Forty call-note memories said the same thing, correctly, about the past. The assistant that drafted her renewal outreach in November weighed forty consistent memories against one changed field and wrote her a warm, specific, well-grounded email about the new dispatch board.
Every sentence was traceable to a real source. It was still wrong: the person it was written for no longer exists. I call this The Stale Self: the system stays loyal to a version of the customer that has expired.
Identity resolution, extraction and retrieval all worked. A correction landed in one representation, never propagated, and no process existed whose job was to notice.
Propagation also fails in milliseconds. In a governed personalization engine I built, a successful write did not always mean the next step could read it, so important state gets a bounded read-after-write check before anything renders from it. A correct value that arrives late still produces a wrong page. The self-hosted memory system I built taught the harsher version: under a saturated database, a save once reported success while writing nothing, and a failed embedding once returned success with an empty vector. Both now fail loudly. Success must mean landed.
Memory goes wrong in four ways, and only one is the model's fault
When personalization gets a fact wrong, the instinct is to blame hallucination. That is usually not the reason. In customer memory, errors enter through four doors: the world changes (the fact was true when stored); sources disagree (CRM, enrichment vendor and support transcript each say something different); extraction errs (a model stored the wrong value, or an inference as a fact); and derived views lag (summaries, scores and embeddings were computed from facts that have since changed).
Only the third involves a model misreading anything, and its fix is usually a better property description. The other three are maintenance, which memory tooling tends not to ship. Anthropic's memory tool documentation is candid: the model can create, update and delete memory files, while expiry appears under security considerations as your job, with the advice to "periodically delete memory files that haven't been accessed in a long time" [1]. LongMemEval, the benchmark that explicitly tests knowledge updates, found assistants losing roughly 30% accuracy across sustained interactions [2]. And OpenAI's own illustration, from the launch of ChatGPT's Dreaming feature, is a memory still saying "You're going to Singapore in July" after the trip is over [3].
The writing is automatic. The forgetting is your job.
Every fact has a half-life
A freshness policy states, per property, how long a value can be relied on and what happens when that time runs out. Different facts lose trustworthiness at different rates, and the rate belongs to the kind of fact, not to the database.
People ask me for the "decay rate" of B2B contact data. The often-quoted 22.5% a year has no primary source I can find, so I do not use it [4]. The defensible anchor is labor data: the U.S. Bureau of Labor Statistics reported median employee tenure of 4.1 years in January 2026, with 20.6% of wage and salary workers at their employer for a year or less [5]. Tenure counts employer changes, not role changes. Elena never changed employer, so she would not appear in that statistic at all. Treat it as a floor on how fast contact records go stale.
A starting freshness table for Larkspur. The windows are illustrative, to be calibrated against your own correction data:
| Property | Volatility | Trust window | When the window expires |
|---|---|---|---|
| Legal company name, domain | Low | 12 to 24 months | Re-verify on next enrichment; keep using |
| Contact's employer | Medium | 6 to 12 months | Re-verify before outreach that assumes it |
| Contact's role | Medium to high | 3 to 6 months | Stop naming the role; write for the account |
| Stated priorities, current initiative | High | 60 to 90 days | Phrase as history ("in spring you mentioned...") or drop |
| Open support escalation | Very high | Hours to days | Re-read from the system of record at decision time |
| Contact preferences, opt-outs | None | Until the person changes it | Never expire silently |
The last column is a behavior, not a deletion: most expired facts change how they are used before they are removed. And some facts never go stale on a timer; an opt-out does not weaken because nobody has looked at it for a year.
That is where I part ways with access-based expiry. Deleting what has not been read in a long time is a sensible default for an agent's working notes [1]. For customer memory it is backwards in both directions: a rarely read fact (a do-not-call request) can be the most important thing on the record, and a frequently read fact can be the stalest. Expire by the kind of fact and its evidence, not by traffic.
The rule that falls out is the handbook's accuracy standard in operational form: if a fact used to personalize is older than its freshness policy, re-verify it or fall back to an honestly general statement.
A fact without a source is a rumor with a timestamp
Here is the sentence this chapter turns on: you cannot keep a memory true if you cannot trace each fact back to where it came from. Freshness, precedence, correction, deletion and curation all depend on it.
Each material fact already carries the provenance envelope from Chapter 7: route, source and source reference, observation time, method, confidence, corroboration, permitted purposes, allowed surfaces, and a review date. Keeping memory true adds these fields to it:
| Field | What it answers |
|---|---|
recorded_at | When did it enter memory? (Chapter 7's observed_at says when we saw it.) |
valid_from, valid_to | When was it true in the world? |
extractor | Which model, prompt and schema version turned the source into this value? |
kind | Is this value a fact, or a model's inference or summary derived from other facts? (Chapter 7's route says how it arrived.) |
sensitivity | Does it need stricter handling than its purposes alone imply? |
status | Current, superseded, disputed, or under correction? |
Figure 11.1. Chapter 7's envelope says where a fact came from; the extension says when it was true, who extracted it, and whether it still stands.
Separating when a fact was true from when the system learned it is not pedantry: Zep's temporal knowledge graph is built on exactly this bi-temporal model, and when a new fact contradicts an old one it invalidates the old edge by closing its validity window rather than deleting it [6]. W3C's PROV model offers a standard vocabulary for the same idea [7]. Truth has a time, and so does knowledge of it.
The extractor field is the one teams skip and the one that pays back first. When a model or schema change degrades one property, the extractor version tells you exactly which values to replay. Without it, every correction is a full re-ingestion or a guess.
Append or replace is a decision about truth
Every property in a typed memory schema answers a quiet question: when a new value arrives, does it replace the old one or join it? Replace fits properties with one current value: role, account owner, plan tier, renewal date. Append fits histories: objections raised, products evaluated, events attended.
The mistake is treating replace as overwrite. In a customer system of record, replace should mean supersede: the old value's validity closes and it stays as history. Elena's old role is not wrong. It is last year's truth, and it lets the system say "you ran dispatch, so you know what the board costs to operate," which is what makes a renewal conversation feel continuous rather than robotic.
Figure 11.2. Supersede closes the old value's validity window instead of deleting it: the old role is last year's truth, not an error.
Decide the semantics per property, in its definition. A property whose update rule nobody chose fails silently in one direction or the other.
Documents need the same rule. In the memory system I built, an agent's document update that named no mode once replaced the whole body, and the history the documentation promised held only the mode, not the old text. Now agent updates append by default, the replaced body is snapshotted in the same transaction as any overwrite, and a failed snapshot fails the update.
When sources disagree, precedence decides, not the latest write
Last-write-wins is the default in most pipelines and the wrong default for customer memory. The latest write is often the least authoritative: a nightly enrichment job, an agent's inference from an email signature. Resolve contradictions with Chapter 7's Signal Ladder, applied per property, rather than a fixed ranking of sources:
- Purpose gates first. A value not permitted for this use is out, however good it is.
- Authority depends on the kind of fact. The system of record owns transactions and contract terms (billing for plan tier, contracts for renewal date). The person's declaration owns preferences. The person's own recent artifacts (a signature, a reply, a changed login domain) own circumstances such as role and employer. Third-party values wait until something first-hand agrees with them, and an inference may inform a decision but never outranks a first-hand source.
- Freshness can override authority for properties that change. This is where the freshness windows above do their work: authority holds until its window expires, then it must be re-verified or yield. A fresh first-hand artifact can beat a stale declared title; an inferred channel preference never beats a declared one, however recent.
- Corroborated confidence breaks ties.
Three rules keep this honest. A contradiction is a signal, not only a conflict: a fresh third-party title contradicting a two-year-old declared role most likely means the person's situation changed, which is exactly when the summary should be rewritten. Count independent sources, not agreeing writers: three agents repeating one unverified claim are one source echoing. Generated content is never its own source: a summary the system wrote must not be re-ingested as evidence for the facts it summarized.
How well does automated conflict handling work? In the evaluation I published of the governed memory system I build, across 30 synthetic conflict pairs (a company changing its database, a contact switching roles), the fresh fact appeared in 83.3% of answers, only the fresh fact in 33.3%, and the stale fact alone once [8]. Most of the gap was acceptable: answers led with the new fact and mentioned the old one as trajectory. It is a small synthetic set, an upper bound, not a production rate. The lesson: detection is tractable; what the answer should say about history is a product decision.
Dreaming: curating the customer while nobody is asking
Policies say what is true. Something still has to find the entities whose memory changed, reconcile them, and rewrite the views that depend on them. The field has converged on a name for that work, and neither the name nor the mechanism is mine. Credit, in order:
- Reflection. Generative Agents (Park et al., 2023) periodically synthesized an agent's memory stream "into higher-level reflections" [9].
- Sleep-time compute. Lin et al. (2025) let models "think offline about contexts before queries are presented," cutting test-time compute about 5x at equal accuracy on two reasoning benchmarks and average cost per query 2.5x when amortized across related queries; the benefit tracked how predictable the queries were [10].
- Sleep-time agents. Letta (2025) made the primary agent unable to edit its core memory and gave that job to a background agent [11], later adding git-versioned memory and the phrase "agent dreaming" [12].
- Dreams. Anthropic's Managed Agents research preview (announced May 19, 2026) reads a memory store and up to 100 past sessions and writes a new store, duplicates merged and stale or contradicted entries replaced; the input is never modified, so the output can be reviewed or discarded [13]. Anthropic reports Harvey's completion rates rose about 6x in its tests [14].
- ChatGPT Dreaming (June 4, 2026) synthesizes across a user's conversations and rewrites what the assistant remembers, turning the Singapore plan into "You went to Singapore in July 2026" [3].
The biological metaphor is only a metaphor: complementary learning systems theory pairs fast capture of episodes with slow extraction of structure, and sleep appears to help consolidate memories [15][16]. The useful part is the division of labor: capture fast, reorganize when nobody is asking.
Every system above dreams about one agent or one assistant user: one memory, one owner, one reader. Customer memory is a shared system of record that dozens of agents and several teams read and write, about people who never open the tool. The field taught the agent to dream about itself. The customer is the harder subject, and the one where the value is.
Applying dreaming to the customer adds five requirements:
- It runs across writers, not one agent's notes.
- It sits on identity resolution. Consolidating memory for a wrongly merged entity makes the error permanent and fluent. Curation must respect merge confidence (Chapter 9). In mine, deterministic rules propose duplicate pairs, a model only judges them, verdicts below 0.9 confidence are refused, and every merge can be undone for 90 days by default [21].
- It writes typed properties with provenance, not only prose. A superseded
rolewith a closed validity window is what the decision layer and governance can act on. - It is governed. What a pass may change is policy: merge duplicate observations freely, propose a role change with evidence, never touch consent, contract terms or system-of-record fields.
- It is non-destructive and reviewable. It proposes a new state as a diff that can be reviewed and rolled back.
The dreaming cycle as I recommend building it, and as I built it into a self-hosted memory system (dreaming shipped there on July 30, 2026, after the releases above; the controls that make it safe to leave unattended followed on August 11 [21]):
- Select entities with new evidence, facts past their window, or open contradictions. Never re-dream the whole base by default. Mine runs on a schedule when the job queue is idle, yields to live requests, and re-dreams a record only when it changed or the job's objective did. Curation's own writes never trigger curation. A scope filter must fail closed: mine once silently ignored an unrecognized collection filter and ran against the whole organization; an unknown scope now fails the run.
- Gather the entity's properties, recent memories, summary, and their provenance, plus a preamble: which company it works for (the organization's own context documents), today's date, which is what lets it write "ran fleet operations until spring" instead of "runs," and the stage it is running at. A missing context document does not stop the pass; the miss is reported.
- Consolidate duplicate observations into one memory with all its sources attached.
- Reconcile contradictions with the Signal Ladder, per property; mark the unresolvable as
disputed. - Re-synthesize the summary from current facts, phrasing superseded facts as history.
- Expire what is past its window by superseding it. Deletion does not happen here: the curator I built is never given a delete tool, and erasure runs through its own path.
- Propose, gate, apply. A new curation job only proposes until an operator promotes it. After that, low-consequence changes apply automatically; high-consequence ones queue for review or sit outside the curator's reach; every change is logged with the cycle that made it.
Figure 11.3. Dreaming about the customer curates only changed entities and ends in a gated, reviewable diff, never an overwrite.
Building it taught me three things about step 7. A job earns the right to write. A new one only proposes, five records a pass; after a completed pass an operator may promote it to a pilot (ten records, three writes each), then to full. You approve the judge, not each verdict. An emergency stop drops a job back to propose-only at once, but it disarms writes, not spend: a job still on a schedule keeps running, and consuming model tokens, until someone also disables it. The gate lives in code, not in the prompt. Write permission is checked in the call path on every write, so a model cannot talk its way past it and a stop lands mid-run. Close every path around the gate, not only the main one: a pass scheduled through the general scheduler without a job definition would have run with none of the controls, so that path now refuses it. And the curator writes through identity, never raw ids, after an early pass let a model put an email in an id field and create an orphan record. The system, not the model, stamps each write with its job and run, and a ledger records every applied, proposed and refused write by property name, never value. Correct, do not pile. A curator that appends a note beside a wrong fact makes memory longer, not truer. Mine corrects the current value and closes the old one's validity window: the supersede rule, applied by the curator.
Run the Stale Self through it. The May title change selects Elena. Reconcile finds her declared role past its window, supersedes it, and flags the change as material. Re-synthesize writes: "Elena ran fleet operations until spring and now leads procurement; she owns the renewal." The November email is written for the person who exists.
Forgetting is a feature you have to build
The GDPR's storage limitation principle requires personal data to be kept in identifiable form no longer than its purposes need [17]; Chapter 19 covers the law. The engineering point: deletion must cascade through every derived representation. When a source is deleted or corrected, the properties extracted from it, the memories citing it, the summary built from them and the embeddings of all three must follow. A system that deletes the transcript but keeps the summary sentence it produced has not forgotten anything. Derivation lineage in the provenance envelope is what makes that cascade testable. The curator's run records count: in the system I built, erasing a person also scrubs them. And the cheapest deletion is the write you refused: sensitive inferences and details with no authorized purpose should never enter memory.
The schema needs the same maintenance: descriptions that extracted well under one model drift under the next. Treat it as a living document: replay extraction on a sample, classify each property as correct, missed, low-confidence or inaccurate, revise the weak ones, validate, then backfill using the extractor field [8].
What this does not do
Curation cannot create truth. If the only evidence about Elena is forty notes from her old job, no dreaming pass discovers her promotion; a fresh source has to arrive. Curation also introduces its own errors: a merge can fuse two different facts, and a summarizer can promote a hedge into a certainty. That is why proposals are reviewable. Nor does curation referee itself: two jobs can fight over one property, one retiring a value the other re-derives every night. Mine detects that oscillation and reports it; which job stands down is a person's call.
Be careful with numbers here. Most memory results are vendor-run, and they disagree, as the 2025 LoCoMo dispute among Mem0, Zep and Letta showed (Chapter 10) [18][19][20]. Popular benchmarks mostly measure context management, and almost none measure whether a system notices that a fact stopped being true. Harvey's 6x and OpenAI's reported internal recall gains are claims from the companies involved. Evaluate on your own freshness cases: real customers whose facts you know changed. I have published no correction rate for my own curator, so I quote none; its ledger of applied, proposed and refused writes is where yours should come from.
I also disagree with rewriting memory in place, even correctly. For a personal assistant, "went to Singapore" is a fine rewrite. For a customer system of record, the old state is evidence: it explains past decisions and lets you roll back a bad pass. Supersede; do not overwrite.
At scale
At Larkspur's size (250,000 contacts) the question is how to avoid curating everything. If 1 to 3% of contacts receive material new evidence on a given day, that is 2,500 to 7,500 entity cycles, not 250,000. At an assumed $0.005 to $0.05 per cycle (illustrative; it depends on model, memory size and caching, and should be recomputed with Chapter 12's cost model), the daily bill runs from about $12 to $375. The selection rate carries the weight: a nightly full-base pass multiplies cost roughly 30 to 100 times for little gain. It is sleep-time compute's amortization logic, applied per customer [10]. Granularity matters too: when my curator moved from one run per page to one run per record, the first nightly passes were the ones to watch, bounded by the stage caps until re-dreaming only what changed settled them [21].
Scale also changes ownership. With six teams (marketing, sales, customer success, support, product, events) writing to one memory, the per-property authority rules are an organizational agreement, not a config file, and changes to it deserve the review a schema migration gets.
Failure story: The Echo Summary
An agent reading a sales email inferred that another Larkspur customer "may be evaluating competitors," and stored it correctly labeled as an inference. The nightly summary included it, hedged. Two weeks later a curation pass rebuilt the summary from the previous summary plus new notes, and the hedge fell away. A retention agent read "the customer is evaluating competitors," found it matched the stored inference, counted two sources, and sent a discount to an account that had never considered leaving.
Nothing lied. A guess was re-read by the system that wrote it until it sounded like a fact. The fix: rebuild summaries from source-level facts, never from prior summaries; keep inference labels through every rewrite; count only independent sources. A memory that reads its own summaries will eventually believe them.
Patterns
Provenance Envelope. Problem: facts cannot be corrected, expired or deleted because nobody knows where they came from. Solution: every material fact carries Chapter 7's envelope (route, source, method, observation time, confidence, purposes, surfaces), extended with recorded and valid times, extractor version, kind, sensitivity and status. Tradeoff: heavier writes; pays back on the first correction cascade or model change.
Freshness Policy. Problem: old facts are used with the confidence of new ones. Solution: per-property trust windows with an explicit behavior at expiry. Tradeoff: windows need calibration from real corrections; too short wastes enrichment spend, too long produces Stale Selves.
Non-destructive Dream (Propose, Gate, Apply). Problem: memory accumulates duplicates, contradictions and stale views. Solution: event-selected, entity-scoped background curation that emits a reviewable diff; new jobs start propose-only and an operator promotes them; the write gate lives in code; low-consequence changes auto-apply, high-consequence ones queue; history is kept and the curator holds no delete tool. Tradeoff: review capacity becomes the bottleneck if the gate is too tight.
Leader questions
- Pick one customer fact our AI used this week. Where did it come from, and when was it last verified?
- When a customer changes role, how many days until every agent that writes to them knows?
- For our ten most-used properties, who owns the rule when sources disagree?
- If a customer asks us to delete their data, can we show the summaries and embeddings built from it are gone?
- What does our background curation change without a human, and how would we roll it back?
Build checklist
- Every property declares its update semantic (supersede or append) and a freshness window with an expiry behavior.
- Every material fact carries a provenance envelope, including extractor version and fact kind.
- Authority rules (Chapter 7's Signal Ladder) are defined per property; last-write-wins is off for customer facts.
- Unresolved contradictions are stored as
disputed, not silently resolved. - Summaries are rebuilt from source-level facts, never from prior summaries.
- Curation outputs are diffs with a review gate and a rollback path; new curation jobs start propose-only and an operator promotes them; no path starts a pass without those controls.
- Deletion and correction cascade through every derived representation, with a test that proves it.
- A freshness evaluation set of real, known changes runs on every pipeline or model change.
Metrics to watch
- Stale-use rate: share of personalized outputs that used a fact past its trust window.
- Time to propagate: days from a change in a source to every representation reflecting it.
- Contradiction backlog: open
disputedfacts, and their age. - Curation acceptance rate: share of proposed changes applied after review; a falling rate means the curator is drifting.
- Per-job digest: stage, records covered, writes applied, proposed and refused, failures, quarantined records, and whether a pass awaits review. In the system I built, this is the operator's review surface.
- Cascade completeness: share of corrected or deleted sources with zero surviving derived references.
Reader Q&A
Is dreaming just a batch job with a new name? Partly, and it is fine to say so. The new part is that language models can read unstructured evidence, reconcile it, and rewrite narrative views. The old part (scheduling, idempotency, review, rollback) is ordinary engineering, and skipping it is how curation becomes a new source of errors.
Should the curator use our strongest model? For consolidation, usually not. For reconciling high-consequence contradictions and rewriting account summaries, often yes. Route by step (Chapter 12).
How often should we re-verify? Let the freshness table set the schedule and your corrections tune it. If corrections to a property cluster at 90 days, a 180-day window is too long.
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 11
concepts:
- name: Freshness Policy
definition: "Per-property trust window plus the behavior when it expires (re-verify, rephrase as history, generalize, delete)."
- name: Provenance Envelope
definition: "Chapter 7's per-value envelope (route, source, source_ref, observed_at, method, confidence, corroborated_by, permitted_purposes, allowed_surfaces, review_after), extended for keeping memory true with recorded_at, valid_from/valid_to, extractor version, kind (fact, inference, or derived summary), sensitivity, and status."
- name: Per-property precedence
definition: "Contradictions are resolved by applying Chapter 7's Signal Ladder to each property: purpose gate, authority by kind of fact, freshness overriding authority for volatile properties, corroborated confidence for ties; a fresh first-hand artifact can beat a stale declared value."
- name: Supersede, not overwrite
definition: "Replace semantics close the old value's validity window and keep it as history rather than deleting it."
- name: Dreaming (applied to the customer)
definition: "Continuous, event-selected background curation of shared customer memory: consolidate, reconcile, re-synthesize, expire; governed, identity-aware, provenance-preserving, non-destructive and reviewable. Builds on reflection (Park et al. 2023), sleep-time compute (Lin et al. 2025), Letta sleep-time agents, Anthropic Dreams and ChatGPT Dreaming."
- name: Correction Cascade
definition: "A correction or deletion of a source propagates to every derived property, memory, summary and embedding, including the curator's own run records."
- name: Curation Stage Ramp
definition: "Each curation job's stage caps what it may do (propose-only, pilot, full); only an operator promotes, one stage at a time, after a completed pass; an emergency stop always works; the gate is enforced in code, not the prompt."
decision_rules:
- if: "a fact used to personalize is older than its freshness window"
then: "re-verify it or fall back to an honestly general statement"
- if: "two sources disagree about a property"
then: "apply Chapter 7's Signal Ladder for that property (purpose gate, authority by kind of fact, freshness override for volatile properties, corroborated confidence for ties); store unresolved conflicts as disputed"
- if: "a higher-authority value is past its freshness window and a lower-authority source contradicts it"
then: "treat the contradiction as a likely change; re-verify and trigger re-synthesis of dependent summaries"
- if: "several agents report the same claim"
then: "count independent original sources, not agent writes"
- if: "content was generated by the system (summary, inference)"
then: "never ingest it as evidence for the facts it was derived from"
- if: "a curation pass would change a high-consequence property (consent, contract, identity merge)"
then: "queue the change for review; do not auto-apply"
- if: "a source is deleted or corrected"
then: "cascade the change to all derived representations and verify none survive"
- if: "a new background curation job is created"
then: "start it propose-only; promote one stage at a time after reviewing a pass; give it no delete tool"
- if: "a downstream step renders from state that was just written"
then: "await the write and run a bounded read-after-write check before rendering"
- if: "a curation job is stopped in an emergency"
then: "also disable or unschedule it; a stop disarms writes but a scheduled job keeps spending"
- if: "a curation scope filter does not resolve"
then: "fail the run; never widen the scope to the whole base"
assessment_questions:
- "Which customer properties have a defined trust window, and what happens when it expires?"
- "Can you trace any stored customer fact to its source, observation time and extractor version?"
- "What rule decides which system wins when CRM, enrichment and conversation data disagree?"
- "Does any process rewrite account summaries when underlying facts change? How is it selected, reviewed and rolled back?"
- "Can you prove a deleted source no longer influences summaries or embeddings?"
patterns: [Provenance Envelope, Freshness Policy, Non-destructive Dream]
anti_patterns: [The Stale Self, The Echo Summary, Last-Write-Wins, Access-Based Expiry for Customer Facts, Overwrite in Place, Silent Success, Oscillating Curators]
maturity_dimension: identity_and_memoryReferences
- Anthropic. "Memory tool." Claude Platform documentation (accessed 2026-09-26). https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool
- Wu, D. et al. "LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory." ICLR 2025. arXiv:2410.10813. https://arxiv.org/abs/2410.10813
- OpenAI. "Dreaming: Better memory for a more helpful ChatGPT." 2026-06-04. https://openai.com/index/chatgpt-memory-dreaming/ (primary page not directly retrievable in this research pass; details confirmed via secondary coverage)
- The "22.5% per year" B2B data decay figure is commonly cited (for example by HubSpot, attributing MarketingSherpa); no primary study located. Not used as fact.
- U.S. Bureau of Labor Statistics. "Employee Tenure in 2026." Released 2026-09-24. https://www.bls.gov/news.release/tenure.nr0.htm
- Rasmussen, P. et al. "Zep: A Temporal Knowledge Graph Architecture for Agent Memory." arXiv:2501.13956, 2025. https://arxiv.org/abs/2501.13956
- W3C. "PROV-DM: The PROV Data Model." W3C Recommendation, 2013. https://www.w3.org/TR/prov-dm/
- Taheri, H. "Governed Memory: A Production Architecture for Multi-Agent Workflows." arXiv:2603.17787, 2026. https://arxiv.org/abs/2603.17787 ; results summarized in "99.6% Fact Recall, 74.8% on LoCoMo: What the Numbers Actually Mean," hamedtaheri.com, 2026-03-14.
- Park, J. S. et al. "Generative Agents: Interactive Simulacra of Human Behavior." UIST 2023. arXiv:2304.03442. https://arxiv.org/abs/2304.03442
- Lin, K. et al. "Sleep-time Compute: Beyond Inference Scaling at Test-time." arXiv:2504.13171, 2025-04-17. https://arxiv.org/abs/2504.13171
- Letta. "Sleep-time Compute." 2025-04-21. https://www.letta.com/blog/sleep-time-compute
- Letta. "Context Repositories: Git-based Memory for Coding Agents," 2026-02-12, https://www.letta.com/blog/context-repositories/ ; "Memory Models: Towards Agents That Learn," 2026-06-25, https://www.letta.com/blog/towards-agents-that-learn/
- Anthropic. "Dreams." Claude Managed Agents documentation, research preview (accessed 2026-09-26). https://platform.claude.com/docs/en/managed-agents/dreams
- Anthropic. "New in Claude Managed Agents." 2026-05-19. https://claude.com/blog/new-in-claude-managed-agents
- Kumaran, D., Hassabis, D., McClelland, J. L. "What Learning Systems do Intelligent Agents Need? Complementary Learning Systems Theory Updated." Trends in Cognitive Sciences 20(7), 2016. https://web.stanford.edu/~jlmcc/papers/KumaranHassabisMcC16CLSUpdate.pdf
- Diekelmann, S., Born, J. "The memory function of sleep." Nature Reviews Neuroscience 11:114-126, 2010. https://www.nature.com/articles/nrn2762
- Regulation (EU) 2016/679 (GDPR), Article 5(1)(e). https://eur-lex.europa.eu/eli/reg/2016/679/oj
- Chhikara, P. et al. "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory." arXiv:2504.19413, 2025. https://arxiv.org/abs/2504.19413
- Zep. "Lies, Damn Lies, & Statistics: Is Mem0 Really SOTA in Agent Memory?" 2025-05-06. https://blog.getzep.com/lies-damn-lies-statistics-is-mem0-really-sota-in-agent-memory/ ; Mem0 response, GitHub issue getzep/zep-papers #5, 2025-05-08. https://github.com/getzep/zep-papers/issues/5
- Letta. "Benchmarking AI Agent Memory: Is a Filesystem All You Need?" 2025-08-12. https://www.letta.com/blog/benchmarking-ai-agent-memory
- Author's release notes for the self-hosted memory system in the text: dreaming 2026-07-30; stages, promotion gate, stop, ledger and per-record runs 2026-08-11; merge pass 2026-09-07; save and document-history fixes 2026-08-04 and 2026-09-16; checked against the 2026-09-16 release. [PRODUCT-NAMING: link the public changelog if the product is named.]
This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised