Questions this chapter answers
- Where does the money in an AI memory system actually go?
- Our pilot was cheap and production is not. Why does cost climb with adoption?
- Should we just use a cheaper model?
- What is the right unit for budgeting AI personalization?
- Which savings are safe, and which quietly cost us accuracy?
The short answer
A memory system is not expensive because it stores a lot. Storage is the cheap part. It is expensive because of reading: a model reads raw sources to build the memory, and then agents read the memory, plus instructions, plus their own earlier steps, on every call they make. In the pipelines I instrument, the input tokens outnumber the output tokens by roughly 33 to 1 on enrichment and 77 to 1 on generation. The writing is a rounding error. The reading is the product.
That gives leaders four levers, and none of them is "pick the cheapest model":
- Read raw data once, at write time, on the cheap path. Extraction that can wait an hour can run through provider batch APIs at half price, with prompt caching on top.
- Serve compact memory, not raw history. An agent that needs an answer should receive a cited paragraph, not forty documents.
- Lay out every prompt so the stable part is cached. Instructions and policy first, this record last.
- Route each step to the smallest model that passes it. Judgment on the strong model, extraction and classification on small ones.
Budget in cost per accepted decision, not cost per token or per seat. That unit includes retries, the strong-model reading, background curation, and the human review that is often larger than the model bill. Metered AI behaves like electricity, not like software licenses: cost rises with use. If you budget it like seats, success will look like an overrun.
The caveat: each saving can also cut accuracy. Measure cost and acceptance together, per step.
The bill is the reading
I meter input and output tokens separately on every step of the pipelines I run. The first time I saw the split, I assumed a bug. An enrichment step read about 10,551 tokens to write 322. A content-generation step read about 54,161 tokens to write 704. That is simply what the work looks like.
Figure 12.1. Drawn to scale, output is a sliver; what a model reads is the bill.
Three structures showed up, and each is a design decision. There is a floor: a six-word test prompt still billed about 7,700 input tokens, because instructions and guidelines load before the question does. There is re-sending: inside one multi-step generation, per-step input climbed from 10,033 to 13,749 tokens while each step added only 700 to 850 new ones; about 86 percent of that call's billed input was text the model had already read. And the records rhyme: across one campaign, shared instruction context ran near 5,900 tokens per record while the genuinely unique data ran near 190.
Now picture Larkspur Systems (a fictional composite company used throughout this handbook: field-service software, about 40,000 customer accounts and 250,000 contacts). Its renewal program started as a pilot: a renewal agent on 300 customer accounts, costing a few hundred dollars a month. The team extends it to all 40,000 customer accounts and adds a research agent and a review agent. The month-two invoice is not 130 times the pilot; it is larger, because every agent re-reads the same account history, at every step, on the strongest model available. Nobody misused anything. Call it the re-read tax.
Where the money actually goes
A memory-backed personalization system spends in four places.
Write time: extraction. A model reads calls, emails, tickets and documents and turns them into typed properties, evidence-level memories and summaries (Chapter 10). Per record, this is the largest single model call the memory layer makes. In my own systems, the LLM cost of extracting a memory is on the order of fifty times the cost the memory layer spends serving that record back over its life.
Read time: the consuming agent. This is where the 50:1 above can mislead you. Serving memory from a database is cheap; what is expensive is the agent that receives it and reads it, with its instructions, at premium rates, on every step. Across the whole system, including the agents, this read side usually dominates, and it grows with every agent you add. Anthropic reports that its agents use about 4 times the tokens of a chat interaction and multi-agent systems about 15 times [1]. Both statements are true at once: the memory layer's own bill is dominated by writes, and the system's bill is dominated by what agents read.
Background: curation. Consolidating, deduplicating and re-synthesizing (the "dreaming" cycle in Chapter 11) is model work nobody waits for, so it can run on the cheapest path.
Outside the model. Data subscriptions, identity resolution, storage, and above all human review, which (as the worked example shows) can cost more per decision than all the tokens combined.
Here is the sentence this chapter turns on: memory saves money not by storing more, but by being read less: read raw once, cheaply, and then served compact to every expensive reader after that.
Figure 12.2. Memory saves money by being read less: read raw once on the cheap path, then serve a compact brief to every expensive reader.
Cost per accepted decision is the unit
Tokens are the wrong unit because nobody knows what a token is worth. Seats are wrong because a metered system costs more exactly as it becomes useful. Microsoft reportedly wound down most internal Claude Code licenses in one large engineering group in 2026 and moved engineers to its own tool; the same reporting cites Uber engineers spending $500 to $2,000 a month each on tokens and exhausting a year's AI coding budget in four months [2]. Nothing there was misuse. It was a metered tool on a seat-shaped budget.
The unit that survives is the decision that shipped: one person, one chosen action (including "do nothing"), accepted by your quality bar. The planning shape:
cost per accepted decision = (write share + read cost + curation share + review share) / acceptance rate
Dividing by acceptance rate matters more than it looks. On one production batch I ran, fewer than a thousand records became about 3,700 generation runs once retries were counted. Every failed attempt is paid for and produces nothing.
Worked example (Larkspur, ranges, list prices as of September 2026). One decision per account per month: "what, if anything, should happen next on this renewal?" That is 40,000 decisions a month. Prices are one provider's published list rates: a small model at $1/$5 per million input/output tokens, a mid-tier at $2/$10, a frontier model at $4/$20, batch at 50 percent off, cache reads at 0.05 to 0.1 times the input rate [3][4]. Every volume below is an assumption, stated so you can replace it.
| Component | Assumption | Naive design | Engineered design |
|---|---|---|---|
| Write (per account-month) | 5k to 20k tokens of new source material, 4k-token schema prefix, mid-tier model | Sync, no cache: $0.03 to $0.06 | Batch plus cache: $0.01 to $0.03 |
| Read (per decision) | 4 agent steps on the frontier model | Raw history re-read each step, 60k to 200k input tokens: $0.26 to $0.84 | 2k to 3k-token memory brief, stable prefix cached, 30k to 50k input mostly cache reads: $0.04 to $0.10 |
| Curation (per account-month) | Weekly pass on the 30 percent of accounts that changed, batch | None (memory decays instead) | $0.01 to $0.04 |
| Token cost per decision | $0.29 to $0.90 | $0.06 to $0.17 | |
| Human review (per decision, averaged) | 2 to 10 percent sampled, 2 to 4 minutes each, $60 to $100 loaded hourly cost | $0.04 to $0.67 | $0.04 to $0.67 |
At 40,000 decisions a month, the naive design spends roughly $12,000 to $36,000 on tokens and the engineered one roughly $2,400 to $6,800. Before dividing by acceptance rate.
Figure 12.3. Architecture shrinks the token cost per decision; human review stays the same and can become the largest line.
Two assumptions carry the weight: tokens read by the expensive model per decision, and the review rate, which can exceed the whole token bill and does not fall when models get cheaper.
Note what the table does not say. A B2B renewal decision is worth far more than $0.90; token cost rarely makes one decision unprofitable. What kills programs is the curve: cost equals entities, times decisions, times agents per decision, times steps, times tokens per step, times price. Every factor is rising at once. That growth is multiplicative, not exponential (I have called it exponential in my own writing; the word is wrong, and the multiplicative version is bad enough).
Pay again only for what failed
Retries are the factor you control most directly. The rule: a retry buys back the smallest thing that failed. In a governed personalization engine I built, that is running code [21]:
- Conditional best-of-two. The first email set is scored deterministically (blank, bad length, repeated openings, missing approved proof). A clean set ships; only a flawed one triggers a second generation, and the lower-penalty set wins. Better quality without always paying twice.
- Targeted recovery. Failed sequence slots are regenerated, not the sequence; a missing page is rebuilt from the preserved research and score, not the whole record.
- Delivery retry without regeneration. If the endpoint is down, resend the exact completed payload. Regenerating pays twice and returns a different answer.
- Circuit breaking and idempotency. A failing research source stops being called for the rest of the batch; a duplicate request claims the existing work instead of paying again.
Read once, at write time, at a discount
Extraction is the easiest place to save: CRM imports, nightly ingest, backfills and re-extraction after a schema change can all tolerate an hour.
Providers price for that. As of September 2026, Anthropic's Message Batches API and OpenAI's Batch API both charge 50 percent of the synchronous rate, with results within 24 hours and most Anthropic batches finishing in under an hour [4][5]. Prompt caching stacks on top: the schema, extraction instructions and examples are identical across thousands of calls, so after the first write they are read from cache at around a tenth of the input price [3][6]. On extraction workloads where the cacheable prefix is 60 to 85 percent of input, I measure the combined path at roughly 35 to 40 percent of the synchronous price, for the same extraction [7].
The cost is latency variance. "Within 24 hours" is a ceiling, not a promise, and Anthropic's own docs say cache hits inside batches are best effort, ranging from 30 to 98 percent depending on traffic [4]. So batch belongs to queued work, and sync to writes a live agent is waiting on (Chapter 17). The engineering job is making the discount reachable through the same call: one flag on the write, not a second integration. In a self-hosted memory system I built, a failed batch is also requeued to the normal worker, so the discount never costs a lost write [21].
The cheapest extraction is the one you skip. A value already known exactly (a campaign ID, a score, a renewal date from billing) is written as a typed property, never passed through a model: no call, and no probabilistic error added to data that was right [21]. Then scope the rest: in that system a save runs one extraction call per property collection plus one for free-form memories, so limiting a save to the collections its source can inform is a direct cut.
A prompt is a layout
Caching turns a prompt from a message into a layout: a cache only matches an identical prefix, so anything that varies early invalidates everything behind it.
So order by volatility. Stable first: system instructions, governance and brand policy, the extraction schema, examples. Then slower-moving context: the account brief. Last: this step, this question, this new email. On the workload described earlier, that ordering alone put the cacheable ceiling between 85 and 94 percent of billed input, and the model never saw a different problem.
A timestamp or a user name in the first line of a system prompt silently defeats the whole scheme. A short prompt is not a cheap prompt; a well-ordered one is. Fan-out makes this concrete: in my memory system the shared curation preamble sits in the system prompt so every per-record pass re-sends an identical prefix, and some providers cache only where the request marks an explicit breakpoint, so verify the hit rate rather than assuming it.
The context window is a budget, not a bucket
The strongest argument against stuffing context is not cost. It is accuracy. Liu and colleagues showed that models use information at the start and end of a long context better than information in the middle [8]. Chroma's 2025 study of 18 models found performance "consistently degrades with increasing input length," even on simple tasks, and that "even a single distractor reduces performance" [9]. A bigger window is a bigger invoice and, past a point, a worse answer.
Treat context as a budget per step. A workable allocation for a personalized decision:
| Slot | Content | Typical size | Cached? |
|---|---|---|---|
| Instructions and policy | Role, rules, output contract | 3k to 6k | Yes |
| Entity brief | Synthesized memory with citations | 1k to 3k | Per record, across steps |
| Task evidence | The few items this step needs | 0.5k to 3k | No |
| Working history | Compacted prior steps | Capped | Partly |
| Output reserve | The answer | 0.5k to 2k | n/a |
Three techniques keep inside it.
Compaction and synthesis. Summarize prior steps and raw history into short, cited state. Anthropic reports that context editing cut token use by 84 percent on a 100-turn internal evaluation [10]; vendor-internal, but directionally consistent with what I see. Chapter 10's finding that quality saturated at around seven well-chosen memories per entity points the same way.
Progressive delivery. Do not re-send what the agent already has. Tracking which policies were delivered in a session, and sending only what is new, cut token use by about half across our multi-step workflows with no measured quality loss, entirely by removing repetition [11].
Intent over query. When an agent asks memory a question, let the memory layer answer it on a cheap model and return a cited paragraph, rather than returning twenty raw items for the premium model to compress. In one worked comparison, the premium model's cost per call fell to 5 to 10 percent of the query pattern, and over a 30-call loop the agent carried about 7,500 tokens of retrieved context instead of about 90,000 [12].
Right model, right step
The price spread is the reason routing works. As of September 2026, one provider's list prices run from $1 to $10 per million input tokens between its smallest and largest general models, a tenfold gap, with the same gap on output [3]. FrugalGPT reported matching the best single model with up to 98 percent cost reduction by cascading cheaper models first [13]. RouteLLM reported cost cuts of more than two times without quality loss, using routers trained on preference data [14].
Take those as existence proofs measured on public benchmarks, not forecasts for your steps. What transfers is the principle:
- Small models for narrow, well-specified steps: classification, extraction, routing, change detection.
- The strong model where judgment lives: the ambiguous decision, the contradiction, the irreversible action.
- The more precisely a step is specified, the smaller the model it can run on.
- Measure success rate per step, not per agent; an average hides the step that is quietly failing.
Make it configuration, not code. In my memory system each function (extraction, recall, synthesis, generation, embeddings) names its own model, and so does each curation job: cheap for high-volume deduplication, stronger for deciding whether two records are one person [21]. Any function can run on a local model, which removes the per-token bill, not the cost: you pay in hardware and in the concurrency one local instance cannot sustain.
A cheap model that fails half the time is the most expensive option you have: you pay for the call, the retry, and the human who cleans up. Also check the unit you are comparing. Anthropic notes that its newer tokenizer produces about 30 percent more tokens for the same text [3]; a lower price per token is not automatically a lower price per task.
Curation runs on cheap time
Curation should run on the cheapest path you have: batch, cached, smaller models, only on entities that changed. Lin and colleagues' sleep-time compute work reports about five times less test-time compute at equal accuracy, and about 2.5 times lower amortized cost per query, when predictable questions are prepared offline [15]. The benefit scales with how predictable the questions are, and customer memory is unusually predictable: the same agents ask about the same accounts every week. Curation also pays back at read time, because a curated brief is shorter than the history it replaces.
Cheap time still needs a budget, because nobody watches background work run. Most controls in my memory system's idle-time consolidation (Chapter 11) are cost controls [21]. It runs only when the queue is quiet and yields to live requests. A record is reviewed again only if it, or the job's objective, changed. A record that fails three passes in a row is quarantined for a week, so one poison record cannot burn tokens nightly. One lesson: moving to one pass per record gave each record its own write budget and audit trail, and made the first nightly runs cost more until convergence settled them. Another: the emergency stop disarms a job's writes, not its schedule. A still-enabled job keeps running in propose-only mode, paying for model calls on every pass, until someone also disables it. A kill switch for writes is not a kill switch for spend.
What this does not do
Token savings are not accuracy. Memory vendors advertise compaction. Mem0's paper reports more than 90 percent token savings against full context [16], but the LoCoMo results behind such claims are disputed in both directions, and a plain full-context baseline held its own (Chapter 10) [17][18]. The lesson I take is narrower: when histories are short, full context is a perfectly good design, and compaction is an accuracy risk you must measure, not a free win.
Cheaper tokens do not mean a smaller bill. Epoch AI estimates that the price of reaching a fixed capability has fallen between 9 and 900 times a year, depending on the task [19]. Spend still rises, because agents and fleets use more tokens. Planning on price declines to fix an architecture is planning on the wrong curve.
At scale
Across millions of entities and many teams, three things change. Cost becomes an attribution problem: without per-step, per-agent metering, nobody knows whose agent is re-reading the account history. Shared memory becomes the main saving: six teams reading one curated brief beats six teams re-deriving it. And concurrency becomes a quality variable. Past a provider's limits, calls often come back marked successful with degraded work; in our pipelines visible abort rates rose from about 2 percent at forty concurrent generations to about 19 percent at eleven thousand [20]. Ramp in steps and hold the acceptance bar constant.
Failure story: The Cheap-Model Swap
Larkspur's renewal program (fictional) triples its budget in a quarter. The Friday-meeting fix: move every step to the smallest model. The token bill drops by about 80 percent. Over the next six weeks the acceptance rate falls, retries double, the review queue grows, and two renewal emails cite a contract term the account never had. Measured per accepted decision, the program now costs more than before, and it is less trusted.
The mistake was optimizing the price of a token instead of the cost of a decision. The fix was boring: route per step, batch the extraction, serve compact briefs, keep the strong model where judgment lived.
Patterns
Read Once, Serve Compact. Problem: every agent re-reads raw history. Forces: agents need depth; depth is expensive to read repeatedly. Solution: extract at write time on the batch path, curate in the background, serve a cited brief; raw sources stay one lookup away. Tradeoffs: write-time latency; compaction can drop what mattered, so measure accuracy against a full-context baseline.
Stable-First Layout. Problem: repeated instructions billed at full price. Solution: order prompts by volatility, stable to volatile, and keep variable content out of the prefix. Tradeoffs: discipline across teams; cache lifetimes are short.
Right Model, Right Step. Problem: one model for the whole job overpays or underperforms. Solution: specify steps tightly, route narrow steps to small models, escalate on low confidence, measure success per step. Tradeoffs: more moving parts; routers need evaluation on your data.
Leader questions
- What is our cost per accepted decision today, including retries and human review?
- What share of our input tokens is text the model has already read?
- Which of our AI writes are genuinely waiting on a user, and which could run in batch?
- Is our AI budget shaped like seats while the bill is shaped like a meter?
Build checklist
- Meter input and output tokens separately, per step, per agent, per decision.
- Route every non-interactive write through batch.
- Order every prompt stable-first; lint for variable content in the prefix.
- Serve entity briefs with citations instead of raw retrieval to premium models.
- Assign a model per step, with a per-step success metric.
- Write known structured values as typed properties; never send them through extraction.
- Retry the smallest failed unit: the slot, the artifact, or the exact completed payload.
- Run curation on batch, only for changed entities.
- Ramp batch sizes in steps with a fixed acceptance definition.
Metrics to watch
- Cost per accepted decision (and its spread by decision type).
- Cache-read share of billed input tokens.
- Input tokens read by the strongest model per decision.
- Retry amplification: runs per accepted output.
- Review cost as a share of total cost per decision.
Reader Q&A
Won't million-token windows make all this unnecessary? No. Long windows are billed per token, and accuracy still degrades with length and distractors [8][9]. Long windows are useful for rare, deep reads, not as a default.
Is batch safe for customer-facing work? For the write that prepares memory, usually. For a reply a customer is waiting on, no.
Why not just negotiate a volume discount? Do that too. It lowers the price; architecture lowers the quantity, and quantity is the part that grows.
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 12
concepts:
- name: Cost per accepted decision
definition: "Total cost (write share, read, curation share, human review, retries) divided by decisions that pass the quality bar."
- name: The bill is the reading
definition: "In LLM systems, input tokens (what models read) dominate cost, usually by one to two orders of magnitude over output."
- name: Context budget
definition: "A per-step allocation of tokens across instructions, entity brief, task evidence, working history and output, treated as a limit rather than a container."
- name: Right model, right step
definition: "Assign the smallest model that passes each well-specified step; reserve the strongest model for judgment."
- name: Read once, serve compact
definition: "Extract raw sources once at write time on the cheapest path, then serve curated, cited briefs to every consuming agent."
decision_rules:
- if: "a write is not awaited by a user or live agent"
then: "route it through a provider batch API with prompt caching"
- if: "prompt content varies per record or step"
then: "place it after all stable instructions and policy"
- if: "a premium agent needs an answer about an entity"
then: "request a cited synthesized brief, not raw items"
- if: "a model change lowers token cost"
then: "accept it only if acceptance rate per step is unchanged or better"
- if: "a record partly fails or its delivery fails"
then: "retry only the failed unit or resend the completed payload; never regenerate the whole record"
- if: "a value is already known exactly from a system of record"
then: "write it as a typed property without model extraction"
- if: "histories per entity are short"
then: "prefer full context over compaction until measured otherwise"
assessment_questions:
- "What is your cost per accepted decision, and does it include human review and retries?"
- "What share of input tokens is cached or previously read?"
- "Which steps run on which models, and what is each step's success rate?"
- "What share of writes run synchronously that could run in batch?"
- "Is your AI budget seat-shaped or usage-shaped?"
patterns: [Read Once Serve Compact, Stable-First Layout, Right Model Right Step]
anti_patterns: [The Cheap-Model Swap, The Re-Read Tax, Seat-Shaped Budget, The Infinite Context Problem, Replay the Whole Pipeline]
maturity_dimension: identity_and_memory
prices_as_of: "2026-09"References
- Anthropic, "How we built our multi-agent research system," June 2025. https://www.anthropic.com/engineering/multi-agent-research-system
- The Next Web, "Microsoft's Claude Code retreat" (citing Windows Central, The Verge, The Information), 2026. https://thenextweb.com/news/microsoft-claude-code-retreat-ai-cost
- Anthropic, Pricing (model, caching, batch, tokenizer note), accessed 2026-09-26. https://platform.claude.com/docs/en/about-claude/pricing
- Anthropic, Batch processing docs, accessed 2026-09-26. https://platform.claude.com/docs/en/build-with-claude/batch-processing
- OpenAI, Batch API guide, accessed 2026-09-26. https://developers.openai.com/api/docs/guides/batch
- OpenAI, Prompt caching guide, accessed 2026-09-26. https://developers.openai.com/api/docs/guides/prompt-caching
- H. Taheri, "65% Off Memorization," 2026-06-03. https://hamedtaheri.com/articles/65-percent-off-memorization/
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts," TACL 2024. https://arxiv.org/abs/2307.03172
- Hong, Troynikov, Huber, "Context Rot," Chroma, 2025-07-14. https://www.trychroma.com/research/context-rot
- Anthropic, "Managing context on the Claude Developer Platform," 2025-09-29. https://claude.com/blog/context-management
- H. Taheri, "Progressive Context Delivery," 2026-03-14. https://hamedtaheri.com/articles/progressive-context-delivery/
- H. Taheri, "Think and Execute For Me," 2026-06-03. https://hamedtaheri.com/articles/think-and-execute-for-me/
- Chen, Zaharia, Zou, "FrugalGPT," 2023. https://arxiv.org/abs/2305.05176
- Ong et al., "RouteLLM," ICLR 2025. https://arxiv.org/abs/2406.18665
- Lin et al., "Sleep-time Compute," 2025. https://arxiv.org/abs/2504.13171
- Chhikara et al., "Mem0," 2025. https://arxiv.org/abs/2504.19413
- Zep, "Lies, Damn Lies, & Statistics: Is Mem0 Really SOTA?" 2025-05-06. https://blog.getzep.com/lies-damn-lies-statistics-is-mem0-really-sota-in-agent-memory/
- Letta, "Benchmarking AI Agent Memory," 2025-08-12. https://www.letta.com/blog/benchmarking-ai-agent-memory
- Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks," 2025-03. https://epoch.ai/data-insights/llm-inference-price-trends
- H. Taheri, first-party production measurements (fan-out abort rates, retry amplification, token anatomy), 2026-08 to 2026-09; unpublished.
- H. Taheri, first-party system documentation: capabilities reference for a governed personalization engine (2026-08-25, live capabilities only) and release documentation for a self-hosted governed memory system (release 0.8.2, verified 2026-09-16); unpublished. [PRODUCT-NAMING: name the systems here if Hamed decides to]
This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised