# 13. Agents That Run for Weeks

*The question: How do long-running agents use and maintain memory?*

> **Questions this chapter answers**
> - Our agent works fine in a demo. Why does it fall apart on a job that takes six weeks?
> - What exactly should an agent remember between sessions, and where should that live?
> - If the agent crashes on step nine, does it come back at step nine or step one?
> - How do agents hand work to each other, and to people, without losing what they learned?
> - Should we split the job across many subagents to make it faster?

## The short answer

An agent that works a customer over weeks is not a long conversation. It is a case: a goal, a plan, commitments made along the way, and a customer who keeps changing. The context window cannot hold any of that for long, because every session, crash, deploy, and handoff empties it. So the design question is not "how big is the context" but "where does the case live when the agent is not running?"

The answer is three kinds of memory with three lifetimes. Working memory is the context window, and it lasts one session. Task memory is the case file: the plan, which steps are done, what was promised, what is still open. It lasts as long as the job. Entity memory is what the organization knows about the customer, and it outlives every agent that touches it. Most teams build the first, borrow the third, and forget the second entirely. That missing middle is why agents restart from zero.

Three rules follow. Persist the outcome of every step, so a failure resumes instead of restarting. Hand off state, not reports: agents and people read and write the same case file. And fan out to isolate work, not to go faster, because each extra agent multiplies cost and each shared write can multiply one error.

For a leader, the test takes a minute. Ask your team to kill a running agent in staging. If it comes back at step one, state lives in the wrong place, and every crash is billed twice.

## Week four, and the agent starts over

Larkspur Systems (a fictional company, a composite of many real ones) runs a renewal agent on its larger fleet accounts. Take one of them, a regional haulage operator we will call Kestrel Haulage (also fictional). The renewal is six weeks out.

In week one the agent does good work. It sees in the telemetry that route optimization has barely been touched since a spring outage, and in the support history that Kestrel's new head of operations opened two tickets about overnight sync. It drafts a check-in; the customer success manager sends it. The operations head replies: the sync issue is fixed, thanks, but please do not offer a discount again; last quarter finance took it as a sign the product was worth less. The success manager promises to hold the price and schedule a route-optimization session before the renewal call.

In week four Larkspur deploys a new version of its agent service, and the renewal agent's session ends mid-run. It restarts with the CRM record and nothing else. It re-researches the account, finds the old sync tickets, and drafts an apology for a solved problem. It sees low route-optimization usage and, following its at-risk playbook, proposes a discount.

The success manager catches it, this time. Notice what was lost. Not data: every fact is still in an email thread somewhere. What was lost is the case. The agent had learned three things no system of record held: the discount was off the table, a promise had been made, and the sync issue was closed. None of it was written anywhere the next session would read.

## A long job is many short sessions linked by state

The instinct is to treat this as a context problem: a bigger window, or a summary when the window fills. Both help inside one session. Neither survives a crash, a deploy, a second agent, or a person picking up the case.

Anthropic's engineering team described the same problem for coding agents in terms any operations leader will recognize: "Imagine a software project staffed by engineers working in shifts, where each new engineer arrives with no memory of what happened on the previous shift" [1]. Their fix was not a larger window. It was a progress file read at the start of every session, a JSON feature list with a pass or fail field, and git history as the record of what was done. They say plainly that "compaction isn't sufficient" [1]. Compaction summarizes a conversation near the limit and restarts with the summary [2]. It is a way to keep going, not a way to remember.

The window is also the wrong home for cost reasons. Anthropic treats context as "a finite resource with diminishing marginal returns" [2], and Chroma's tests across 18 models found recall degrading as inputs grow [3]. An agent carrying six weeks of history pays for all of it on every step and reads it worse the longer it gets.

Here is the sentence this chapter turns on. **An agent that runs for weeks does not think for weeks. It thinks for minutes at a time, many times, and the weeks live in the state between.**

## Three memories, three lifetimes

Separate what the agent must remember by how long it must last.

| Memory | What it holds | Lifetime | Where it lives | Who may write |
|---|---|---|---|---|
| Working | The current step's instructions, retrieved context, tool results | One session | Context window | The running agent |
| Task | Goal, plan, completed steps and their outcomes, commitments made, open questions, next action, deadline | The job | A durable case record outside the model | The agents and people assigned to the case |
| Entity | What the organization knows about the customer: properties, evidence, synthesis (Chapter 10) | The relationship | Unified Customer Memory | Governed writers, with provenance |

Most frameworks handle working memory, and Chapters 10 and 11 cover entity memory. Task memory is the one teams skip, because in a demo the task fits in one session. In production it is the difference between an agent and an expensive restart loop.

![A six-week timeline with three rows. Working memory appears as many short sessions, one of which a week-four deploy cuts off. Task memory is a single case record spanning the whole job, so it survives the deploy. Entity memory, the customer memory, extends before and after the job and across every agent; the boundary test is that a fact which still matters after the job closes goes to entity memory with source and date, and otherwise to the case record.](/images/handbook/ch12-three-memories-three-lifetimes.svg)

*Figure 13.1. Working memory lasts a session, task memory lasts the job, entity memory lasts the relationship.*

The boundary matters. "The customer asked us not to offer a discount" is about Kestrel; every agent that touches Kestrel should see it, so it belongs in entity memory with source and date. "Check-in drafted, waiting for the success manager to send it" is about this job, so it belongs in the case record. The test: if the fact would still matter after the job closes, it is entity memory. Write each to the right place and a new session, agent, or person can pick up the case in one read.

## Resume, do not restart

The oldest part of this problem is not about AI. Payment systems and order pipelines have survived crashes for decades by persisting the outcome of each step as it completes, so a failure replays from the last completed step. The name is durable execution.

What changed is the price of redoing a step. A re-run workflow step used to cost milliseconds. A re-run agent step costs a model call, retrieval, perhaps a web search, and the tokens to re-read everything accumulated so far. In a May 2026 interview, Temporal's SVP of Engineering, Preeti Somal, described customers returning to build version 2.0 of agents they had already shipped because they "didn't take care of the plumbing" [4]. Anthropic's multi-agent research system was built the other way: because agents "maintain state across many tool calls," it resumes "from where the agent was when the errors occurred," and the lead agent saves its plan to external memory before its context is truncated [5].

Three properties make resume real:

1. **Checkpoint outcomes, not transcripts.** Record what each step produced (the draft, the score, the decision, the message ID) and the evidence it used. A transcript says what was said; a checkpoint says what is true.
2. **Make external actions idempotent.** If step seven sent an email and crashed before recording it, a replay must not send it again. Give every outbound action a key the downstream system can deduplicate on.
3. **Re-ground on resume.** The checkpoint says where the job was; entity memory says where the customer is now. Read both before acting.

In the memory system I built, a job is a database row, not a process. If its worker dies, the job outlives its timeout, goes back in the queue, and runs again. Running again is the recovery path, which is why property two is not optional.

The personalization engine I built journals finer states (accepted, processing, delivered, artifact pending, recovery attempted, terminal error), so recovery rebuilds the smallest missing unit. Page missing but score delivered: regenerate the page, keep the score. Output right but destination down: resend the exact payload rather than generate a second, different answer. The rule: **retry faults, not decisions.**

A crash that costs one step is an incident. A crash that costs the whole job is a business model you did not choose.

![A job of nine steps where steps one through eight each checkpoint their outcome and evidence into a case record, and the agent crashes on step nine. Restart brings the agent back at step one, so every completed step runs again and the crash is billed twice. Resume brings it back at step nine from the case record, so the crash costs one step. Three properties make resume real: checkpoint outcomes rather than transcripts, make external actions idempotent so a replay never sends twice, and re-ground on resume by reading the case record and entity memory.](/images/handbook/ch12-resume-or-restart.svg)

*Figure 13.2. With checkpointed outcomes, a crash on step nine resumes at step nine instead of step one.*

## Hand off state, not reports

Long jobs involve more than one worker. The renewal agent hands research to a specialist, waits on a person, and eventually passes the case to the account executive. Every handoff is a place where understanding leaks.

The default is the report: agent A writes a summary and passes it to agent B. The summary is lossy by design. It drops the queries tried and abandoned, the weak evidence, the caveat that seemed minor. Chain four reports and the fourth agent reasons from a summary of a summary.

When I rebuilt our own agent API, we replaced one long prompt with an ordered list of steps sharing one accumulated context, so step three sees step one's raw tool results rather than a paraphrase. We got specialist behavior without handoff loss. Across weeks and agents, the same principle has to move outside the model: every worker reads and writes one shared record and passes a reference, not a retelling. Anthropic's research system does the same: subagents "store their work in external systems, then pass lightweight references back to the coordinator" [5].

Shared state changes what the system can see. When a usage analyst records that route optimization stalled after an outage, and a support analyst records that the outage tickets are closed, the renewal strategist can reason over both without either specialist being asked to connect them. The capability lives in the shared record, not in any one agent. The same holds for people: the success manager who opens the Kestrel case reads the same record the agent reads. A handoff that requires someone to write a briefing is a handoff that will eventually be skipped.

![Two panels compare handoffs. In the report chain, agent A passes a summary to B, B to C, and C to D, losing abandoned queries, weak evidence, and minor caveats at each hop, so the fourth agent reasons from a summary of a summary. In the shared-state design, a usage analyst, a support analyst, a renewal strategist, and a success manager all read and write one shared case record with typed outputs and provenance, passing references instead of retellings.](/images/handbook/ch12-state-not-reports.svg)

*Figure 13.3. Reports lose evidence at every hop; a shared record lets every worker, human or machine, reason over the same facts.*

## Fan out for isolation, not for speed

Every major agent platform can now spawn subagents in bulk. The question is when you should.

The cost is documented. Anthropic reports that agents use about 4 times the tokens of a chat interaction and multi-agent systems about 15 times [5]. Each subagent carries its own instructions, tool definitions, and baseline context, and re-reads them every turn. Parallelism feels faster; it is not always faster end to end, and it reliably costs more.

The benefit is real but different. A subagent can read forty documents and return two paragraphs, keeping forty documents out of the main context for the rest of the job. Anthropic describes subagents exploring "tens of thousands of tokens or more" and returning "often 1,000-2,000 tokens" [2]. That is the justification: isolation. A fan-out that exists because two things could happen at once is usually paying twice for one answer.

Isolation also means scope. Each specialist sees only the memory its job needs and holds no more authority than its task requires. The usage analyst does not need contract terms; the drafting agent cannot change who the customer's signer is. Every retrieval is scoped to the entity before any similarity search runs. In an adversarial test I published (100 entities in the same industry with similar roles, overlapping names, and similar deal sizes; 500 queries, 3,800 results), filtering by entity key before vector search produced zero true cross-entity leakage; every flagged result was a false positive from a shared name token [9][10]. Embedding distance alone cannot promise that, because similar customers produce near-identical embeddings.

Authority needs the same scoping when the agent runs on someone else's infrastructure. When my memory system hands a long job to a model provider's background service, the hosted agent never sees the organization's key; it gets a credential for that one job, limited to three memory tools, expiring in a day, revoked when the job ends. What the hosted agent reads is tokenized, and the arguments it sends back are detokenized on the way in, so a real value echoed by the model lands on the right record instead of an orphan keyed by a fake identity. One constraint pulls the other way: those tools work only if the provider's cloud can reach the memory endpoint. A deployment kept on a private network gets a hosted job with no memory access, so the job records a warning rather than running blind without saying so.

For work repeated across thousands of records, the right unit is smaller still: one structured subagent per record, running the same authored instruction chain, grounded in that record's memory, writing typed outputs back. Autonomy is the right default for one open-ended task. It is the wrong default for the same task repeated fifty thousand times.

## Drift is the default; re-grounding is the design

Left alone long enough, agents drift. They misread a situation and commit to it. They treat their own earlier conclusion as fact. They report progress with more confidence than their results earn. Keeping long-running agents on task is most of the engineering work, and anyone who says it is solved is selling something.

Some drift is mechanical: the research agent in Chapter 10 that spent almost a quarter of its output tokens on content it had already seen. The fix is retrieval that remembers what it already delivered within a run and reports when a source is exhausted, so the agent stops asking. The same applies to rules: in my published evaluation, tracking which governance context a session had already received cut re-sent context by about half over five steps [9].

Most drift is semantic, and it is fixed by re-grounding at fixed points rather than hoping the agent notices:

- **At every resume**, re-read entity memory and the case record.
- **Before any consequential action**, re-check recorded commitments and policy. The Kestrel discount dies here.
- **At every handoff**, read the shared record, not the sender's summary.
- **On a schedule**, compare the plan against what changed. A week-one plan can be wrong by week four without any agent making an error.

The context window forgets by accident. A well-designed agent forgets on purpose and remembers on schedule.

## Shared memory multiplies learning, and error

A shared record multiplies in both directions. One unverified write, read by three agents, looks like three confirmations (the failure story below shows how). That is not corroboration. It is one claim echoing through the architecture.

Deliberate corruption of agent memory, usually called memory poisoning, is an active research area. AgentPoison (Chen et al., 2024) showed that a very small fraction of tampered entries in an agent's memory or knowledge base could steer its behavior while leaving normal performance nearly unchanged [6]. For a builder the lesson is integrity controls, and they are the same controls that stop honest mistakes:

- **Separate observation, inference, and authoritative fact.** An agent may propose that someone is the new signer; only a verified source may make it a property.
- **Treat content an agent reads as data, never as instructions** about what to remember or do.
- **Count independent sources, not agreeing agents.** Provenance (Chapter 11) makes echoes visible.
- **Mark derived content** so a generated summary never becomes its own evidence.
- **Write through identity, never raw IDs.** Validating autonomous writes in my own system surfaced a model that put an email address in the record-ID field, creating an orphan, and an update that quietly turned a contact into a generic entity. The fix was code: give the agent each record's identity (email or domain), coerce ID-shaped mistakes, and keep the existing type.
- **Scope write access by role**, stamp each write with a job identity the model cannot see or forge, and keep rollback history for high-impact properties.

Long-running systems also accumulate duplicates, stale entries, and contradictions, which is why Chapter 11's background curation (dreaming) matters here. Anthropic's Managed Agents Dreams, a research preview, reads a memory store plus up to 100 past session transcripts and writes a new, reorganized store, leaving the input untouched for review [7]. That is the right instinct for customer memory: curation should produce a reviewable proposal, not a silent overwrite. One rule from running it: a curation pass's own writes must never wake the next pass, or the loop feeds on itself.

## What this does not do

It does not make agents reliable over long horizons by itself. METR finds the length of task frontier models complete at 50% success has doubled roughly every seven months for six years, and projects week-long software projects within two to four years if the trend holds [8]. Read it carefully. It is measured mostly on software tasks, at a 50% threshold, and a process that must work every night needs far more than one success in two. It also answers a different question. A six-week renewal is not a six-week task. It is dozens of short tasks linked by state, and the reliability that matters is the linkage.

Nor does multi-agent design guarantee better results. Anthropic's widely quoted 90.2% gain for its multi-agent research system was measured on its internal research evaluation, strongest on breadth-first questions, and the same post reports that token usage alone explained 80% of the variance in its browsing evaluation [5]. Much of the gain was spending more tokens. A real result, and a narrow one.

Where things stand, as of September 2026:

| Capability | Status |
|---|---|
| Durable execution, checkpointing, idempotent actions | Ships now; decades-old engineering |
| Case records, progress files, external agent memory | Ships now (vendor tools and in-house builds) |
| Background curation of agent memory with human review | Early production and research previews |
| Agents that manage a multi-week customer relationship with no human checkpoints | Not production practice; forecast at best |

## At scale

At one account, task memory is a file. At 40,000 accounts it is a system of thousands of open cases, most of them waiting on a condition rather than running. The architecture that holds up treats cases as data: a worker is dispatched against a case, checkpoints, and exits; a trigger (a reply, a usage change, a date) wakes the next one. Nothing sits in memory waiting.

Sleeping cases need their own guards. In my systems a schedule fires each occurrence once across the cluster, and an invalid one is disabled rather than retried forever. A job handed to a provider stays a row I own: a sweep reconciles jobs nobody polled, and anything still running after 24 hours is cancelled. A record that fails three passes in a row is quarantined for a week, so one poison record cannot burn tokens nightly. Background work yields queue slots to live requests, so a fan-out cannot starve the customer waiting now.

Three things change with volume. Cost is predictable only if each worker runs a bounded instruction chain in a known token envelope; open-ended autonomy across 40,000 cases is a budget you discover after spending it. Contradictions between agents become statistical certainties, so coordination rules (one owner per customer-facing action, one contact policy) must be enforced by the system, not remembered by each agent. And every error in shared memory reaches every reader, so write gates and provenance matter more than any single agent's accuracy.

## Failure story: Echoed Corroboration

At Larkspur, a support agent logs a caller's claim that he is "the new fleet director"; an enrichment job promotes it to his title; a renewal agent routes the pricing proposal to him, and the actual fleet director hears about it from a colleague. Three systems agreed, and it was one unverified sentence promoted twice. It is Chapter 11's Echo Summary traveling between agents instead of through a summarizer, and the fix is the same: count independent sources, and let no observation become a property without an authoritative one.

## Patterns

**Resume-or-Restart.** *Problem:* long jobs die mid-run and restart from zero. *Forces:* agent steps are expensive to repeat; external actions must never repeat. *Solution:* checkpoint each step's outcome and evidence in a durable case record; key outbound actions for idempotency; re-ground on resume. *Tradeoff:* state placement is decided before step one; retrofitting is a rebuild.

**Shared State over Messages.** *Problem:* reports lose evidence at each hop. *Solution:* all workers, human and machine, read and write one shared record with typed outputs and pass references. *Tradeoff:* shared state needs write permissions and provenance, or it spreads one error to every reader.

**Scope Isolation.** *Problem:* agents see too much, cost too much, and contaminate each other. *Solution:* fan out only to isolate verbose work; give each worker the minimum memory slice and authority; scope retrieval by entity before similarity search. *Tradeoff:* more boundaries to design and test.

## Leader questions

1. If our agent crashes on step nine, where does it come back? Show me.
2. Where does the case live when no agent is running, and can a person read it?
3. Which agents can write to customer memory, and what can each one change?
4. When three agents agree on a fact, how do we know they did not inherit it from one source?
5. Why is this job split across subagents: isolation or speed?

## Build checklist

- [ ] Task memory exists as a durable case record, separate from the context window and from entity memory.
- [ ] Each step's outcome and evidence are checkpointed on completion.
- [ ] Every external action carries an idempotency key.
- [ ] Resumed and handed-off workers re-read entity memory before acting.
- [ ] Commitments made to a customer are written to entity memory with source and date.
- [ ] Agents pass references to shared state, not free-text reports.
- [ ] Retrieval is scoped by entity before similarity search; a leakage test exists.
- [ ] Write permissions are scoped by role; observations cannot promote themselves to properties.
- [ ] Every agent has a turn ceiling and a context budget; every job has a timeout and a maximum age.
- [ ] Faults are retried; completed outputs are resent, not regenerated.
- [ ] The kill test (stop a running agent in staging) runs before every release.

## Metrics to watch

- **Resume point after failure:** share of interrupted jobs that resume at the last completed step.
- **Re-execution cost:** tokens spent redoing steps already completed, as a share of total.
- **Commitment violations:** actions that contradict a recorded customer commitment (target: zero).
- **Single-source agreement rate:** facts cited by several agents that trace to one source.
- **Cost per completed case**, not per call.

## Reader Q&A

**Is a bigger context window a substitute for task memory?** No. It delays the problem inside one session and makes each step more expensive. It does nothing for crashes, deploys, or handoffs.

**Should the transcript be the memory?** Keep transcripts for audit and background curation, not as the first thing the next session reads. A case record tells the next worker what is true; a transcript makes it work that out again.

**How many subagents is too many?** There is no universal number. If you cannot state a subagent's boundary (what it reads, what it writes, what it isolates) in one sentence, it probably should not exist.

**Where should a human sit in a six-week job?** At the consequential moments: before irreversible or customer-facing actions, and wherever evidence is weak. Chapter 18 covers making that review usable.

## For your AI

```yaml
chapter: 13
concepts:
  - name: Working memory
    definition: "The context window: current instructions, retrieved context, and tool results; lasts one session."
  - name: Task memory
    definition: "A durable case record for one job: goal, plan, completed steps and outcomes, commitments, open questions, next action; lasts as long as the job."
  - name: Entity memory
    definition: "What the organization knows about the customer, with provenance; outlives every agent (see Chapter 10)."
  - name: Durable execution
    definition: "Persisting each step's outcome so a failure resumes from the last completed step instead of restarting."
  - name: Re-grounding
    definition: "Re-reading entity memory, the case record, and policy at resume, handoff, and before consequential actions."
  - name: Echoed corroboration
    definition: "Several agents appear to confirm a fact because they inherited one unverified write; the multi-agent form of Chapter 11's Echo Summary."
decision_rules:
  - if: "a fact will still matter after the current job closes"
    then: "write it to entity memory with source and date; otherwise write it to the case record"
  - if: "an agent resumes after a crash, deploy, or handoff"
    then: "load the case record, re-read entity memory, then act; never rebuild state from the transcript"
  - if: "a step performs an external action"
    then: "attach an idempotency key so a replay cannot repeat it"
  - if: "a subagent is proposed only to finish sooner"
    then: "reject it unless it also isolates verbose work from the main context"
  - if: "several agents cite the same fact"
    then: "count independent sources via provenance, not the number of agents"
  - if: "an agent's observation would change a high-impact property"
    then: "require an authoritative source or human review before promotion"
  - if: "a step fails after its output or decision already exists"
    then: "retry the fault (resend the completed output); do not regenerate the decision"
  - if: "a job runs on infrastructure you do not control"
    then: "issue a job-scoped, expiring, least-tool credential and revoke it when the job ends"
assessment_questions:
  - "Where does an agent's in-progress job state live when no agent is running?"
  - "What happens, concretely, when a running agent is killed mid-job?"
  - "Which agents can write to customer memory, and what can each change?"
  - "Are commitments made to customers recorded somewhere every agent reads?"
  - "Is retrieval scoped by entity before similarity search, and is there a leakage test?"
patterns: [Resume-or-Restart, Shared State over Messages, Scope Isolation]
anti_patterns: [The Amnesiac Restart, Echoed Corroboration, Fan-out for Speed, Report Chains]
maturity_dimension: identity_and_memory
```

## References

1. Anthropic Engineering. "Effective harnesses for long-running agents." 2025-11-26. https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
2. Anthropic Engineering. "Effective context engineering for AI agents." 2025-09-29. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
3. Hong, K., Troynikov, A., Huber, J. "Context Rot: How Increasing Input Tokens Impacts LLM Performance." Chroma, 2025-07-14. https://www.trychroma.com/research/context-rot
4. VentureBeat. "AI agents are entering their rebuild era as enterprises confront the reliability problem." May 2026 (interview with Preeti Somal, Temporal). https://venturebeat.com/orchestration/ai-agents-are-entering-their-rebuild-era-as-enterprises-confront-the-reliability-problem
5. Anthropic Engineering. "How we built our multi-agent research system." 2025-06-13. https://www.anthropic.com/engineering/multi-agent-research-system
6. Chen, Z., Xiang, Z., Xiao, C., Song, D., Li, B. "AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases." arXiv:2407.12784, 2024-07-17. https://arxiv.org/abs/2407.12784
7. Anthropic. Claude Managed Agents, "Dreams" (research preview), documentation, accessed 2026-09-26. https://platform.claude.com/docs/en/managed-agents/dreams
8. Kwa, T., et al. (METR). "Measuring AI Ability to Complete Long Tasks." 2025-03-19; arXiv:2503.14499. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
9. Taheri, H. "Governed Memory: A Production Architecture for Multi-Agent Workflows." arXiv:2603.17787, 2026-03-18. https://arxiv.org/abs/2603.17787
10. Taheri, H. "Zero Cross-Entity Leakage Across 3,800 Results." hamedtaheri.com, 2026-03-14. Source of the 100-entity and 3,800-result detail. https://hamedtaheri.com/articles/zero-cross-entity-leakage
