# 17. Now or Later

*The question: When must personalization be real-time, and when is async better?*

> **Questions this chapter answers**
> - Which of our personalization actually has to happen in real time?
> - Is it worth making the customer wait for a smarter answer?
> - How do we get deep, researched personalization onto a page that loads in under a second?
> - What breaks when we move work into queues and background jobs, and how do we stop it?
> - What should the customer see when the smart path fails?

## The short answer

Most personalization does not need to be computed in the moment. It needs to be *served* in the moment. Those are different engineering problems.

Speed matters where a person is waiting: controlled experiments at Google, Bing, and Booking.com measured real losses from added delay on pages people were using. But the deep work (researching an account, deciding what matters to this person) can almost always be done earlier, stored, and read back in milliseconds.

So the design rule is: precompute depth, add live context at the last moment. Keep the real-time path to a lookup plus a little fresh signal. Both major model providers now price patience: their batch interfaces charge half the synchronous rate.

Async is not free. Queues trade slowness for duplication, retry storms, and stale results. The cures are boring: idempotency keys, bounded retries with jitter, freshness windows, and a fallback ladder that makes output less specific when something fails, never less true.

What to fund: a latency budget per channel, a precompute layer keyed to the account or person rather than the request, and a degradation ladder agreed in writing.

---

Tuesday, 10:14 a.m. An operations director at a regional plumbing and HVAC contractor, a customer of Larkspur Systems (a fictional composite: field-service and fleet software, about 40,000 accounts), clicks a renewal email and lands on the pricing page. Larkspur knows her company opened a second depot this spring and added about twenty vans, that her team opened four tickets about route exports, and that her contract renews in 104 days. The page has well under a second to reflect any of that. If it stalls to think, she scrolls past a blank block, or watches it jump when the personalized version finally paints.

In the same minute, in a background worker, an agent is writing the renewal plan for that account. It reads eighteen months of tickets, call notes, telemetry, and news about her company. It will take most of an hour. It should. Nobody is waiting, and every minute of reading makes the plan better.

Same customer, same company, same morning. One piece of personalization is wrong if it is slow. The other is wrong if it is fast.

## Latency is a budget, and each channel sets its own

Start with what is measured. In 2009, Jake Brutlag at Google injected 100 to 400 ms of server delay into search results and saw daily searches per user fall 0.2% to 0.6%; after a 400 ms delay was removed, affected users stayed 0.21% below control for five weeks [1]. Ronny Kohavi's team at Bing isolated speed and estimated every 100 ms of speedup at about 0.6% of revenue [2]. At Booking.com, about 30% more latency cost more than 0.5% in conversion, which pushed the team to precompute and cache predictions where it could [3].

Small per interaction, large in aggregate. One qualification matters more than the headline: slowing Bing's right-hand pane by 250 ms produced no significant change in key metrics [2]. Delay hurts where the person is waiting. Delay nobody is waiting for is nearly free.

A **latency budget** is the time a channel allows before the delay itself changes the outcome. The channel sets it, not your architecture.

| Channel or moment | Working budget (design guidance) | What runs inside it |
|---|---|---|
| Web page, first paint (waiting) | A few hundred ms of server time; "good" Largest Contentful Paint is 2.5 s at p75 [5] | Lookup of precomputed content plus live context |
| In-product interaction (waiting) | About 0.1 s to feel instant; 1 s before flow breaks [4] | Lookup, small model score, rules |
| Chat or agent reply (waiting) | First visible output within about 1 s; the rest streams | Memory retrieval plus streamed generation |
| Triggered email (not waiting) | Minutes | Retrieval, generation, validation |
| Rep brief, account plan (not waiting) | Hours to a day | Multi-step agentic research |
| Direct mail (not waiting) | Days | Everything, plus human review |

The first rows are anchored in human-factors limits (Miller 1968, restated by Nielsen: 0.1 s feels instantaneous, 1 s keeps the flow of thought, 10 s holds attention) [4] and Google's page-experience thresholds [5]. The rest are my working targets, not standards. Write yours down before anyone argues about architecture.

![Six channels grouped by whether a person is waiting. Where someone waits, the budgets are tight: a web page's first paint gets a few hundred ms of server time for a lookup plus live context; an in-product interaction has about 0.1 s to feel instant and 1 s before flow breaks, enough for a lookup, a small score, and rules; a chat or agent reply shows first output in about 1 s and streams the rest. Where nobody waits, budgets widen to minutes for a triggered email, hours to a day for a rep brief or account plan, and days for direct mail, which is enough for retrieval, generation, validation, multi-step agentic research, and human review. The lower rows are working targets, not standards.](/images/handbook/ch15-latency-budget-by-channel.svg)

*Figure 17.1. Where a person is waiting, only lookup and selection fit; everything deeper belongs where nobody is.*

Watch layout shift too; Google's "good" threshold is 0.1 or less [5]. Late personalization that replaces a default block makes the page jump. Late personalization is not just slow. It is visibly unstable. Generated experiences stretch the budget furthest: a generated interface can take a minute or more, so Chapter 16 generates pages ahead of the visit and leaves the live path only to select.

## Do the work before they arrive

If the page has a few hundred milliseconds and account research takes minutes, the page need not be shallow. The research must already be finished.

Netflix described this split in 2013 as three tiers: **offline**, with few limits on data or complexity but prone to staleness; **online**, fresh but bound by a tight service level; and **nearline**, computed on events but stored for later serving. The tiers are not either/or: precompute part of a result offline and leave the cheaper, more context-sensitive parts online [6].

Language models make that split more valuable, because the expensive step got more expensive and more useful at once. The sleep-time compute research makes the economics explicit: processing a known context before the question arrives cut test-time compute about 5x at equal accuracy in its benchmarks, and the benefit grew with how predictable the queries were [7]. Personalization is unusually predictable. You know which accounts renew next quarter, and which facts a pricing-page visit will need.

I call the pattern **Precompute-then-serve**. The expensive understanding runs asynchronously and is written to memory as a stored artifact: an account synthesis, a ranked list of what matters now, pre-generated content zones. The real-time path reads it, adds what could only be known in the moment, and renders.

![Two lanes. In the async lane, where nobody is waiting, an event or a schedule triggers deep research keyed to the entity rather than the request, and the result is written as a stored artifact (account synthesis, ranked facts, content zones) with provenance and a freshness window. In the real-time lane, where a person is waiting a few hundred milliseconds, the system looks up that artifact (a fetch, not a generation), combines it with live context such as the pricing page, the email click, and the phone, selects among precomputed options, and renders. Only the part that depends on the last few seconds runs live.](/images/handbook/ch15-precompute-then-serve.svg)

*Figure 17.2. Precompute-then-serve: the slow understanding happens before the visit, so the visit is a lookup plus a selection.*

In the governed personalization engine I built, the inbound request validates the lead, persists workflow state, claims an idempotency key, dispatches the work, and returns 202 Accepted. Research, scoring, and the personalized landing page are produced in the background and the page link goes back with the callback, so the click is a fetch, not a generation. A production burst taught me not to hold that request open [16]. And research is keyed to the entity, not the request. When the unit was the individual lead, one popular employer was researched close to a hundred times in a day. Keyed to the company with a seven-day freshness window, the first record paid and the rest read, and records classified as research-rich rose from 60.3% to 93.5% as budget moved from known answers to open gaps [8]. The fastest personalization is the one you finished before they arrived.

## Event, schedule, or request

Async work still needs a reason to start. There are three.

- **Event-triggered**: something happened (a cancellation click, a fourth ticket in a week, a champion changing jobs) that changes what should be said. Seconds to minutes.
- **Scheduled**: the calendar is the trigger (renewals in 90 days, a weekly account refresh, the curation pass Chapter 11 calls dreaming). For predictable value and slow-changing inputs.
- **On-request**: someone asks, and the work runs inside their wait. Reserve it for what genuinely depends on the request.

Most teams default to on-request because it is easiest to build, which puts the most work inside the tightest budget.

## Depth is decided early; relevance is decided now

Here is the hinge. The question is never "real-time or batch?" It is: **which part of this decision depends on something that happened in the last few seconds?** That part runs live. Everything else runs earlier.

For the operations director, the new depot, the tickets, and the renewal date were all known yesterday. What is new is that she is on the pricing page, from the renewal email, on a phone. That live context selects among precomputed options: the renewal variant, the mobile layout, the route-export zone. Selection is fast. Understanding was slow, and it already happened.

Real-time is where you use what you know. Async is where you come to know it.

## Async is where the failures move

Moving work out of the request path changes the shape of failure. A slow request fails in front of the customer; a queued job fails silently, repeatedly, or twice.

Many model-call failures are not wrong answers at all. In Datadog's analysis of customer LLM traffic, 2% of model call spans errored in March 2026 and rate limits were almost a third of those; in February, 5% errored and 60% of errors were rate limits [9]. Capacity, not intelligence, is the first thing to design for. When a fleet of workers is refused in the same second and retries on the same interval, the retries land as a synchronized wave on a provider that is still saturated.

The disciplines are dull and non-negotiable:

1. **A concurrency cap** per provider, so parallelism is a number you chose, not one your traffic discovered.
2. **Bounded retries with exponential backoff and jitter**, at one layer of the stack, not several [10].
3. **Retry faults, not decisions.** Retry a timeout, a rate limit, a provider 5xx. Do not retry missing context, an invalid contract, or a policy rejection; that pays for the same "no" twice. Know where it failed, too: in the self-hosted memory system I built, a 429 comes before any write and is always safe to retry after its Retry-After interval; a bare 500 mid-save may already have written, so it is investigated first [17].
4. **Idempotency keys** on every mutating step (intake, generation, send, callback) [11], derived from the work (lead, campaign, step) rather than the request, so a redelivery carries the same key. Queues redeliver and networks fail after the server finished; without a key, each becomes a second generation and possibly a second, different message to the same person.
5. **Recover the smallest failed unit.** If the output is correct but the callback failed, resend the exact payload; regenerating costs money and produces a different message. If only the page is missing, regenerate the page alone. Both need a durable journal of each step's state (accepted, callback delivered, page pending, terminal error), not logs [16].

Refuse at your own front door too. The memory system rejects work it cannot run, on separate budgets so a write overload does not stop reads, rather than accepting everything and silently falling behind [17]. Only bulk work lands in its durable queue; a refused single save is not queued on the caller's behalf, so the caller retries or moves the load to the bulk path.

A queue turns a busy minute into a slow minute. Without idempotency, it turns a slow minute into a duplicate.

## Degrade specificity, never truthfulness

Netflix's 2013 advice for the online tier was to always have a fast fallback, such as a precomputed result [6]. AI personalization needs more than one cached answer, because some of its failures produce plausible output.

The name is deliberately narrow: **Graceful Degradation** is only for dependency and system failures (a throttled provider, a timeout, a missing upstream artifact). It is not Honestly General, the content principle that decides what to say when the evidence is thin (stepped down through Chapter 15's Specificity Ladder); this ladder decides what to serve when the machinery fails. The ladder I use follows the decision modes of Chapter 14:

1. **Agentic**: live reasoning over memory and fresh context. Most specific, slowest, most expensive.
2. **Predictive**: a precomputed or model-scored choice among approved variants. Fast, still personal.
3. **Rule-based**: segment or lifecycle-stage content selected by deterministic rules.
4. **Approved static**: the known-safe default for that zone.
5. **Fail closed**: publish nothing, when even the default would break the contract or the experience.

A missed deadline counts as a system failure. When rich enrichment fails close to its execution deadline, the engine I built emits a deterministic, lower-specificity brief rather than inventing detail or leaving the record stuck [16].

Each zone declares its own ladder, and every output is labeled full, degraded, or failed so it can be counted. A system that falls back silently looks healthy while serving generic content to a large share of its audience. A specific but false message is worse than a general but true one, so the ladder moves toward less specific, never toward less verified.

![A five-rung staircase triggered only by a dependency or system failure (throttled provider, timeout, missing upstream artifact). From most to least specific: agentic, live reasoning over memory and fresh context, labeled full; predictive, a precomputed or scored choice among approved variants; rule-based, segment or lifecycle content chosen by deterministic rules; approved static, the known-safe default for the zone, these three labeled degraded; and fail closed, publishing nothing when even the default would break, labeled failed. Verification is held on every rung. Each zone declares its own ladder, and every output is labeled so fallbacks can be counted.](/images/handbook/ch15-degradation-ladder.svg)

*Figure 17.3. Graceful Degradation steps down in specificity, never in truthfulness, and labels every step so silent fallback shows up in the numbers.*

## What this does not do

It does not make latency statistics portable. Kohavi's own rules of thumb include "your mileage will vary" [2]. Search engines and a travel marketplace do not tell you what 300 ms costs on your B2B pricing page. Run your own slowdown test first.

Two popular numbers should not be used at all. "Every 100 ms cost Amazon 1% of sales" is usually traced to Greg Linden's 2006 slides, with no published method. "Google lost 20% from a 500 ms delay" came from a test that also raised the results per page from 10 to 30; when Bing isolated speed, 500 ms cost about 3% of revenue, "not 20%" [2].

It does not make real-time better whenever you can afford it. Rewriting a page on a single click is often worse than a stable page informed by the last month; the adaptive-menu research in Chapter 3 shows why. Fresh does not mean important.

And a disagreement with a respected source. Jay Kreps argued against the Lambda architecture's separate batch and streaming code paths [12]: use one streaming path and reprocess by replay [13]. I agree with the first half: one code path and one generation contract for both routes. I disagree that replay is cheap here. Replaying a language-model pipeline costs money per record and does not reproduce the same outputs. Store generated artifacts with provenance, and recover missing pieces instead of replaying history.

## At scale

At 250,000 contacts, precompute has its own waste. A page for every contact, when a small fraction will ever visit, is content nobody reads. Precompute where reuse is high (the account synthesis, the ranked facts) and generate the person-specific layer on the event that signals interest.

Staleness scales too. Every artifact needs a freshness window and invalidation events: a job change, a cancellation, a closed ticket. A cache that only accumulates serves a confident wrong answer until it expires.

Background work must not starve the foreground. In the memory system, overnight curation claims queue slots only after waiting live work and is admitted only up to the ceiling minus a reserve (20 slots by default), so it backs off instead of crowding out a live request [17]. The bulk worker yields the same way but never stops: under a sustained burst it still takes a small slice, so an import slows rather than stalls.

Admission limits have their own scale traps, all from sizing that system [17]. **Limits are per process**: each replica counts only its own traffic, so three replicas admit roughly three times every number, and a team that forgets the multiplier sizes wrong. **Tune before you enforce**: a dry-run mode admits everything but records what would have been refused, so ceilings come from real traffic, not guesses. Raising admission without database connection headroom only moves the wait from the front door to the database. And a per-tenant **backlog ceiling** on the bulk queue caps what one runaway producer can claim, but it governs new intake, not work already accepted: when a provider's batch fails and its items are requeued, they go back even past the ceiling, because the caller was already told they were queued.

And cost follows timing. Anthropic's Message Batches API and OpenAI's Batch API each list a 50% discount for work that can wait up to 24 hours (as of 2026-09-26) [14][15]. The async decision is also a pricing decision. Make it per job, with a way back: a batch the provider fails is requeued to the normal worker, so nothing is lost [17].

## Failure story: The Double Send

A composite at Larkspur, built from failure modes I have debugged. At 9:00 a.m. a renewal campaign enqueues 40,000 accounts at once. The provider starts refusing calls; every worker retries on the same fixed interval, keeping it saturated. Some jobs time out after generation but before the send recorded success. The queue redelivers them. The send step has no idempotency key, so a few hundred customers receive two renewal emails, each personalized, each different, one citing a discount the other does not. Nothing in the model was wrong. The orchestration decided, in unison, to try again.

## Patterns

**Precompute-then-serve.** *Problem:* deep personalization does not fit the real-time budget. *Forces:* research is slow; the moment is short; stored results go stale. *Solution:* compute understanding asynchronously per entity, store it with provenance and a freshness window, serve by lookup plus live context. *Tradeoffs:* precompute wasted on people who never arrive; invalidation logic to own.

**Latency Budget.** *Problem:* teams argue architecture without agreeing what "fast" means. *Solution:* a written budget per channel, at a percentile, listing what may run inside it. *Tradeoffs:* needs re-measuring as channels change.

**Graceful Degradation.** *Problem:* the smart path fails in many ways, some plausible. *Solution:* a declared ladder per zone, every output labeled. *Tradeoffs:* more variants to approve and maintain.

## Leader questions

1. Which personalized experiences have a person waiting, and what is each written budget?
2. How much of our real-time path is computation that could have happened yesterday?
3. When the model provider throttles us, what does the customer see?
4. Can any step send the same person two different messages?
5. Under a burst, do we refuse early and say so, or fall behind silently?

## Build checklist

- [ ] Latency budget per channel, measured at p75 or p95.
- [ ] Precomputed artifacts keyed to the entity, with freshness windows and invalidation events.
- [ ] Real-time path limited to lookup, scoring, and selection among approved variants.
- [ ] Concurrency caps; bounded retries with backoff and jitter at one layer; faults retried, decisions not.
- [ ] Idempotency keys on intake, generation, send, and callback.
- [ ] Intake that returns 202 after validate, persist, claim, dispatch; a durable journal so recovery targets the missing step.
- [ ] Admission limits that refuse (429 with Retry-After) before any write; background work yields to live traffic; limits sized per replica and tuned in dry run first.
- [ ] Degradation ladder per zone; every output labeled full, degraded, or failed.
- [ ] Batch interfaces for all work nobody is waiting for.

## Metrics to watch

- **p75 and p95 latency of the personalized path**, against the channel budget.
- **Degraded-output rate** by zone and by cause.
- **Precompute hit rate and staleness**: share of real-time requests served from a fresh artifact.
- **Duplicate-action rate**: sends or writes that share an idempotency key.
- **Stuck and backlog counts** (claims, pending pages, recovery), with any new error signature treated as a warning [16].

## Reader Q&A

**Can we just use a faster model on the live path?** For selection and short generation, sometimes. No model does an hour of research in 200 ms; move the research.

**Is a streamed chat reply real-time personalization?** It is real-time delivery. The personal part should come from memory curated before the conversation started.

**How fresh is fresh enough?** Per fact. A job title can be a week old; a cancellation click cannot be five minutes old. Freshness is set per property (Chapter 11).

## For your AI

```yaml
chapter: 17
concepts:
  - name: Latency Budget
    definition: "The time a channel allows before delay itself changes the outcome; set per channel and measured at a percentile."
  - name: Precompute-then-serve
    definition: "Compute understanding asynchronously at the entity level, store it with provenance and freshness, serve by lookup plus live context."
  - name: Graceful Degradation
    definition: "A declared fallback ladder per zone (agentic, predictive, rule-based, approved static, fail closed) that reduces specificity, never truthfulness. Used only for dependency and system failures; thin evidence is handled by Honestly General and the Specificity Ladder (Chapter 15)."
  - name: Retry faults, not decisions
    definition: "Retry transient infrastructure failures; never retry deterministic rejections."
  - name: Admission control
    definition: "Refuse work you cannot run (429 with Retry-After, before any write) instead of falling behind; background work yields to live traffic."
decision_rules:
  - if: "an inbound request triggers research or generation"
    then: "validate, persist workflow state, claim an idempotency key, dispatch async work, and return 202"
  - if: "the output is complete but delivery to the destination failed"
    then: "retry the exact completed payload; do not regenerate"
  - if: "a person is waiting and the work takes longer than the channel budget"
    then: "move the work to an event-triggered or scheduled job and serve the stored result"
  - if: "part of the decision depends on signals from the last few seconds"
    then: "run only that part live, as selection among precomputed options"
  - if: "a step mutates state or sends to a person"
    then: "require an idempotency key before it runs"
  - if: "a call fails on rate limit, timeout, or provider 5xx"
    then: "retry with bounded exponential backoff and jitter at one layer"
  - if: "a call fails on missing context, invalid contract, or policy rejection"
    then: "do not retry; take the next rung of the degradation ladder"
  - if: "no one is waiting for the result"
    then: "use a batch interface"
  - if: "work already accepted is requeued after a provider failure and the backlog is over its ceiling"
    then: "re-admit it anyway and log a warning; the ceiling governs new intake, not work already promised"
assessment_questions:
  - "Which personalized experiences have a person waiting, and what is each latency budget?"
  - "What is computed on request that could be precomputed on an event or schedule?"
  - "Is research keyed to the request or to the entity?"
  - "What does the customer see when the model provider throttles you?"
  - "Which steps can send or write twice?"
patterns: [Precompute-then-serve, Latency Budget, Graceful Degradation]
anti_patterns: [The Double Send, Synchronized Retry, Silent Fallback, Everything On-Request]
maturity_dimension: governance_and_reliability
```

## References

1. Brutlag, J. (2009-06-22). "Speed Matters for Google Web Search." Google. https://services.google.com/fh/files/blogs/google_delayexp.pdf
2. Kohavi, R. et al. (2014-08-27). "Seven Rules of Thumb for Web Site Experimenters" (KDD 2014 slides). https://exp-platform.com/Documents/2014-08-27ExperimentersRulesOfthumbKDD.pdf ; see also Kohavi et al. (2013), "Online Controlled Experiments at Large Scale," KDD 2013. https://dl.acm.org/doi/10.1145/2487575.2488217
3. Bernardi, L., Mavridis, T., Estevez, P. (2019). "150 Successful Machine Learning Models: 6 Lessons Learned at Booking.com." KDD 2019. https://dl.acm.org/doi/10.1145/3292500.3330744
4. Nielsen, J. (1993, updated 2014). "Response Times: The 3 Important Limits." Nielsen Norman Group, drawing on Miller (1968). https://www.nngroup.com/articles/response-times-3-important-limits/
5. Google, "Web Vitals" (last updated 2024-10-31; accessed 2026-09-26). https://web.dev/articles/vitals
6. Amatriain, X., Basilico, J. (2013-03-27). "System Architectures for Personalization and Recommendation." Netflix Tech Blog. https://netflixtechblog.com/system-architectures-for-personalization-and-recommendation-e081aa94b5d8
7. Lin, K. et al. (2025). "Sleep-time Compute: Beyond Inference Scaling at Test-time." arXiv:2504.13171. https://arxiv.org/abs/2504.13171
8. First-party measurement on a production enrichment pipeline operated by the author, 2026-08 (entity-level research cache).
9. Datadog (2026-07-30). "State of AI Engineering." https://www.datadoghq.com/state-of-ai-engineering/
10. AWS Well-Architected Framework, "Control and limit retry calls." https://docs.aws.amazon.com/wellarchitected/2025-02-25/framework/rel_mitigate_interaction_failure_limit_retries.html
11. AWS Well-Architected Framework, "Make mutating operations idempotent." https://docs.aws.amazon.com/wellarchitected/latest/framework/rel_prevent_interaction_failure_idempotent.html
12. Marz, N. (2011). "How to beat the CAP theorem." https://nathanmarz.com/blog/how-to-beat-the-cap-theorem.html
13. Kreps, J. (2014-07-02). "Questioning the Lambda Architecture." O'Reilly Radar. https://www.oreilly.com/radar/questioning-the-lambda-architecture/
14. Anthropic, "Batch processing" (accessed 2026-09-26). https://platform.claude.com/docs/en/build-with-claude/batch-processing
15. OpenAI, "Batch API" guide (accessed 2026-09-26). https://developers.openai.com/api/docs/guides/batch
16. First-party: capabilities reference for a governed personalization engine built by the author (reliability, delivery, and recovery sections), 2026-08-25. Unpublished.
17. First-party: operator documentation for a self-hosted governed memory system built by the author (capacity and limits, features, changelog), release 0.8.2, verified 2026-09-16.
