# 22. Deploying It

*The question: How does a company actually build and roll this out?*

> **Questions this chapter answers**
> - What do we build first, and in what order?
> - Should we build, buy, or own the personalization layer?
> - What team does this take, and who owns what?
> - What does the first 90 days look like, and how do we know we are ready to widen?
> - How do we tell where we are, and what to improve next, without chasing the newest tool?

## The short answer

A personalization engine is not a model with a prompt. It is a production system in which the model is one function among many: intake, identity, governed context, planning, generation, validation, channel compilation, delivery, recovery, quality review, and governance. Most of the deployment effort goes into the parts around the model, and most first-week failures happen there too.

So deploy it in the order the failures arrive. Write the contract before touching data: what the experience must do, which claims are allowed, what a fallback looks like, which rules code enforces. Build the shared mechanics once (identity, trust rules, idempotency, choke-point protections, workflow state) so every later campaign inherits them. Run a real but small first cohort as part of the build, not after it. Widen only when the release thresholds you wrote down in advance say so, and let automation propose improvements while people keep the right to ship them.

On build versus buy: rent what is commodity, own what compounds. The memory, the routing decisions, and the reviewed examples get more valuable with use, and whoever holds that layer holds the economics.

For a leader, three questions decide most of the outcome. Is there a written contract for the first decision? Can the team say, for any record, what state it is in and what was delivered? Will the second campaign cost clearly less to launch than the first? If not, you have built a pilot, not a platform.

The phases, in order: contract, foundations, shadow and first cohort, supervised expansion, a second campaign on the same mechanics. Only then does turning the engine into a product begin.

## What broke first

*Larkspur Systems is a fictional composite company used throughout this handbook. The failures below are composites of failure classes documented in my published engineering write-up [1]; the counts are illustrative.*

Larkspur switched on its engine for renewal season on a Monday. After the inbound fast win (Chapter 6), this was the engine's first renewal campaign: an email, a landing page, and a brief for the account manager, for each fleet-operations contact whose contract renewed within 120 days. Chapter 20 tells what the content review found. This is what the operations dashboard found, and it found it first.

By Monday afternoon, contacts who had signed up with personal email addresses were receiving pages that described "operations at your scale" using the headcount of their email provider. The domain-to-company step had trusted a free-mail domain as an employer.

By Tuesday, one research provider had exhausted its credits. Nothing stopped it, so every record sent it the same failing request, waited for the timeout, and moved on with thinner context. Quality dropped all day without an error anyone was watching.

By Wednesday, the CRM showed every record as "complete." A reconciliation script found that a few dozen landing pages had been generated, validated, and stored, but never received a public URL. The callback had reported success because the callback succeeded. The artifact the customer needed did not exist.

None of these was a model failure. The model wrote well all week. The failures lived in identity, dependency handling, and delivery state: the parts of a deployment that demos never exercise. Each has a standing guardrail in the engine I built: free-mail domains suppress firmographics, an exhausted provider trips a circuit breaker, and a missing page is regenerated alone, with the score and research untouched.

## The model is the smallest part of the deployment

In 2015 a group of Google engineers argued that machine learning's quick wins are not free: real-world ML systems accumulate "massive ongoing maintenance costs" through entanglement, hidden feedback loops, and data dependencies [2]. The model is a small box inside a large system. Generative personalization makes the box more capable and the system around it more important, because the box now writes what customers read.

Here is the architecture I use, top to bottom [1]: the **campaign contract**; **input validation and tenant routing**; **identity resolution**; **research and first-party context**; **data trust and provenance**; the **governed customer context** (facts with source, confidence, recency, and allowed surfaces); a **content plan** that splits the experience into typed zones and picks the personalization tier; **generation**, **self-review**, and **deterministic validation**; **channel compilation** for email, web, seller playbook, paid, print, and events; **delivery state**, **retry**, and **targeted recovery**; and finally **population QA** and **human-supervised optimization**, which feed the next version.

![The reference architecture drawn as one flow of twelve steps. It runs from the campaign contract through input validation and tenant routing, identity resolution, research and first-party context, data trust and provenance, the governed customer context, and a content plan of typed zones and tier. Only then comes the model step, generation and self-review, highlighted as one box among many. After it come deterministic validation, channel compilation, delivery state with retry and recovery, and population QA with supervised optimization, which loops back to feed the next version of the contract.](/images/handbook/ch20-reference-architecture.svg)

*Figure 22.1. The model is one step in the middle; everything before and after it decides whether its output may reach a person.*

Each layer exists because a failure there is different from a failure anywhere else. A wrong company match is not a writing problem. A missing page URL is not a research problem. If the system cannot tell these classes apart, every defect becomes another sentence in the prompt, and that does not scale.

In code, the model is one call in the middle. Identity is resolved first, trust policy filters context before the model sees it, the tier is a policy output, deterministic validation has the final word on hard rules, and fallback is an explicit state [1]. Chapters 9, 10, 15, and 20 build those pieces. This chapter puts them into production in the right order.

Here is the sentence this chapter turns on: **the model is a function call; the deployment is everything that decides whether its output is allowed to reach a person.**

## Start with the contract, not the customer record

Before touching data, I write the contract. It turns "personalize this" into something testable: who the audience is, which surfaces are active, what level of personalization is expected, which capabilities and offers are approved, which claims are prohibited, which facts need corroboration, what a valid fallback looks like, and which rules are advisory versus enforced in code [1].

Where does the contract come from? In the engine I built, it starts as a customer intake: the ICP, positioning, the economic story, the claims allowed and their sources, banned language, examples, legal copy, the CTA, and brand behavior. The intake becomes an approved guideline, the guideline a content-zone schema, the schema a template and channel configuration, and only then does anything get generated. That chain runs today as an operating process; turning it into a standardized, repeatable workflow is still in progress.

The contract is the only thing QA can judge against. A company-level page that never names an individual is correct for one campaign and a defect for another. Without a contract, the release decision falls back on a vague sense of good writing, and vague senses do not produce GO or NO GO.

Make it machine-readable where possible (surfaces, tier, prohibited claim classes, length limits, risk tier, fallback policy) and let a prose guideline carry the rest. Put the release thresholds in it too, written before the first record is generated, so the team cannot talk itself into a launch.

## Shared mechanics are platform; campaign judgment is configuration

The decision that most determines what the second campaign costs is where each requirement lives [1].

**Shared mechanics** behave the same everywhere: sanitizing malformed output, escaping HTML, normalizing identifiers, idempotency, classifying retryable failures, hard length limits, stripping internal metadata, blocking unsafe URL schemes, preventing cross-tenant routing, enforcing wire schemas. They live in code, at shared choke points, and get stronger over time.

**Campaign judgment** belongs in versioned guidelines and configuration: voice, approved capabilities, persona priorities, CTA strategy, recency windows, examples, proof hierarchy.

The reason is blast radius. A campaign wanting a more technical tone should not need a global engine deploy; a rule every campaign needs should not depend on twenty prompt authors remembering. For every requirement, ask: *what kind of control is this?* Models reason, guidelines express judgment, code enforces invariants.

![Three campaigns sit on top of one shared platform. Each campaign carries its own judgment in versioned guidelines and configuration: voice, approved capabilities, persona priorities, CTA strategy, recency windows, examples, and proof hierarchy, changed per campaign without a global engine deploy. Every campaign inherits the same shared mechanics, enforced in code at shared choke points: sanitizing output, escaping HTML, normalizing identifiers, idempotency, retry classification, hard length limits, stripping internal fields, blocking unsafe URLs, tenant routing, and wire schemas. Models reason, guidelines express judgment, code enforces invariants.](/images/handbook/ch20-shared-mechanics-campaign-judgment.svg)

*Figure 22.2. Put campaign intent in configuration and invariants in shared code, so each new campaign inherits the protections instead of forking the engine.*

The cheapest reliability investment in the deployment: find the last shared function before each important boundary (renderer, customer serializer, URL parser, tenant check) and make it protective. New campaigns inherit the protection without anyone remembering to add it.

## Delivery is half the job

A model call completing is not an experience succeeding. The engine needs explicit workflow state for every record: accepted, context ready, generated, validated, persisted, rendered, delivered, confirmed [1]. With that state, Larkspur's missing pages are a rendering defect found by a query, not a mystery found by a customer.

The state keeps five kinds of success apart, because each fails on its own: **generation** (did a step produce output?), **content** (is it complete and safe?), **persistence** (can the next step read it yet?), **delivery** (did the destination get the right artifact?), and **campaign** (does the population behave as the contract says?), and each needs a different recovery. Persistence is the one teams skip: a successful write is not always readable milliseconds later, so the engine I built confirms state that feeds a customer page with a bounded read before rendering.

Chapter 17 argues the operational disciplines in detail; here they are as a deployment checklist:

- **Idempotency keys** on intake, generation jobs, callbacks, and recovery, because exactly-once delivery is hard in distributed systems [3].
- **Bounded retries with backoff and jitter**, at one layer [4]. Retry a fault. Never retry a decision.
- **Circuit breakers per provider**, so an exhausted dependency trips once instead of failing on every record.
- **Three result classes**: full, degraded, and fail closed. Specificity should degrade before reliability does.
- **Recovery that targets the missing artifact.** If the page exists and the callback failed, resend the callback. Do not rewrite the account.
- **Separate internal and customer wire types**, so adding an internal field never publishes it. The wire schema also carries the destination's field limits: where a downstream system caps a field, generate to a budget below the cap, then validate and truncate before the payload leaves.

None of this is specific to AI. That is the point. Once the model is treated as an unreliable dependency, ordinary distributed-systems engineering does most of the work.

## Watch business state, not only logs

Logs catch exceptions. They do not catch an API that returns 200 while a required artifact is missing, a callback that completes with degraded content, or a provider that quietly returns worse data [1]. Larkspur's Tuesday failure produced no errors at all.

So I monitor two classes of signal: **system signals** (errors, timeouts, queue depth, provider failures) and **workflow signals** (stuck records, missing artifacts, fallback rate, recovery saturation, defect classes). The second class matters more, because silent degradation is the common failure. And a monitor needs a heartbeat: one that silently stops must not look like a healthy system.

Cost belongs in the same view. Personalization gets expensive through fan-out, not any single call: several providers, a dozen zone generations, an audit, three emails, a page and a playbook per record, multiplied by retries. The rule I deploy with: spend model intelligence on uncertainty and judgment, and use ordinary software for everything deterministic. Cheap checks run first; a failed callback never triggers regeneration; semantic QA runs on a worklist and widens only when risk justifies it.

Evaluation modes need the same discipline. In a notification engine I built for my own product, separate from the personalization engine (Chapter 16 sets out the notification decision), full AI evaluation cost about $0.004 per event, mostly spent deciding not to notify anyone. Right for calibration, wrong for production, where a daily per-user digest cost about $0.01 per user per day [5]. Deployment is where calibration mode should end.

## One campaign cannot corrupt another

Once the engine serves more than one campaign, a new class of defect appears: correct content in the wrong context. A company that appears in two campaigns must not share one page. A guideline found by a loose name search can be the wrong version or the wrong tenant. A callback can belong to a different downstream organization.

The protection is boring and complete: scope every durable artifact by the keys that define uniqueness (tenant, campaign, account or person, artifact type, version), and resolve authentication and tenant routing before any enrichment starts [1]. No model evaluation will catch a cross-campaign defect. Only keys will.

## The first cohort is part of the build

A campaign is not ready because it passed a handful of test records. The first few hundred real records reveal behavior nobody anticipated. In the engine I built, an operator widens the cohort in steps (synthetic records, then 10 to 25 real ones, 50 to 100, a few hundred) and at each step uses AI to reason across the population, asking what repeats and what belongs in code rather than another guideline. Chapter 3 frames this as qualitative research on the machine's own output. For high-value campaigns, a domain expert reviews the first outputs; that is calibration, not a person reading every record. Chapter 20 gives the method: population QA, semantic review on a worklist, root-cause classification by layer, and every fix validated on replay, held-out, and regression sets.

For deployment, three decisions follow. First, staff the first cohort as a build phase with engineers attached, not as a launch with a support rota. Second, gate widening on explicit verdicts (GO, GO WITH FIXES, NO GO) against thresholds written in the contract. Google's SRE practice of canarying, comparing a small slice against a control before widening, is the right mental model [6], judged here on content quality and delivery completeness as well as errors. Third, decide in advance who may promote changes. Today, in my engine, a person diagnoses and decides; the design I would adopt as automation grows is Chapter 18's human-supervised autonomy, with two approvals and no production credentials on the analysis worker. NIST's AI RMF treats governance as continuous across the lifecycle rather than a single gate [7]; this is that idea as an operating model.

## Deployment phases

Here is the order the architecture implies. It is an ordering, not a promise of duration.

| Phase | What gets built or decided | Exit criterion |
|---|---|---|
| 0. Decide and contract | The first decisions to personalize (Chapter 6); the contract; each requirement sorted by control type; release thresholds | A contract a reviewer could fail a record against |
| 1. Foundations | Intake with idempotency; identity and governed context; shared mechanics at choke points; workflow journal; wire schemas | Any record's state and provenance can be answered by query |
| 2. Shadow and first cohort | Synthetic records, then real ones in widening cohorts, delivered to internal reviewers or a small live slice; expert review of first outputs where the stakes justify it; population QA; layer diagnosis | GO verdict against the written thresholds |
| 3. Supervised expansion | Wider cohorts against a control or holdout (Chapter 21); business-state monitoring; cost ceilings; recovery tested on purpose | Stable defect, fallback, and cost rates at the wider volume |
| 4. Second campaign and supervised improvement | A second campaign on the same shared mechanics; scheduled analysis with two approvals | The second campaign launches for clearly less effort than the first |

A 90-day sketch under this order, for a mid-market company with one decision in scope: roughly the first three weeks on Phase 0 and discovery, the next five on Phase 1, then about three weeks of first cohort, and the final two on the expansion decision and the second campaign's contract.

### After the second campaign: the productization sequence

Deploying an engine and turning it into a product are different jobs. The sequence I am following for the engine I built is: **standardize, measure, productize operator workflows, expose safe agent control, broaden channel compilation, increase supervised autonomy**. Most of it is ahead of me; read it as direction, not results.

Standardize means one machine-readable campaign manifest (surfaces, personalization mode, templates, guidelines, research mode, quality policy, delivery requirements) that runtime, QA, onboarding, agents, and reporting all read. Measure means every report says what was expected, delivered, degraded, and recovered, and which version produced it. The human-run population review becomes a repeatable, read-only step that ends in a proposal. Agents get stable domain operations instead of raw cloud access. New channels come next, print first with stricter preflight, and autonomy widens last, with production behind a human approval (Chapter 18).

The six-month target: **repeatable** (a new campaign does not need the original architect), **measurable**, **AI-operable**, **human-controlled**, **multi-channel**. A target, not a result.

A warning about discovery: in the manual process I used to run, understanding a client's systems well enough to design an integration took six to ten weeks [8], mostly rediscovering what the code already knew. Point an AI assistant at the codebase and data model first, and spend the meetings on decisions.

## Build, buy, or own the layer

Most AI decisions in customer operations get made as a feature comparison: which vendor's agent looks best. That question hides the one that shapes the architecture. Which parts of this are commodity enough to rent, and which are strategic enough that you should own the intelligence and the economics behind them?

Renting is frequently right: the vendor carries the risk, the support burden, and the reliability commitment. The published evidence leans toward buying, with caveats. MIT's Project NANDA reported that, in its interview sample, external partnerships with "learning-capable, customized tools" reached deployment about 67% of the time, against about 33% for internal builds [9]. I take the direction seriously and the number lightly. The sample is 52 organizations, outcomes are self-reported, success definitions varied, and the authors themselves warn that the gap "may reflect organizational capabilities rather than implementation approach alone." The same report names the real barrier as a learning gap: tools that do not retain feedback, adapt to context, or improve over time [9]. That finding cuts across build versus buy. Whatever you choose, the learning has to accumulate somewhere you can use it.

That is my test. Three loops compound inside a personalization layer. **Memory**: every identity resolved and fact written down once is work the next operation does not repeat. **Routing**: every step priced to its real risk, a small model on narrow work and a large one where judgment lives, is a decision made once. **Acceptance**: every output a person reviews becomes a labeled example of what good looks like in your business. A rented layer can reset all three at the contract boundary. Rent the commodity; own what compounds.

Either way, ask five questions, in order of what a wrong answer costs [10]: What can it do that cannot be undone? Who approved each action, and can you prove it six months later? Who sets its ceilings? Does it work per step, not just on average? When it fails, does it degrade or collapse? Building does not exempt you from answering them.

If you own the layer, keep workers replaceable: the model underneath will change fast (GPT-3.5-level inference fell from $20 to $0.07 per million tokens between November 2022 and October 2024 [11]). And consider where it runs. Deploying into the customer's own infrastructure, with their keys and audit logs, lets their security team inspect the controls instead of taking a vendor's word for them.

The self-hosted memory system I built shows the shape (architecture in my paper [16]). It is one stateless container on the customer's own Postgres with pgvector, and that database holds everything: memories, governance documents, schema, job queue, relationship graph, usage counters. It runs on any major cloud, Kubernetes, one on-prem machine, or air-gapped with a signed offline license and local models. Models are chosen per function (small for extraction, stronger for generation, or local for all of it). The vendor's only call is a license check carrying usage counts, never content; the offline license removes even that.

That keeps the memory and routing loops the customer's by construction, and moves work to them: their database is the single point of failure unless run highly available, and the embedding dimension is fixed at first boot, so a different-sized embedding model means re-embedding everything. Nor does self-hosted mean nothing leaves: a hosted model still receives the content it processes. Redaction and tokenization shrink that (Chapter 19); only local models send nothing. [PRODUCT-NAMING: named worked example of a governed engine and a self-hosted memory system deployed in the customer's own infrastructure, if Hamed decides to name them.]

## The team is small, and the roles are specific

A lean team can run this; it cannot skip the roles. When agents do the work, coordination people once carried becomes schemas, memory, evaluations, and permissions [12]. Responsibility moves; it does not disappear.

| Role | Owns |
|---|---|
| Product or system owner | Purpose, the decisions in scope, success criteria, the go decision |
| Campaign owner | The contract, guidelines, and examples for each campaign |
| Identity and memory owner | Matching, provenance, freshness, correction, deletion |
| Platform engineer | Shared mechanics, delivery state, recovery, cost controls |
| Evaluation owner | QA method, release thresholds, the three validation sets |
| Governance approver | Irreversible-action list, the second approval, policy precedence |
| Channel owners | Each surface's compiler policy and final-artifact checks |

In a small company, one person holds several of these. The separations that must survive are the ones between proposing and approving, and between the person who writes a guideline and the evaluation that judges it.

## The same design at three scales

The design does not require enterprise complexity. It requires the separations.

| Organization | Scope | What matters most | Practical shape |
|---|---|---|---|
| Enterprise | Many campaigns, tenants, regions, and channels | Governance plane, campaign isolation, recovery at volume | Separate services per the reference pattern; request-serving and worker pools scaled separately; dedicated owners per role |
| Mid-market (Larkspur) | A few decisions across CRM, warehouse, product telemetry, support | One identity spine and one governed context feeding two or three channels | One pipeline with clear modules; two replicas on a managed database; a handful of people holding several roles |
| Small business | One owner, one CRM, a few channels | Fallbacks and one human approval before anything irreversible | One process performing separated steps, on one machine and one database |

The small version may run in one process. Evidence, inference, generation, validation, and authority still stay separate.

Scaling out stays boring when processes hold no state. In the self-hosted system, replicas coordinate only through the database (workers claim jobs with row locks that skip rows already held; the scheduler advances each schedule with a conditional update), so every job fires once however many replicas run. Request-serving and background workers can be split and scaled on their own signals (workers on queue depth), and in a container, set job concurrency explicitly: auto-sizing reads the host's cores, not the container's limit, and over-provisions. The container is rarely the ceiling: provider rate limits bind first, and as replicas multiply, database connections run out before CPU.

## Score maturity on six dimensions, not one ladder

A single ladder misleads: a company can have excellent decisioning and fragmented identity. Score each dimension on its own, then fix the weakest one that blocks the decision you care about.

![A grid of six maturity dimensions (understanding, identity and memory, decisioning, experience generation, learning, governance and reliability), each with its own five levels, from Generic to Real-time and uncertainty-aware for understanding, and from Ad hoc to Consequence-aware autonomy for governance. The example from the text is marked: the same company shows excellent decisioning at the top of its row while identity sits at Fragmented, Level 1. Each dimension is scored on its own, and the weakest one blocking the target decision is fixed first.](/images/handbook/ch20-six-dimension-maturity.svg)

*Figure 22.3. Maturity is a profile, not a ladder: one company can be excellent at decisioning and fragmented on identity.*

| Dimension | Level 1 | Level 2 | Level 3 | Level 4 | Level 5 |
|---|---|---|---|---|---|
| Understanding | Generic | Segment | Individual | Contextual | Real-time, uncertainty-aware |
| Identity and memory | Fragmented | Unified profile | Graph and context | Durable memory | Governed, adaptive memory |
| Decisioning | Static | Rules | Predictive | Optimized | Reasoning within constraints |
| Experience generation | Static | Variants | Dynamic selection | Generative | Interactive, adaptive |
| Learning | Manual reporting | A/B tests | Causal measurement | Bandits, online learning | Governed continuous improvement |
| Governance and reliability | Ad hoc | Policy checks | Observable decisions | Automated guardrails | Consequence-aware autonomy |

The companion Assess skill interviews a team and scores these six dimensions with evidence; the Plan skill turns the scores into a build sequence and a 90-day plan. Both are built to show gaps, not to flatter.

## What this does not do

**It does not make the decision worth making.** A reliable engine pointed at a low-value decision delivers low value reliably. Chapter 6 comes before this one for a reason.

**The phases are an ordering, not a guarantee.** I can show why each step comes before the next. I cannot promise how long each takes in your company, and neither can anyone quoting a universal timeline.

**The failure-rate headlines are weaker than they sound.** Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025 [13], and that over 40% of agentic AI projects will be canceled by the end of 2027 [14]. Those are forecasts. MIT's "95% of organizations are getting zero return" comes from the preliminary report above [9]. The causes they name (unclear value, cost, weak risk controls, systems that do not learn) fit this chapter; the percentages should not be quoted as findings.

**My evidence is my own deployments,** not a controlled comparison of architectures. I describe what repeatedly broke and what fixed it; I cannot give you a population-level success rate for this approach. The productization sequence and its six-month target are a plan, not an outcome. And a vendor with strong memory, governance, and export rights can beat an owned build: the test is what accumulates and who can use it, not who wrote the code.

## At scale

Across millions of people and many teams, four things change. Campaign count, not record count, becomes the growth axis, so the share of work done by shared mechanics decides cost per launch. Recovery becomes a standing service, because some records are always mid-failure. Governance needs its own versioned plane, so a policy change is a reviewed release with a rollback target [1]. And the person who can answer "why did this customer receive this?" must exist, per campaign and surface, before the question is asked.

## Failure story: The Campaign Fork

This is Larkspur's history from before the engine (fictional, as before). An early AI campaign was a success, built fast: one long prompt, a few trust rules written inline, a renderer with the campaign's CTA hard-coded. For the second campaign, a different team copied the repository and edited the prompt. The third copied the second. By the fifth, a fix for free-mail domains existed in two forks and not in three, a tone change requested for one campaign had leaked into another through a shared template, and nobody could say which version of which guideline had produced a given email.

Nothing about the model had changed. Every campaign had become a new engine, and every engine had to learn every lesson again. That history is the reason the engine in this chapter starts from shared mechanics.

## Patterns

**Contract First.** *Problem:* QA and release decisions have nothing to judge against. *Solution:* a machine-readable contract per campaign and surface, with release thresholds, written before generation. *Tradeoff:* slower first week; much faster every week after.

**Shared Mechanics, Configured Judgment.** *Problem:* each new campaign forks the engine. *Solution:* invariants in code at shared choke points; campaign intent in versioned guidelines and configuration. *Tradeoff:* discipline about where each requirement lives, enforced in review.

**Own What Compounds.** *Problem:* the layer that accumulates memory, routing decisions, and reviewed examples resets at every contract boundary. *Solution:* rent commodity capabilities; own, or secure export and control of, the compounding layer. *Tradeoff:* more responsibility for reliability and support.

## Leader questions

1. Show me the contract for the first decision. Could a reviewer fail a record against it?
2. For any record, can you tell me its workflow state, exactly what was delivered, and which version produced it?
3. Which protections live at shared choke points, and which live only in prompts?
4. Who may promote a change to production, and what credentials does the automated analysis hold?
5. What will the second campaign cost to launch compared with the first?

## Build checklist

- [ ] Contract per campaign and surface, including release thresholds and fallback policy
- [ ] Every requirement sorted by control type: guideline, deterministic, or both
- [ ] Idempotency keys on intake, generation, callbacks, and recovery
- [ ] Workflow journal with per-record state from accepted to confirmed
- [ ] Circuit breakers and bounded retries per provider
- [ ] Full, degraded, and fail-closed result classes, labeled in telemetry
- [ ] Artifacts scoped by tenant, campaign, entity, artifact type, and version
- [ ] Business-state monitoring with a heartbeat
- [ ] First cohort scheduled as a build phase with a GO / GO WITH FIXES / NO GO gate
- [ ] Two approvals for production changes; no production credentials on analysis workers

## Metrics to watch

1. **Delivery completeness**: share of records whose expected artifacts were confirmed delivered.
2. **Fallback and fail-closed rates** by surface, trended against identity and evidence quality.
3. **Cost per accepted artifact**, including retries and recovery, not just model calls.
4. **Time and effort to launch the next campaign**, the clearest sign of platform versus pilot.
5. **Defect recurrence** after a fix, by root-cause layer.

## Reader Q&A

**Can we start without identity resolution and add it later?** You can start with a narrow scope, but not without the rule that personalization specificity falls when identity confidence falls. Without it, the engine personalizes confidently to the wrong entity from the first day.

**Do we need a multi-agent system?** Rarely at first. A focused function per zone is often enough. Anthropic's engineering guidance says the same: find the simplest solution possible and add complexity only when it demonstrably improves outcomes [15].

**What if our vendor will not let us export memory or review data?** Then price that in. You are renting the compounding loops as well as the software.

## For your AI

```yaml
chapter: 22
concepts:
  - name: Experience Contract
    definition: "Machine-readable specification of audience, surfaces, personalization tier, approved and prohibited claims, fallback policy, enforcement type per rule, and release thresholds."
  - name: Shared Mechanics
    definition: "Invariants enforced in code at shared choke points that every campaign inherits: sanitizing, escaping, idempotency, retry classification, internal-field stripping, tenant routing, wire schemas."
  - name: Campaign Judgment
    definition: "Voice, priorities, examples, and strategy held in versioned guidelines and configuration, editable per campaign without an engine deploy."
  - name: Workflow State
    definition: "Per-record status from accepted to confirmed delivery, used to diagnose failures and target recovery."
  - name: Own What Compounds
    definition: "Rent commodity capabilities; own or control the layer where memory, routing decisions, and reviewed examples accumulate."
  - name: Five Kinds of Success
    definition: "Generation, content, persistence, delivery, and campaign success, tracked separately because each fails on its own and needs a different recovery."
  - name: Productization Sequence
    definition: "After the second campaign: standardize (one campaign manifest), measure, productize operator workflows, expose safe agent control, broaden channel compilation, increase supervised autonomy."
  - name: Six-Dimension Maturity
    definition: "Independent scores for understanding, identity and memory, decisioning, experience generation, learning, and governance and reliability."
decision_rules:
  - if: "no written contract with release thresholds exists for a decision"
    then: "do not generate for real records; write the contract first"
  - if: "a requirement must hold for every campaign"
    then: "enforce it in code at a shared choke point, not in campaign prompts"
  - if: "a requirement expresses one campaign's intent"
    then: "put it in versioned campaign guidelines or configuration"
  - if: "a failure is a deterministic outcome (missing context, invalid contract, policy rejection)"
    then: "do not retry; fall back or fail closed"
  - if: "an artifact exists but delivery failed"
    then: "redeliver the artifact; never regenerate or re-research"
  - if: "a capability accumulates memory, routing, or reviewed examples"
    then: "own it or secure export and control before renting it"
  - if: "a write feeds a customer-facing artifact"
    then: "confirm it with a bounded read before rendering; a successful write is not yet a readable one"
  - if: "policy or residency forbids vendor storage of customer data"
    then: "run the layer on the customer's infrastructure and database; use local models if no content may leave"
  - if: "the second campaign costs about as much to launch as the first"
    then: "stop adding campaigns and extract shared mechanics"
assessment_questions:
  - "Which decision is first, and is there a written contract for it?"
  - "Can you query any record's workflow state and delivered artifacts?"
  - "Which protections live at shared choke points versus in prompts?"
  - "Who approves production changes, and what credentials do automated workers hold?"
  - "Which parts of the stack are rented, and can you export memory and review data?"
  - "How do you score on each of the six maturity dimensions, with evidence?"
patterns: [Contract First, Shared Mechanics Configured Judgment, Own What Compounds, Recover The Smallest Unit, Human-Supervised Autonomy]
anti_patterns: [The Campaign Fork, Whole-Pipeline Replay, Logs-Only Observability, Demo-Driven Launch, Autonomous Production Mutation]
maturity_dimension: governance_and_reliability
handoff:
  assess_skill: "score the six maturity dimensions"
  plan_skill: "produce build sequence, phase plan, team roles, and 90-day plan from this chapter"
```

## References

1. Taheri, H. (2026-08-23). "How I Build an Enterprise Personalization Engine That Can Be Trusted at Scale." https://hamedtaheri.com/articles/building-enterprise-personalization-engine
2. Sculley, D. et al. (2015). "Hidden Technical Debt in Machine Learning Systems." Advances in Neural Information Processing Systems 28 (NIPS 2015). https://papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
3. AWS Well-Architected Framework, "REL04-BP04 Make mutating operations idempotent," accessed 2026-09-26. https://docs.aws.amazon.com/wellarchitected/latest/framework/rel_prevent_interaction_failure_idempotent.html
4. AWS Well-Architected Framework, "Control and limit retry calls," accessed 2026-09-26. https://docs.aws.amazon.com/wellarchitected/2025-02-25/framework/rel_mitigate_interaction_failure_limit_retries.html
5. Taheri, H. (2026-03-07). "Dogfooding governed memory: building smart notifications for our own product." https://hamedtaheri.com/articles/dogfooding-governed-memory
6. Google, *The Site Reliability Workbook*, Chapter 18, "Canarying Releases." https://sre.google/workbook/canarying-releases/
7. NIST (2023-01-26). *Artificial Intelligence Risk Management Framework (AI RMF 1.0)*, NIST AI 100-1. https://www.nist.gov/itl/ai-risk-management-framework
8. Taheri, H. (2026-03-10). "Encoding Solution Architecture Into an AI Skill." https://hamedtaheri.com/articles/solution-architect-skill-deep-dive
9. Challapally, A., Pease, C., Raskar, R., Chari, P. (2025-07). *The GenAI Divide: State of AI in Business 2025*. MIT Project NANDA (preliminary findings). https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
10. Taheri, H. (2026-08-10). LinkedIn carousel on five agent-platform questions.
11. Stanford HAI (2025). *AI Index Report 2025*. https://hai.stanford.edu/ai-index/2025-ai-index-report
12. Taheri, H. (2026-06-29). "The Lean Agentic Company." https://hamedtaheri.com/articles/the-lean-agentic-company
13. Gartner (2024-07-29). "Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025." https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
14. Gartner (2025-06-25). "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027." https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
15. Schluntz, E., Zhang, B. (2024-12-19). "Building effective agents." Anthropic Engineering. https://www.anthropic.com/engineering/building-effective-agents
16. Taheri, H. (2026-03-18). "Governed Memory: A Production Architecture for Multi-Agent Workflows." arXiv:2603.17787. https://arxiv.org/abs/2603.17787
