Closing: The Generative Experience, and Your Plan

Where is this going, what will still be true when it gets there, and what do I do on Monday?

Working draft, revised September 2026 · 14 min read · Markdown for your AI

Questions this closing answers

  • What does a company look like a year after it takes depth seriously?
  • Will websites and products be generated for each person? What ships now, and what is still a demo?
  • If models keep changing every few months, what should we build that will not expire?
  • How do I hand this handbook to my own AI and get a plan back?

The short answer

The next step in personalization is not a better sentence. It is a whole experience composed for one person: the page, the offer, the product screen, the conversation. Parts of that ship today; much of it is still a lab demo, and I label which is which below, as of September 2026. What will not change with the next model is the work underneath: a memory that holds an accurate, current model of each customer, a decision layer that chooses what should happen next (including nothing), governance that runs inside the system, and measurement that proves cause. A generated interface only raises the stakes on all four, because every component on a composed page is a claim about the person looking at it. The practical move is to give this handbook to your AI, let the Assess skill score where you stand, and let the Plan skill draft your first three decisions and a 90-day build. Then do the unglamorous part yourself.

Tuesday, 9:00 a.m., a year later

Larkspur Systems (the fictional composite company that has run through this handbook) still sends a quarterly newsletter. It goes out at nine on a Tuesday, as it did a year ago. It goes to fewer people now, and it is honestly general: a product release note for contacts about whom Larkspur knows nothing that should change the message. Nobody pretends it is personalization. (All Larkspur details and numbers here are illustrative.)

The four people from the first page of this handbook got something else.

The customer whose renewal was thirty days out, and who had stopped opening the reporting module in spring, did not receive a renewal email first. Priya, her account manager, received a brief overnight: the drop in reporting use, the two tickets that preceded it, the name of the admin who had built the original dashboards and left in April. The brief cited its sources. Priya called. The email that followed was written after the call, and it said what Priya had promised.

The operations director with three open tickets about route sync received nothing from marketing that week. The decision layer suppressed the campaign she had been scored for, flagged the escalation to her account manager, and requeued the offer for thirty days after the fix held. When the offer came, it opened with the fix, not the feature.

The new head of operations, whose name had surfaced from a support signature, was not added to a nurture track. Memory resolved him to the account and marked the fact's source and date. His first message came from his account manager and explained what his team already used, because a new leader's first question is usually what he has inherited.

The account that had doubled its fleet was not sent a discount. The Discount Mistake from Chapter 2 was now a rule in the decision layer. It got a plan-fit review and a conversation about capacity.

None of this came from a smarter model. Larkspur's year looked like the table of contents. It fixed identity after a wrong merge. It built one memory that every agent reads, with a curation cycle that keeps it true. It put a contact budget and one arbitration point above six teams' programs. It turned its guidelines into checks that run before and after generation. Its first cohort of 2,000 messages failed a population-level scan in ways no reviewer caught, and the fixes were in four layers, none of them "a better prompt." Its open-rate chart went up while pipeline quality fell, and that is when it started a holdout. Some decisions still cannot be proven at Larkspur's volume, and the team says so.

A year in, the most visible change is not what Larkspur sends. It is what Larkspur no longer sends.

A year later, each of the four Larkspur (fictional) customers from the introduction got a decision instead of the newsletter. The renewing account's account manager was briefed overnight with cited sources, called first, and sent the email after the call. The angry operations director's campaign was suppressed, her escalation flagged to her account manager, and the offer requeued for 30 days after the fix held, opening with the fix. The new head of operations was resolved to the account with the fact's source and date, skipped the nurture track, and heard first from his account manager; the account that doubled its fleet got no discount, because the Discount Mistake is now a decision-layer rule, but a plan-fit review and a capacity conversation. The newsletter still goes out, to fewer people and honestly general.

Figure 23.1. A year later, each of the four got a decision rather than the newsletter; the biggest change is what Larkspur no longer sends.

The page is next

A year ago, Larkspur's unit of personalization was the message. The unit is moving to the experience.

The old message was a letter. It was written once, it said the same thing to everyone, and once sent it could not respond. What has been arriving for a few years is a conversation, one that can be different with each person and can answer the question the letter never heard. The same shift is now reaching interfaces. A page, a dashboard, a form, or an onboarding flow can be composed at the moment a specific person asks for it.

Here is where that stands, labeled, as of September 2026 (Chapter 16 works through the mechanics).

CapabilityStatusEvidence
Per-person text and images in campaignsShips nowRoutine in production. In a field test of more than 173,000 ad impressions, AI-generated images earned up to 50% higher click-through than stock images [1].
Interactive UI returned inside AI assistantsShips nowMCP Apps became an official Model Context Protocol extension on January 26, 2026: tools return dashboards, forms, and workflows that render in the conversation, with support announced in Claude, ChatGPT, Goose, and VS Code Insiders [2].
Interfaces generated per prompt in consumer searchShips, as an experimentGoogle rolled out generative UI in the Gemini app and, for Pro and Ultra subscribers in the US, in Search AI Mode on November 18, 2025 [3].
Declarative component catalogs for agent UIsEarly specificationGoogle's A2UI (v0.9 in a June 17, 2026 post) has agents send JSON describing what to render; the host renders only trusted components from a predefined catalog [4].
Fully generated pages per visitor, at web latency, on a company's own siteLab demoGoogle's own researchers report generation can take "a minute or more" with "occasional inaccuracies", and raters still preferred pages made by human experts [3].
Most customer-facing surfaces composed per person from memory, governed components, and a decision layerForecast (mine)An inference from the rows above, not a finding.

Two cautions about the forecast. The first is speed. METR's measure of how long a software task an AI agent can complete at 50% reliability has doubled roughly every seven months since 2019 [5], and METR itself warned in January 2026 that a 50% horizon of X hours "does not mean we can delegate tasks under X hours to AIs" [6]. Capability curves are real; they are not delivery dates.

The second is the claim you will hear most often: that every website will soon be generated fresh for every visitor. I do not accept it as stated. Google's evaluation, the most careful public one I know, still ranked expert-built pages first [3]. My expectation is narrower and, I think, more useful: a stable frame, with the parts that should vary composed per person from a governed catalog. The industry is converging on the same shape; the declarative approaches above restrict agents to allow-listed components for safety and consistency [4]. It is also the shape I ship. In the governed personalization engine I built, a landing page is a fixed shell: brand, navigation, and legal copy never vary, and roughly 8 to 12 declared text zones do. Each zone carries a contract (the context it may use, the context it must never use, a word limit, the checks it must pass) and approved segment copy to fall back on.

That shape is the whole handbook in miniature. The depth is precomputed and stored (Chapters 10 and 17). The component is chosen by a decision (Chapter 14). Every claim a component makes about the person must be grounded or left out (Chapter 15). The catalog is governed like brand guidelines, because it is brand guidelines (Chapter 18). The generated page does not replace the architecture. It is the architecture's newest output.

Where I am taking it next is direction, not results, as of September 2026. The same governed context should reach paper: an approved print template with declared zones and a deterministic preflight, whose QR code continues into a personalized page (Playbook M3). And it should start before identity: personalize to the campaign context a visitor arrived with (the audience, industry, and offer the ad was built for), then to the known person and account once a form tells you who they are. The second half runs today; the first is a pilot design.

What stays true when the model changes

Here is the sentence this closing turns on. Models will change every few months; the model of your customer is the asset you keep.

Four things in this handbook do not depend on which model you run.

Memory. An accurate, current, sourced model of each person, resolved to the right identity and curated continuously. A better model reading bad memory writes better-sounding mistakes.

Decisions. What should happen next for this person, chosen from options that include waiting and doing nothing, under hard constraints. Generation answers "how do we say it"; the decision layer answers "should we say anything."

Governance. Rules executed as code paths before and after the probabilistic step, with a trace of what was known, decided, and why. An interface composed per person needs this more than an email does, because it has more places to be wrong.

Measurement. A counterfactual, a holdout, and outcomes that are not proxies. When every surface adapts, the only customer who tells you the truth is the one who did not get the adaptive version.

In the self-hosted memory system I built, that separation is literal. Memory and governance documents live in the customer's own Postgres; the model is a configuration entry per function, from several providers or a fully local model. Swapping the generation model is a config edit, and the memory does not move (changing the embedding model does mean re-embedding). Memory gives the next model continuity; governance gives it boundaries.

Models are swappable: today's model gives way to the next one and the one after, every few months. Four layers do not depend on which model runs. Memory is an accurate, current, sourced model of each person; decisions choose what happens next, including waiting and doing nothing; governance runs rules as code paths before and after generation, with a trace; measurement uses a counterfactual, a holdout, and outcomes that are not proxies. The standard, specific and true or honestly general, does not expire either.

Figure 23.2. Models change every few months; memory, decisions, governance, and measurement are the durable four you keep.

The same boundary holds when AI improves the system itself. Today, in systems I run, AI proposes: a memory-curation pass starts by only proposing changes, and an operator promotes it stage by stage, with a stop that works mid-run. The direction I am building toward, not yet shipped, lets AI diagnose a campaign, fix it outside production, and validate the fix on replayed failures and held-out examples; production stays a separate human approval. High reasoning autonomy, bounded execution authority, explicit production control.

There is one more thing that stays true, and I will say it once. A system precise enough to compose an experience that serves one person is precise enough to compose one that pressures them, and which one it builds is set by an objective that a builder chooses.

Specific and true, or honestly general. That standard does not expire either.

Your plan: give this handbook to your AI

You do not need to read this handbook front to back, and you do not need to translate it into a plan by hand.

Step 1. Give it the text. Point your assistant at the full-text file or the chapters you care about. Every chapter ends with a For your AI block: stable concept names, decision rules, assessment questions, and named patterns and anti-patterns. Those blocks are what the skills read.

Step 2. Run Assess. Install the Assess skill from the companion repo. It interviews you about your industry, channels, data sources, identity situation, stack, regions, team, and current personalization. It scores you on six maturity dimensions (understanding, identity and memory, decisioning, experience generation, learning, governance and reliability), each with evidence and gaps. Bring real artifacts: one actual customer record, one live campaign, one recent incident. Answers from artifacts are more honest than answers from memory.

Step 3. Run Plan. It takes your assessment and the handbook and drafts:

  • your first three decisions to personalize, scored by value, data readiness, reversibility, and time to evidence (one fast win, one memory-building bet, one measurement foundation);
  • a starter memory schema and freshness policy;
  • starter governance rules and an irreversible-action list;
  • a privacy checklist for your regions (dated, and not legal advice);
  • an accuracy stack and a first-cohort protocol;
  • a measurement plan with a sized holdout;
  • a 90-day build sequence and the playbooks that fit your functions.

It is built to make gaps visible, not to flatter. If your identity layer cannot support the decision you most want, the plan will say so and put identity first. It will not promise revenue; it will show value as ranges, with the assumption carrying the weight made visible.

Step 4. Argue with it. The plan is a draft for your team, not a decision. Take it to the people who own the data, the channels, and the risk. The best use of the output is the meeting it starts.

One at a time

In 1993, Peppers and Rogers asked companies to build relationships "one customer at a time" [7]. The phrase described something the economics could not yet deliver: understanding each person cost a skilled human's hours. That cost has fallen. The standard has not moved. Being right about one person still takes a sourced memory, a decision that is allowed to say "not now," and rules that hold when nobody is watching.

You will build this for hundreds of thousands of people. Each of them will meet it alone, one at a time.

Leader questions

  1. Which of our customer-facing surfaces is closest to being composed per person, and which of its components makes a claim we cannot source?
  2. If we switched model providers next quarter, what would we lose? If the answer is "a lot", where does our customer understanding actually live?
  3. What did we stop sending this year because a decision said not to?
  4. Which personalization result would survive a holdout?
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 23
concepts:
  - name: Generative Experience
    definition: "An interface, page, offer, or flow composed for one person at request time, from memory, a decision, and governed components."
  - name: Governed Component Catalog
    definition: "An allow-listed set of UI components an agent may compose, each with a contract for what it may claim; governed like brand guidelines."
  - name: Stable Frame, Variable Parts
    definition: "Keep navigation and structure stable; compose only the parts that should vary per person."
  - name: Durable Four
    definition: "Memory, decisions, governance, and measurement: the layers that outlast any model choice."
decision_rules:
  - if: "a capability is described as the future of personalization"
    then: "label it ships-now, lab demo, or forecast, with an as-of date, before planning on it"
  - if: "a surface is to be generated per person"
    then: "precompute depth, choose components by decision, ground every personal claim, and render only from a governed catalog"
  - if: "a generated component would state a fact about the person without a sourced memory"
    then: "omit it or render the honestly general version"
  - if: "a model change is proposed"
    then: "confirm memory, decisions, governance, and measurement are provider-independent before switching"
  - if: "an AI agent proposes a change to live rules, memory, or a campaign"
    then: "let it reason, implement outside production, and validate; require an explicit human approval to promote to production"
assessment_questions:
  - "Which customer-facing surfaces could vary per person, and what decision would choose each variant?"
  - "Where does customer understanding live today: in prompts, in one vendor's tool, or in a memory you own?"
  - "Is there a holdout on each adaptive surface?"
  - "Which claims about a person can a generated interface make, and who approves the component catalog?"
patterns: [Stable Frame Variable Parts, Governed Component Catalog, Precompute-then-serve]
anti_patterns: [Every Page Generated, The Proxy Trap, The Hallucinated Detail]
maturity_dimension: experience_generation
skills_handoff:
  assess: "interview -> scores on six maturity dimensions with evidence and gaps"
  plan: "assessment + handbook -> first three decisions, memory schema, freshness policy, governance and irreversible-action list, privacy checklist, accuracy stack, measurement plan, 90-day sequence, playbooks"

References

  1. Hartmann, J., Exner, Y., and Domdey, S. (2025). "The power of generative marketing." International Journal of Research in Marketing 42(1). https://www.sciencedirect.com/science/article/pii/S0167811624000843
  2. Model Context Protocol Blog, "MCP Apps: Bringing UI Capabilities To MCP Clients," 2026-01-26. https://blog.modelcontextprotocol.io/posts/2026-01-26-mcp-apps/
  3. Leviathan, Y., Valevski, D., Natchu, V., Matias, Y., "Generative UI: A rich, custom, visual interactive user experience for any prompt," Google Research, 2025-11-18. https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/
  4. Google Developers Blog, "A2UI + MCP Apps: Combining the best of declarative and custom agentic UIs," 2026-06-17. https://developers.googleblog.com/a2ui-and-mcp-apps/
  5. METR, "Measuring AI Ability to Complete Long Software Tasks," 2025-03-19. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
  6. METR, "Clarifying limitations of time horizon," 2026-01-22. https://metr.org/notes/2026-01-22-time-horizon-limitations/
  7. Peppers, D., and Rogers, M. (1993). The One to One Future: Building Relationships One Customer at a Time. Currency/Doubleday.

This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.

Get chapters by email as they are revised

Prefer LinkedIn? Subscribe to the newsletter there instead.