Questions this chapter answers
- Should we personalize whole pages, or sections inside a page we control?
- What is generative UI, and is it ready for our product and our website?
- How do we personalize notifications without wearing people out?
- How do we keep an interface usable when it is different for every person?
- What must never be generated, and what happens when generation fails on screen?
The short answer
A message is read once. An experience is used: a page someone navigates, a product screen they work in every morning, an alert that reaches their lock screen at 6:40 a.m. When AI composes those for one person, three things change. The output has a layout, so it can be broken, inaccessible, or unsafe to render. It is repeated, so a change the person did not expect costs them time every day. And a platform increasingly stands between you and the person, filtering, bundling, even rewriting what you send.
The practical path climbs a ladder. First, personalized zones inside a page you govern: this ships today and holds most of the value. Second, whole pages and product surfaces composed per person from a constrained vocabulary: this ships at large consumer companies, with small, real gains. Third, interfaces generated per person or per request: preferred to plain text by raters, and still too slow and unpredictable for the core of most products as of September 2026. Beside the ladder sits the notification, which is not content at all. It is a decision: whether to interrupt, with what, when, and how, with "not at all" as a scored option.
For a leader, two tests. "When the page changes for one customer, what stays exactly the same for everyone, and who decided that?" And: "How many people turned our notifications off last quarter, and do we count that as a cost?" If nobody can answer the first, the experience will drift. If nobody tracks the second, the channel is being spent without a budget.
The summary that rewrote the news
In December 2024, iPhone users with Apple Intelligence, whose notification summaries had launched that October [30], saw a notification attributed to BBC News reporting that Luigi Mangione, the man charged with killing UnitedHealthcare's chief executive, had shot himself. He had not, and the BBC had never said so. Apple's on-device model had condensed a stack of BBC alerts into one line and invented a fact on the way. Other summaries had declared a darts final won before it was played and turned a warrant into an arrest [1].
Apple first promised clearer labels. Then, in the iOS 18.3 betas of January 2025, it switched summaries off for news and entertainment apps, italicized the rest, let people disable them per app from the lock screen, and warned in Settings that the feature "may contain errors" [2]. Six months later, in the iOS 26 betas, news summaries returned, labeled "Summarized by Apple Intelligence," with per-category controls [3]. When Google shipped notification summaries on Pixel phones in November 2025, it scoped them to long chats and group conversations, not news [4].
Notice who carried the error. The notification said BBC News. The BBC wrote none of it. A platform had composed a personal experience on the BBC's behalf, one lock screen at a time, and the sender's name was the one attached to the invention.
Personalization used to end at the message you wrote. Now the unit is the experience a person moves through, and parts of it are composed by you, parts by your model, and parts by someone else's.
An experience is a message the person has to live inside
The Decision Layer (Chapter 14) chooses what should happen, including nothing. Writing for One (Chapter 15) turns a message into typed zones with contracts, so every personal claim is grounded or left out. That machinery carries over. What changes is the object. An experience has four properties a message does not.
It has a layout. A generated sentence can be false. A generated interface can also be unreadable to a screen reader, unsafe to render, or visually broken.
It is interactive. The person acts inside it, so the system makes decisions while someone is waiting.
It is repeated. A product screen is used daily, and every change the person did not predict costs them the time it takes to find what moved. A novel message is a small pleasure; a novel interface is often a tax.
It is mediated. Browsers, operating systems, assistants, and app hosts rank, bundle, summarize, and sometimes render your content inside their own frames.
Here is the sentence this chapter turns on: the more of an experience you generate, the more of it you must hold still. Everything that follows decides which parts may vary, how much, when, and with what fallback, so the parts that vary serve the person and the parts that stay put keep the experience learnable, safe, and accountable.
The Generative Experience Ladder
Three rungs compose what the person sees. They differ in what varies, who owns the frame, how fast the result must arrive, and how much evidence of production readiness exists; climb only as far as the decision needs. A fourth surface, the notification, reaches the person uninvited. It sits beside the ladder because it is a decision before it is content.
Figure 16.1. The Generative Experience Ladder: each rung varies more of the experience, and each one needs a stronger frame, renderer, and measurement to stay trustworthy.
Rung 1: zones inside a governed page
The first rung is a page designed once, with declared places where content varies per person. It predates language models. Netflix treats each title's artwork as a personalized element of a fixed slot, chosen per member by a contextual bandit, and shows one image per title per member at a time so titles stay recognizable [5]. A 2016 Yahoo paper, best paper at WSDM, learned positions, image sizes, and styles for a whole results page "as long as variations are within business and design constraints," and beat the leading presentation algorithm on real traffic [6]. Commercial testing platforms expose the same shape as named placements: the page asks for a slot's variation, the server resolves targeting and allocation, the page renders the result [7].
Language models add a second verb: a slot once filled by selection can now be filled by generation inside a contract. Chapter 15 covers the contract for one zone; the experience question is how the page holds together. In the governed personalization engine I built, a landing page is a branded template with roughly 8 to 12 declared zones, and legal copy and navigation never vary. Three properties make it an experience rather than a long email:
- The link persists. A personal page has an address the person can return to or forward. It must render the same governed content each time, or change only for a stated reason.
- The identity is scoped. The same company can sit in two campaigns, so the engine scopes page identity by campaign; one program cannot overwrite another's page for the same account.
- Every zone has a fallback. A zone that fails its checks renders approved segment copy. The page never shows a hole or a half-repaired sentence.
Select first, generate second. For each zone, ask whether choosing among five approved variants would serve the person almost as well as writing a sixth. Usually it would, and selection is cheaper, faster, testable by ordinary methods, and impossible to hallucinate. Reserve generation for the sentence that connects this person's stated problem to the offer.
Rung 2: whole pages and product surfaces, per person
The second rung composes the entire page or screen for one person, from a vocabulary that is still constrained. The clearest public example is Netflix. In a paper accepted at RecSys 2026, its engineers describe GenPage, a single transformer that treats the member and request context as a prompt and generates the whole multi-row homepage, replacing a multi-stage recommender stack. Against a mature production system, the A/B test showed +0.24% on the core engagement metric, with serving latency 20% lower [8].
Three details matter more than the headline. The gain is small because the baseline was excellent; that is what real gains on mature surfaces look like. The model generates rows and titles, not free-form interface; the vocabulary is a grid. And offline, enriching the prompt with more context improved the model more than scaling it from 120 million to 900 million parameters [8]. What the system knew about the person mattered more than how big the model was: the memory argument of Part IV, made by a homepage.
For most companies this rung is less exotic than it sounds. At Larkspur Systems (the fictional field-service software company used throughout this handbook), the renewal page an account manager sends is composed from six approved modules. Which four appear, in which order, and what their zones say depends on the account's memory: offline sync for accounts whose tickets mention rural sites, the route-sync fix for the roughly 60 accounts that hit the defect. The frame, the pricing table, and the contract terms are the same for everyone.
Rung 3: interfaces generated per person
The third rung generates the interface itself: layout, components, and interaction, composed for this person or request, with no designer having drawn it.
It exists. In November 2025, Google shipped an experiment in the Gemini app ("dynamic view") that designs and codes a custom interface in real time for a prompt [11]. Google's paper found generated interfaces "overwhelmingly preferred" to text-only answers and at least comparable to expert-made websites in about half of comparisons, but pages built by human experts still ranked first, and generation time was excluded [9]. Stanford researchers reported "up to" a 72% improvement in human preference over chat, a best case rather than an average [12].
For builders, the protocols matter more than the demos, and they point one way. A2UI, an open Google project, has the agent send a declarative JSON description, "not executable code," which the client renders from its own catalog of trusted, pre-approved components [14]. MCP Apps, co-developed by Anthropic, OpenAI, and the MCP-UI maintainers, lets a tool return a pre-declared interface template rendered in a sandboxed frame, with user consent for actions the interface initiates; it became an official extension in January 2026 [15], and OpenAI's apps in ChatGPT build on the same foundation [16]. The precedent is Microsoft's Adaptive Cards: "Purely declarative: No code is needed or allowed," with the host owning the look and feel [17]. I state the convergence as a rule: catalog, not code. The model chooses components and fills their fields. The host owns the components, the rendering, and the permissions.
The limits are specific.
Latency. Google's researchers note generation can take a minute or more [10]. Chapter 17 put an in-product interaction at about 0.1 seconds to feel instant, and Booking.com measured that about 30% more latency cost more than 0.5% in conversion [19]. A screen that arrives in 40 seconds is a different product, not a slower one.
Learnability. Nielsen Norman Group, broadly sympathetic to generative UI, names the cost: users rely on stable conventions and must relearn an interface that changes [13]. The chapter What Designers Knew traces the adaptive-interface research behind that warning.
Accessibility. Models routinely omit alt text, landmarks, and form labels. One team cut the inaccessibility rate of generated interfaces by 87.5% and 58.1% under two auditors, but only by training for it [18]. Accessible is engineered, not default.
Testing. You can QA a template. You cannot QA in advance an interface that did not exist until the person asked. Quality moves to the catalog, the renderer, and samples of what was actually shown (Chapter 20).
My labels, as of September 2026: interactive interfaces inside AI assistants ship now; generated interfaces in consumer apps ship as experiments; fully generated pages per visitor on a company's own site, at web latency, are a lab demo (the Closing gives the forecast). Vercel's 2024 SDK for streaming generated React components has since had that line of development paused [31]: the first architecture for a new capability is rarely the lasting one.
The notification is a decision, not a message
A notification is the most personal surface most companies own and the one they treat most carelessly. It interrupts, in a space the person did not open, and the person can revoke it for good. When Android made notifications a runtime opt-in in 2022, one push provider's data across more than 16 million devices found gaming apps lost nearly a third of opted-in users over two years and news apps 19%, while finance and transport lost least [28]. Vendor data, but the direction is clear: people close a channel that is spent carelessly.
So I treat a notification as four decisions, in order, with a legitimate "no" at each.
Figure 16.2. The notification decision: IF, WHAT, WHEN, HOW, with "not at all" scored like any other option and every opt-out counted as a cost.
IF: should this person be interrupted at all? LinkedIn's team wrote that sending on every eligible event "would result in too many notifications which would in turn annoy members," and built a near-real-time system that decides per member under a volume constraint [22]. Pinterest's system sets each user's weekly volume to optimize long-term engagement; it reduced volume while improving notification click-through and site engagement [23]. Both teams arrived at Chapter 14's do-nothing option, applied to interruption.
WHAT: which message? Duolingo's reminder system is the best-documented answer. Yancey and Settles treated the choice among hand-written templates as a bandit with two twists: some templates are ineligible for some people at some times ("sleeping" arms), and a template's effect recovers only after it rests, because novelty wears off. Against a random-template system that years of A/B tests had already made strong, a two-week experiment raised daily active users 0.5%, lessons 0.4%, and new users' recurring retention 2% [20]. Small numbers, and real ones. They came with a rule: the team could tune timing, templates, images, and localization but "could not increase the quantity of notifications without strong justification and CEO approval" [21]. It won by choosing better, not by sending more.
WHEN: at a time the person can act. A small 2014 study logged an average of 63.5 notifications a day per person [24]. In a within-subject experiment with 221 people, a week with alerts on raised inattention and hyperactivity symptoms compared with a week with them off [25]. People who went a day without push notifications were less distracted and more productive, and also more anxious about missing something [26]. The person is already managing a flood; sometimes the right "when" is tomorrow's digest.
HOW: which channel, and how loudly. Since iOS 15, each notification carries an interruption level: passive (delivered silently), active (the default), time sensitive (can break through Focus and the scheduled summary, if the person allows it), and critical (bypasses the ringer, by entitlement only). Apple's guidance is blunt: "Do not overuse their interruptive nature," and people can turn off time-sensitive alerts or all of an app's notifications [27]. The channel (push, in-app, email, or the account manager instead) is arbitrated against the same per-person budget as every other program (Chapter 14; Playbook P9).
Two consequences follow from the platform in the middle. Write notifications that survive summarization: fact first, self-contained, entity and verb in the same sentence, no meaning carried by tone or by sequence with other alerts. The platform can rewrite your words; it cannot take your name off them. Treat permission as the budget: the balance that matters is how many people still let you interrupt them, and every disable or revoked permission is a withdrawal you cannot easily reverse.
In the engine I built, notifications are not yet a shipped part of the personalization product; the design I am pursuing applies these four decisions to seller and operator alerts, drawing on the same account memory and governance as the pages and emails.
Hold the frame still
The Product playbook applies this to in-product slots. The general form, for any rung:
The frame is fixed. Navigation, control locations, pricing, legal terms, and every control for cancelling, downgrading, consent, or turning something off stay out of the adaptive layer. They are the map.
Change happens at predictable moments. A part may update when a page loads or a session begins, never mid-task. Personalization that swaps a block after first paint moves things under the person's cursor (Chapter 17).
The person can see why and undo it. A generated module can carry a short, honest reason ("Because your team works rural sites") and a way to hide or reset it.
Identity is stable. Netflix's one-image-per-title rule [5] exists so a person recognizes yesterday's title today. Products, prices, and people need the same stability.
What was shown is logged, with the decision and the template, catalog, and guideline versions. Support cannot help with a screen nobody can reproduce.
Generate the parts; govern the frame.
Rendering is where content becomes a security boundary
A generated sentence in an email can be false. A generated value that reaches a renderer can also be dangerous: a script, a link to somewhere it should not go, markup that breaks the page. Treat every model output as untrusted input at the point where it becomes an experience.
Figure 16.3. The render gate: generated output is untrusted input, and anything that fails a check becomes the approved fallback, never a repaired fragment.
In the engine I built, this runs before any content reaches a rendered page: each value is coerced to its zone's type, HTML is escaped, URLs are parsed, only approved schemes pass, and a failure renders the zone's safe fallback. A javascript: or data: link is refused however confidently the model returned it. The protocols build the same idea into the platform: A2UI sends data to a catalog the client owns [14]; MCP Apps renders pre-declared templates in a sandbox and lets hosts block suspicious content before render [15].
Then review the rendered surface, not only the stored data. Fields can be correct while the live page is wrong: a stale template or cache, the wrong page identity, a hidden zone, malformed markup. In my engine, inspecting the actual page is part of the operator's QA workflow for that reason (Chapter 20).
The strictest version of this boundary is paper, which cannot be patched after it is mailed. The direction I am taking the engine, not yet a shipped product, uses the same governed context for print: an approved template with declared zones, a deterministic preflight stricter than for web copy, and a QR code that continues into a personalized page whose visit becomes a signal for the next decision (Playbook M3). The cost of an unrecallable error sets how strict the gate must be.
Measure the experience, not the message
Chapter 21's tools apply, with three adjustments.
The outcome is the task, not the click. A page or screen succeeds when the person finishes what they came for: a form completed, a route dispatched, a renewal signed. A novel generated module attracts clicks for a while whether or not it helps. Measure task completion and time on task, and watch experienced users, who pay first when the frame moves.
Effects are small and variants many. GenPage's +0.24% [8] and Duolingo's +0.5% [20] were detectable because the populations were enormous. Test a few strong hypotheses about the frame and the catalog, let a bandit choose within slots, and keep a permanent holdout on the static experience. Do not trust offline wins: Booking.com found essentially no correlation between offline model gains and business gains across 23 comparisons [19].
Engagement is the constraint. In a two-year randomized trial in 18 Tennessee middle schools, 96% of students tried an AI tutor, but the median student used it in only 17% of practice sessions where they made a mistake; "the binding constraint appears to be engagement" [29]. A personal experience nobody uses has no effect. Measure use before lift.
For notifications, count the costs the channel hides: disables, quiet deliveries, revoked permissions, and unsubscribes, per program. A program that lifts this week's sessions while raising the disable rate is borrowing from next quarter, the ads-blindness shape from Chapter 21.
Cost and latency: generate ahead, select live
Experiences are visited more than once, and someone is usually waiting. Take an illustrative known lead who opens a personal page four times. Generating per visit costs four generations, each on the waiting path. Generating once when the lead is accepted costs one, off the waiting path, and turns each visit into a fetch. That is how the known-lead page runs in the engine I built: the submission is acknowledged at once, the page is generated into zones in the background, validated, and ready before the person clicks.
The same logic prices the rungs. Selection inside zones is nearly free per view. A catalog-composed page costs one generation per person per meaningful change in their record. A generated interface costs a generation per request, plus the latency and the rendering risk. Notifications are the cheapest to generate and the most expensive to get wrong, because the cost is paid in attention and in permissions you cannot buy back.
What this does not do
It does not show that whole-page generation beats sections. I found no public controlled comparison, on a commercial website, of zones in a fixed layout against generating the whole page. GenPage is the closest, and it generates a constrained grid [8]. My preference for zones inside a governed frame is a design judgment from building these systems and from the usability evidence, not a measured result.
It does not make generated interfaces better than designed ones. You will hear that users prefer generative UI. Google's evaluation, the most careful public one, found generated interfaces beat text, and expert-built sites still ranked first, with speed excluded [9]. The "72%" is a best case [12]. I share NN/g's direction toward interfaces shaped around outcomes [13]; I do not accept the stronger claim that every page will soon be generated fresh for every visitor.
It does not fix a bad default or a bad volume. If most users need the same adaptation, change the default for everyone. And a clever notification sent too often is still too often; Duolingo's gains came inside a volume its CEO guarded [21].
The evidence on onboarding is thin. I found no rigorous public study of LLM-generated or per-person onboarding flows. Treat large vendor-claimed activation lifts as unverified until your own holdout confirms them.
At scale
At a few hundred customers, a designer can look at every personalized page. At 250,000 contacts the questions become statistical: what share of renders ran full, degraded, fallback, or empty; which zone falls back most and which data source is behind it; how many values the render gate refused, and why. A sudden rise in fallbacks usually means an upstream source failed; a sudden fall usually means a check stopped running.
Scale also brings the platform in. At millions of notifications, operating-system models summarize, group, and demote your content, and the signal you get back is partly theirs. Log what you sent, at which interruption level, and what you intended, so a platform change can be told apart from yours. And build accessibility checks into the render gate, not a quarterly audit: at scale nobody reviews a render before the person sees it.
Failure story: Borrowed Urgency
Larkspur Systems is a fictional composite company; the numbers here are illustrative.
Larkspur's mobile app launched "smart alerts" for dispatchers. A model wrote each alert from the job's context, and a well-meant rule marked any alert about a customer job as time sensitive, so it would break through Focus and the scheduled summary. In testing, alerts were read faster. In the first month, the median dispatcher received about fourteen time-sensitive alerts a day, most of the "technician running ten minutes late" kind.
Dispatchers did what the operating system lets them do in one step: they turned off time-sensitive alerts for Larkspur, and many moved the app into the scheduled summary.
Five weeks later, the route-sync defect recurred overnight for a group of accounts. The alert that mattered, "Re-sync routes before 7 a.m. or technicians will see yesterday's jobs," was generated correctly, grounded correctly, and marked time sensitive. It arrived at 8:00 a.m., inside a summary, under three late-technician notices.
No single alert was wrong. The system had spent an interruption level it did not own, and the people it interrupted took it back. The fix was a per-dispatcher interruption budget, time sensitive reserved for a short approved list (a sync failure, a safety issue, a job cancelled within the hour), and "not at all" in every alert decision. The anti-pattern: borrowed urgency, using the loudest channel for ordinary news until the person turns the volume down for everything.
Patterns
Catalog, Not Code. Problem: a model that generates interface markup can produce anything, including unsafe, inaccessible, or off-brand screens. Forces: generation adds real value in composition; the host must stay accountable for what renders. Solution: the model selects components from a governed catalog and fills typed fields; the host owns rendering, styling, accessibility, and permissions; every output passes a render gate with a fallback. Tradeoffs: expressiveness is limited to the catalog, which must be designed, versioned, and extended deliberately.
The Four-Question Notification. Problem: notifications are written as messages and sent whenever an event fires. Forces: interruptions have value and cost; permission is revocable; platforms summarize and filter. Solution: decide IF, WHAT, WHEN, and HOW in order, with "not at all" scored at every step, a per-person interruption budget shared across programs, a recency penalty on templates, and the loudest levels reserved for an approved list. Tradeoffs: fewer sends and slower-looking short-term metrics; requires counting disables as costs.
Rendered-Variant Log. Problem: per-person experiences cannot be reproduced, debugged, or measured without knowing what each person saw. Forces: storage cost; privacy of logged content. Solution: log, for every render, the variant, the deciding policy, the template, catalog and guideline versions, the tier (full, degraded, fallback, empty), and any refusals by the render gate, under the same retention rules as the memory it came from. Tradeoffs: log volume; requires discipline about what is stored and for how long (Chapter 19).
Leader questions
- When a page or screen changes for one customer, what stays exactly the same for everyone, and who owns that list?
- Which rung of the ladder are we on for each surface, and what evidence says we should climb?
- Does any model output reach a renderer without type coercion, escaping, URL checks, and a fallback?
- How many people disabled or quieted our notifications last quarter, per program, and who counts that as a cost?
- Can support see exactly what a given customer was shown yesterday, and why?
Build checklist
- Every personalized surface has a written frame: the parts that never vary, including all cancel, downgrade, consent, and turn-off controls.
- Every zone or slot has a contract (Chapter 15) and an approved fallback; selection is tried before generation.
- Generated interfaces use a governed component catalog; no model output is rendered as code.
- A render gate coerces, escapes, checks URLs and schemes, confirms catalog membership, and falls back on failure.
- Accessibility checks run on rendered output, not only on templates.
- Per-person pages are generated ahead of the visit and served as a fetch; the live path only selects.
- Personal pages have persistent, campaign-scoped identity so one program cannot overwrite another.
- Notifications go through IF, WHAT, WHEN, HOW, with "not at all" logged as an outcome and a per-person budget shared across programs.
- Time-sensitive and critical levels are reserved for an approved event list.
- Every render and every notification decision is logged with versions and tier, and the rendered surface is reviewed, not only stored data.
Metrics to watch
- Task completion and time on task for the person's job, with experienced users tracked separately.
- Tier distribution per surface: full, degraded, fallback, empty, and render-gate refusals.
- Notification permission health: disable, quiet-delivery, and revocation rates per program, against a holdout.
- Do-nothing share of notification decisions, watched for sudden drops.
- Accessibility defect rate on sampled rendered output.
Reader Q&A
Should we start with generative UI in our product? Start with zones and slots. Generated interfaces are worth prototyping for exploratory tasks where people are already waiting for an assistant (analysis, planning, configuration), not for the screen someone uses fifty times a day.
Can we let the model decide when to send a notification? Let a model score the value and the timing; keep the budget, quiet hours, approved interruption levels, and the do-nothing option in deterministic policy. The model proposes; policy decides.
What about summaries the platform writes of our notifications? You cannot control them, so write for them: fact first, self-contained, unambiguous. If a notification is safety- or money-critical, say it in a way that cannot be summarized into its opposite, and consider a second channel.
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 16
concepts:
- name: Generative Experience Ladder
definition: "Rungs of composing an experience per person: zones inside a governed page; whole pages and product surfaces from a constrained vocabulary; interfaces generated per person or request. Notifications sit beside the ladder as a decision."
- name: Stable Frame, Generated Parts
definition: "Navigation, controls, pricing, legal terms, and all cancel, downgrade, consent, and turn-off controls stay fixed; only declared parts vary, at predictable moments, with explanation and undo."
- name: Catalog, Not Code
definition: "Models select and fill components from a host-owned, governed catalog; the host renders, styles, and permissions them. No model output is rendered as code."
- name: Render Gate
definition: "Coerce to type, escape, parse URLs and allow approved schemes, confirm catalog membership, validate against the zone contract; on failure render the approved fallback, then nothing."
- name: Notification Decision
definition: "IF, WHAT, WHEN, HOW, decided in order, with not at all scored as an option and a per-person interruption budget shared across programs."
- name: Permission as Budget
definition: "The notification channel's real balance is how many people still allow interruption; disables and revocations are costs."
- name: Summary-Proof Notification
definition: "A notification written so it stays true when a platform model bundles or condenses it: fact first, self-contained, unambiguous."
decision_rules:
- if: "a zone can be served nearly as well by selecting an approved variant"
then: "select; generate only where no approved variant fits the person's situation"
- if: "a proposed adaptation would move navigation, a control, pricing, legal terms, or a cancel or consent control"
then: "do not adapt it; keep it in the fixed frame"
- if: "any model output is about to be rendered"
then: "pass it through the render gate; on failure render the approved fallback, never a repaired fragment"
- if: "an interface is generated for a surface with an interactive latency budget (about 0.1 s)"
then: "do not generate at request time; precompute or select from a catalog"
- if: "a per-person page will be visited more than once"
then: "generate it once ahead of the visit, store it with campaign-scoped identity, and serve each visit as a fetch"
- if: "an event could trigger a notification"
then: "decide IF, WHAT, WHEN, HOW in order; score not at all; check the person's interruption budget"
- if: "a notification would use a time-sensitive or critical interruption level"
then: "allow it only for events on an approved list"
- if: "notification disable or quiet-delivery rates rise for a program"
then: "treat it as a cost against that program's lift, and reduce volume before rewriting copy"
- if: "the surface is physical or otherwise unrecallable"
then: "apply stricter deterministic preflight and approval than for web copy"
assessment_questions:
- "Which surfaces are personalized today, and on which rung of the ladder is each?"
- "What is fixed in each surface's frame, and is that list written down?"
- "Does model output ever reach a renderer without coercion, escaping, URL checks, and a fallback?"
- "Are per-person pages generated ahead of the visit or on request?"
- "How are notification decisions made, and is not sending an explicit, logged outcome?"
- "What are the notification disable and revocation rates per program?"
- "Can support reproduce what a given person was shown and why?"
patterns: [Catalog Not Code, Four-Question Notification, Rendered-Variant Log, Stable Frame Generated Parts, Generate Ahead Select Live, Generation Contract, Do-Nothing Option]
anti_patterns: [Borrowed Urgency, The Shifting Map, Generated Frame, Sentence Surgery, Novelty as Lift, Summarizable Alert]
metrics: [task completion and time on task, tier distribution per surface, notification permission health, do-nothing share, accessibility defect rate]
links: {decision_layer: 14, generation: 15, latency: 17, governance: 18, privacy: 19, accuracy: 20, measurement: 21, product: PR, inbound: M1, physical: M3, coordination: P9, forecast: closing, adaptive_interfaces: 3}
maturity_dimension: experience_generationReferences
- The Register (2025-01-07). "Apple responds to BBC complaint about AI-generated notification summaries." The Register. https://www.theregister.com/2025/01/07/apple_responds_bbc_complaint/
- TechCrunch (2025-01-16). "Apple pauses AI notification summaries for news after generating false alerts." https://techcrunch.com/2025/01/16/apple-pauses-ai-notification-summaries-for-news-after-generating-false-alerts
- MacRumors (2025-07-22). "iOS 26 Beta 4 Reintroduces Notification Summaries for News Apps." https://www.macrumors.com/2025/07/22/ios-26-beta-4-notification-summaries/
- 9to5Google (2025-11-13). "Pixel AI Notification Summaries start rolling out." https://9to5google.com/2025/11/13/pixel-notification-summaries/
- Chandrashekar, A., Amat, F., Basilico, J., and Jebara, T. (2017-12). "Artwork Personalization at Netflix." Netflix Technology Blog. https://netflixtechblog.com/artwork-personalization-c589f074ad76
- Wang, Yin, Jie, Wang, Yamada, Chang, and Mei (2016). "Beyond Ranking: Optimizing Whole-Page Presentation." WSDM 2016, pp. 103-112. https://dl.acm.org/doi/10.1145/2835776.2835824
- Dynamic Yield developer documentation, "Choosing Variations" (Choose API), accessed 2026-09-28. Cited as a representative example of the slot model, not an endorsement. https://dy.dev/docs/choose
- Wang, L., Pan, J., and Baltrunas, L. (2026). "GenPage: Towards End-to-End Generative Homepage Construction at Netflix." arXiv:2606.31031; ACM RecSys 2026. https://arxiv.org/abs/2606.31031
- Leviathan, Valevski, Kalman, Lumen, Segalis, Molad, Pasternak, Natchu, Nygaard, Venkatachary, Manyika, and Matias (2025-11). "Generative UI: LLMs are Effective UI Generators." Google Research. https://generativeui.github.io/static/pdfs/paper.pdf
- Leviathan, Y., Valevski, D., Natchu, V., and Matias, Y. (2025-11-18). "Generative UI: A rich, custom, visual interactive user experience for any prompt." Google Research blog. https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/
- Woodward, J. (2025-11-18). "Gemini 3 brings upgraded smarts and new capabilities to the Gemini app." Google blog. https://blog.google/products-and-platforms/products/gemini/gemini-3-gemini-app/
- Chen, Zhang, Zhang, Shao, and Yang (Stanford SALT) (2025, rev. 2026). "Generative Interfaces for Language Models." arXiv:2508.19227; ACL 2026 Findings. https://arxiv.org/abs/2508.19227
- Moran, K., and Gibbons, S. (2024-03-22). "Generative UI and Outcome-Oriented Design." Nielsen Norman Group. https://www.nngroup.com/articles/generative-ui/
- Google A2UI Team (2025-12-15). "Introducing A2UI: An open project for agent-driven interfaces." Google Developers Blog. https://developers.googleblog.com/introducing-a2ui-an-open-project-for-agent-driven-interfaces/ ; project site https://a2ui.org/ (accessed 2026-09-28)
- Model Context Protocol blog (2025-11-21). "MCP Apps: Extending servers with interactive user interfaces." https://blog.modelcontextprotocol.io/posts/2025-11-21-mcp-apps/ ; and (2026-01-26) "MCP Apps: Bringing UI Capabilities to MCP Clients." https://blog.modelcontextprotocol.io/posts/2026-01-26-mcp-apps/
- OpenAI (2025-10-06). "Introducing apps in ChatGPT and the new Apps SDK." https://openai.com/index/introducing-apps-in-chatgpt/ (confirmed via secondary reports; page not fetched)
- Microsoft Learn. "Adaptive Cards Overview," updated 2025-12. https://learn.microsoft.com/en-us/adaptive-cards/
- Yoon et al. (2025, rev. 2026). "A11yn: Aligning LLMs for Web Accessibility-Aware UI Generation." arXiv:2510.13914. https://arxiv.org/abs/2510.13914
- Bernardi, L., Mavridis, T., and Estevez, P. (2019). "150 Successful Machine Learning Models: 6 Lessons Learned at Booking.com." KDD 2019. https://dl.acm.org/doi/10.1145/3292500.3330744
- Yancey, K. P., and Settles, B. (2020). "A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications." KDD 2020. https://research.duolingo.com/papers/yancey.kdd20.pdf
- Mazal, J. (2023-02-28). "How Duolingo reignited user growth." Lenny's Newsletter. https://www.lennysnewsletter.com/p/how-duolingo-reignited-user-growth
- Gao, Gupta, Yan, Shi, Tao, Xiao, Wang, Yu, Rosales, Muralidharan, and Chatterjee (LinkedIn) (2018). "Near Real-time Optimization of Activity-based Notifications." KDD 2018. https://dl.acm.org/doi/10.1145/3219819.3219880
- Zhao, Narita, Orten, and Egan (Pinterest) (2018). "Notification Volume Control and Optimization System at Pinterest." KDD 2018. https://dl.acm.org/doi/10.1145/3219819.3219906
- Pielot, M., Church, K., and de Oliveira, R. (2014). "An In-Situ Study of Mobile Phone Notifications." MobileHCI 2014. https://dl.acm.org/doi/10.1145/2628363.2628364
- Kushlev, K., Proulx, J., and Dunn, E. W. (2016). "'Silence Your Phones': Smartphone Notifications Increase Inattention and Hyperactivity Symptoms." CHI 2016. https://dl.acm.org/doi/10.1145/2858036.2858359
- Pielot, M., and Rello, L. (2017). "Productive, Anxious, Lonely: 24 Hours Without Push Notifications." MobileHCI 2017; arXiv:1612.02314. https://arxiv.org/abs/1612.02314
- Apple (2021). "Send communication and Time Sensitive notifications." WWDC21 session 10091. https://developer.apple.com/videos/play/wwdc2021/10091/
- Pushwoosh (2024-10-07). "Android 13 impact on opt-in rates." Vendor research. https://www.pushwoosh.com/blog/android-13-opt-in-rates-research/
- Oreopoulos and Low (2026-08). "One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment." EdWorkingPaper 26-1551, Annenberg Institute at Brown University. https://doi.org/10.26300/kner-hv33
- Apple Newsroom (2024-10-28). "Apple Intelligence is available today on iPhone, iPad, and Mac." https://www.apple.com/newsroom/2024/10/apple-intelligence-is-available-today-on-iphone-ipad-and-mac/
- Palmer, Ding, Leiter, Shadcn, Grammel, and Philemon (2024-03-01). "Introducing AI SDK 3.0 with Generative UI support." Vercel. https://vercel.com/blog/ai-sdk-3-generative-ui ; note current Vercel docs mark AI SDK RSC development as paused.
This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised