Questions this chapter answers
- We have forty AI personalization ideas. How do we pick three?
- How much is a personalization idea actually worth before we build it?
- Which channel should go first, and where does generated content need a human to approve it?
- What should we deliberately not do yet?
The short answer
Do not start from use cases. Start from decisions. A decision is something your company chooses again and again for one person at a time, with real options (including doing nothing) and an outcome you can measure: which at-risk account gets which kind of help, what a new inbound lead hears first, whether a customer sees an expansion offer at all.
Price each decision before you build anything. Value at stake is the number of times the decision is made in a year, times the value of making it better once, times the lift you can realistically achieve over what you do today, minus the cost of running it and the risk of getting it wrong. Every term is a range, not a point. Whatever range you write down, rank on its low end, because a decision that only pays at the top of its range is a bet, not a first decision.
Then score the survivors on three things value hides: whether the data exists, whether a wrong decision can be undone, and how long it takes to learn whether it worked. Keep those scores separate. Collapsing them into one "feasibility" number is how companies end up funding whatever is easiest.
Choose three: a fast win, with ready data and evidence in weeks; a memory-building bet, high in value, whose preparation builds the customer memory every later decision reuses; and a measurement foundation, the holdouts and decision logs that prove any of it. Human-assisted channels go first for high-value decisions. Cold, paid, and irreversible channels wait.
The most common mistake is not choosing the wrong technology. It is personalizing the touchpoint that was easiest to automate, which is usually the one that mattered least.
Larkspur Systems (a fictional company, a composite of many real ones) holds its leadership offsite in a hotel meeting room with a whiteboard that fills by eleven. Forty sticky notes, each labeled "AI personalization." AI-written newsletter intros. A website hero that changes by industry. An AI SDR for every prospect in the market. Renewal-risk alerts. A support bot that knows the account. Birthday emails for fleet managers. Personalized event agendas. A pricing page that adapts.
The CMO wants the newsletter: it goes to 250,000 contacts, the tooling exists, it could ship in two weeks. The CRO wants the AI SDR expanded, even though the pilot is making reply rates worse, because the vendor says the next version fixes it. The head of customer success says nothing until the end, then points at one note: renewal risk. "I have six people rebuilding account history by hand before every save call."
Nobody is wrong about their own note. The room simply has no way to compare them.
A use case is a wish; a decision is a unit you can price
"Personalized website" cannot be estimated, funded, or tested. It is a wish. Hidden inside it are several decisions: which of three proof points a returning visitor from a renewing account sees; whether a visitor from an open support escalation sees a sales banner at all; which case study a prospect from a fleet of 500 trucks gets instead of a fleet of five.
A decision qualifies for this chapter when it has three properties:
- It repeats. It is made hundreds or thousands of times a year, each time for one person or account. A one-off strategic choice is not a personalization decision.
- It has options. At least two things could happen, and one of them is usually "nothing." If there is only one possible action, there is nothing to personalize.
- It has a measurable outcome. Renewed or churned, replied or not, resolved in one contact or three. If you cannot name the outcome, you cannot know whether personalizing it helped.
Figure 6.1. Break every use case into the repeated decisions hiding inside it; only those can be priced and tested.
Name the objective each decision serves, for which person: acquire, convert, expand, retain, reactivate, serve, or reduce the cost to serve. "Increase engagement" is not on the list; engagement is a measurement, not an outcome anyone pays for.
The decision inventory is a table with one row per decision: the person, the moment, the options, the outcome, the annual volume, and the owner who would act on a better answer. Larkspur's forty sticky notes collapsed into eleven rows once duplicates merged and wishes were translated; three had no outcome anyone could name and were dropped. Writing the inventory is most of the work: it forces each idea to say what would be chosen differently, for whom, and how you would know.
Write the options with their depth, not only "something or nothing." In the governed personalization engine I built, each campaign declares its personalization mode before anything runs: company-level only, personalized to the named person, or a score and brief for the rep with no customer-facing copy at all. Quality review then judges each record against the declared mode, so a company-level campaign is never faulted for not naming the individual. The row should also say what happens when the evidence for one record is thin: fall back to the account, then the segment, then approved static copy (Chapter 20). A decision whose only options are "fully personalized" and "nothing" fails on every record the data cannot support.
Value at stake is volume times the value of changing the decision
Every growth team already runs a version of this arithmetic. You estimate what a good outcome is worth. You estimate how likely a particular approach is to produce it. You multiply, subtract what the approach costs to run, and spend where the number is highest, stopping where it turns negative. The personalization version adds one discipline: price the change, not the customer.
Value at stake = annual decision volume × value per improved decision × achievable lift, minus cost to run, minus risk.
Three of those terms are routinely misread.
Value per improved decision is not the value of the customer. It is the value of the outcome that changes because the decision changed. A retention offer to a customer who would have renewed anyway is worth nothing, however large the account. Eva Ascarza's field experiments showed that targeting the customers at highest risk of churning can be ineffective, and that the right target is the customers whose behavior the intervention actually changes [1]. Uplift modelers have a name for the worst case, "sleeping dogs": customers a retention contact pushes toward leaving [2]. So the value term is about the margin. It is not "who is this account worth," but "what is this decision worth when we get it right instead of wrong."
Lift is measured against what you do today, not against nothing. Larkspur already sends every at-risk account a generic check-in. A personalized save play has to beat that, not beat silence. And lift should be written in absolute points. A 20% relative lift on a 2% click rate is 0.4 points, which sounds very different in a budget meeting.
Achievable lift is smaller than hoped. Ron Kohavi, who ran experimentation at Microsoft, reported that only about 10 to 20% of ideas tested at Bing improved the metric they targeted, and that most winning changes moved it by 0.1% to 1% [3]. Personalization ideas are ideas. Plan on the low end.
Here is Larkspur's inventory, priced. Every number is illustrative for a fictional company; the ranges and the reasoning are the point.
| Decision | Annual volume | Value per improved decision | Achievable lift (absolute) | Gross value at stake |
|---|---|---|---|---|
| Renewal save: what each at-risk account gets (CSM call, tailored plan, or nothing) | 6,000 at-risk accounts | $6,000 to $15,000 (one retained contract) | 1 to 4 points | $360K to $3.6M |
| Inbound response: who answers a demo request, how fast, with what | 3,000 requests | $6,000 to $15,000 (first-year contract) | 1 to 3 points of close rate | $180K to $1.35M |
| Expansion: whether and how to offer the fleet-tracking add-on | 40,000 accounts | $1,500 to $4,000 per year | 0.2 to 1 point | $120K to $1.6M |
| Support context: next step on a ticket, with account history | 60,000 tickets | $5 to $20 in handling cost | 10 to 30% of tickets handled differently | $30K to $360K |
| AI SDR first touch to new prospects | 100,000 emails | $300 to $1,500 per meeting | 0.1 to 0.5 points (could be negative) | $30K to $750K, before reputation risk |
| Newsletter intros, per contact | 1,000,000 sends | $0.50 to $5 per extra click | 0.1 to 0.5 points | $500 to $25K |
Two things jump out. The ranges are wide, often tenfold, and that is honest: the lift assumption carries most of the weight, and a defensible model shows it instead of hiding it inside one number. And the newsletter, the easiest idea on the whiteboard, sits one to three orders of magnitude below everything else. At a fraction of a cent to two cents of generation per contact (an assumption), its cost could exceed its value.
Figure 6.2. Ranked on the low end of each range, the easiest idea to automate is worth the least.
Here is the sentence this chapter turns on: rank decisions by the value of changing them, not by the value of the customer and not by the ease of automating them.
Four scores, kept separate
Value decides what deserves attention. Three other properties decide what goes first.
Data readiness. Can you make this decision well with what you have today? Larkspur's renewal decision needs support history, usage telemetry, and call notes, which exist in three unjoined systems. The inbound decision needs a form, a web session, and a CRM record, all ready. Readiness is not a reason to skip a high-value decision; it tells you how much memory work sits between you and the first result.
Reversibility. What happens when the decision is wrong? An in-app suggestion can be withdrawn tomorrow. An email cannot be unsent. A quoted price can become an obligation. A cold email that lands in spam damages a sending domain every other team relies on. The less reversible the action, the more governance it needs before it runs unattended (Chapter 18), and the worse it is as a first decision.
Time to evidence. How long until you can tell whether it worked? This is arithmetic, not intuition. To detect a click rate moving from 2.0% to 2.4% with a standard two-sided test (5% significance, 80% power) you need about 21,000 people in each arm. The newsletter has that in one send. To detect a renewal rate moving from 85% to 88% you need about 2,000 accounts per arm, and the outcome arrives only at the renewal date, months after the decision. Chapter 21 covers the methods; for choosing, the point is that measurability is a cost you pay, not a value you get.
One kind of evidence does not wait for a test. Whether the output is correct and appropriate shows up in the first few dozen real records. In the engine I built, an operator reviews the first cohorts of a new campaign (the first dozens, then the first hundred, then the first few hundred), with AI reasoning across the whole population for repeated errors, weak outputs, and fallbacks that fire too often (Chapter 22). Plan that review for every first decision, and do not mistake it for lift.
Here I part ways, partly, with the most quoted enterprise AI report of 2025. MIT's Project NANDA found that about half of generative AI budgets in its sample went to sales and marketing while the clearest returns came from the back office, and said the allocation "reflects easier metric attribution, not actual value" [4]. The diagnosis is useful, within the report's own limits: 52 interviews, 153 surveyed leaders, and a six-month window it concedes may understate success [4]. The lesson I take is not "leave the front office." It is that when time to evidence stands in for value, the easy-to-measure decision always wins the budget.
The same objection applies to value-versus-feasibility grids such as Gartner's "Prism" [5]. "Feasibility" blends readiness, reversibility, and time to evidence into one axis, so a decision that is ready, quick to measure, and irreversible scores as feasible while being the most dangerous thing on the board.
The channel sets how much the machine may say on its own
The channel decides how much generated content can reach a person without review.
- Human-assisted (a rep, CSM, or support agent receives a researched brief and a draft, and decides what to send): the highest tolerance for generation, because a person who knows the account reads it first. High-value, low-readiness decisions start here, and every human edit is evidence of what the system got wrong.
- Owned and automated (email, in-app, website, support replies): generation is fine inside typed zones with grounding checks (Chapter 15). Anything that states a price, a contract term, a commitment, or a regulated claim should require approval or come from a fixed template. So should any statement about the customer's own relationship with you: "your renewal is coming up" or "you already use the fleet module" appears only when memory confirms it; otherwise the copy speaks of an opportunity. In the engine I built, that rule lives in the guidelines and in a code check behind them (Chapter 20).
- Paid (ads and sponsored placements): you control the creative; the platform controls much of the targeting and delivery, under its own policies. Per-person claims have little room here; it is usually a later step.
- Cold outbound to people with no relationship: the least tolerance of all. The recipient has no context for why you know what you know, and the reputation cost is shared by the whole company.
That is why Larkspur's AI SDR pilot is hurting reply rates. The writing is not the problem. The most skeptical channel was chosen first, for the decision with the least memory behind it.
The first three: a fast win, a memory-building bet, a measurement foundation
One first decision is a single point of failure. Forty is a portfolio nobody can staff. Three works, if each plays a different role.
The fast win has ready data, reversible actions, and evidence within weeks. At Larkspur it is inbound response: an agent researches each demo request's account and drafts a brief and first reply; a rep approves and sends. Measurable in a quarter, and visible to the sales team. Known identity is most of what makes it ready. In the engine I built, the flow that runs in production is the known-lead flow: a form or a source says who the person is, and one research pass feeds a score, a brief for the rep, a personalized page, and an email sequence. Personalizing to an anonymous visitor before the form is still pilot design there. Start where identity is handed to you.
The memory-building bet has high value and low readiness, and the work of making it ready builds something reusable. At Larkspur it is renewal save. Joining support history, telemetry, and call notes into one account record is slow, but once it exists, expansion, support context, and every later playbook read from the same memory (Chapter 10). The first decision pays for infrastructure the next five use for free. The CSMs send; the system prepares.
The measurement foundation is less a decision than the plumbing that makes every decision provable: a randomized holdout on each new decision, a small permanent holdout across the whole program, and a log of what was known, decided, and done for each person. If a decision must carry it, pair it with expansion, where 40,000 accounts give enough volume to learn within a quarter and "no offer" is a real option to test.
What Larkspur chose not to do yet matters as much. The AI SDR pilot paused until account memory exists and the decision can be made for fewer, better-researched prospects. The newsletter kept its merge fields. The adaptive pricing page stayed on the whiteboard: high stakes, irreversible, and measurable only slowly.
Figure 6.3. Three roles, not three favorites: a fast win, a memory-building bet, a measurement foundation, and a written "not yet" list.
Choosing three is not the same as saying the other thirty-seven are bad. It says they are later.
What this does not do
The value-at-stake calculation is a way to rank, not a forecast. It is only as good as the lift assumption, and at the start nobody has a measured lift for their own decisions. Treat it as a disciplined argument about which assumption matters, and replace it with measured results as soon as you have any.
Do not import industry averages as your lift. McKinsey's widely quoted finding that personalization "most often drives" a 10 to 15% revenue lift, with a company range of 5 to 25%, describes company-level results in its research, not the lift of any single decision [6]. Gartner's 2019 prediction that 80% of marketers would abandon personalization by 2025 was a prediction, cited far more often than it was ever checked [7]. Neither belongs in your spreadsheet.
Personalization is not a fix for a bad default. If the generic onboarding is broken for everyone, fixing it for everyone will usually beat personalizing it for some. The inventory sometimes reveals that the highest-value decision is not a personalization decision at all.
And some decisions should not be made by a system at all, whatever they score (Chapter 19).
At scale
With one team, the inventory is a spreadsheet. With seven, it is a contested asset: marketing, sales, success, and support all want the same person's attention in the same week, so the inventory must record which decisions compete, and contact policy becomes a decision in its own right (Chapter 14 and the One Customer, One Conversation playbook).
Value per decision also falls as you work down the list, while the cost to govern each one does not. What keeps a program alive at scale is not better lift per decision but lower marginal cost per decision, from shared memory, governance, and measurement. That is why the memory-building bet belongs in the first three, not the fifth.
Failure story: The Streetlight Pilot
A composite of several companies. The team personalized the newsletter because it was easiest to automate: a contact list, an owned channel, a vendor who could ship in a sprint. Open rates rose, and the dashboard celebrated, though some of that rise was not people at all: since 2021, Apple's Mail Privacy Protection preloads content for users who enable it, which registers as an open whether or not anyone reads the email [8]. Clicks moved slightly. Pipeline did not. Meanwhile six CSMs kept rebuilding account history by hand before every save call, on a decision worth far more.
Nothing failed technically. The pilot looked under the streetlight because that is where the light was.
Patterns
Decision Inventory. Problem: ideas cannot be compared, because each team frames its own as a use case. Solution: translate each idea into repeated decisions with person, moment, options (including nothing), outcome, volume, and owner; drop anything without a nameable outcome. Trade-off: it slows the first meeting down, and it kills some favorite ideas.
Low-End Ranking. Problem: value estimates are wide, optimistic, and politically loaded. Solution: estimate every term as a range, rank on the low end, and show which assumption carries the weight. Trade-off: genuinely high-upside bets rank lower; take them deliberately, as bets.
First-Three Portfolio. Problem: one pilot is fragile; many are unstaffable. Forces: leaders need a visible result, infrastructure needs time, and nothing counts without proof. Solution: one fast win, one memory-building bet, one measurement foundation. Trade-off: the bet shows little for a quarter and needs a sponsor willing to wait.
Leader questions
- For each idea we fund, what decision changes, for whom, and how many times a year?
- What lift over today's treatment are we assuming, and is that the assumption carrying the estimate?
- If this decision is wrong, can we undo it, and how many weeks until we know it worked?
- Which of our first three builds something the next five decisions will reuse?
Build checklist
- Every idea translated into a decision with person, moment, options (including nothing), outcome, volume, owner.
- Personalization depth declared per decision, with the fallback when a record's evidence is thin.
- Value at stake estimated as ranges, with lift in absolute points over today's treatment.
- Cost per decision estimated: compute, human review minutes, governance.
- Readiness, reversibility, and time to evidence scored separately, never merged.
- Channel chosen by generation tolerance; approval rules written for prices, terms, and commitments.
- Sample size, holdout, and decision log designed before the first send.
- A written "not yet" list, with the reason for each item.
Metrics to watch
- Measured lift per decision against a holdout, replacing the assumed range as soon as data exists.
- Cost per decision, including human review time.
- Time to evidence: weeks from launch to a readable result, per decision.
- Reuse: how many decisions read from the memory the first bet built.
Reader Q&A
We are a small team. Is three too many? Then keep the roles and shrink the scope. A fast win and a measurement foundation can be the same decision if it has volume. Do not drop measurement to make room.
Should the first decision be customer-facing? Not necessarily. A human-assisted decision, where the system prepares and a person acts, often has the best ratio of value to risk.
What if the highest-value decision has the worst data? That is the memory-building bet. Start it now, in a human-assisted channel, and let the fast win carry the visible results while the memory fills.
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 6
concepts:
- name: Decision
definition: "A choice made repeatedly for one person or account, with at least two options (usually including doing nothing) and a measurable outcome."
- name: Decision Inventory
definition: "One row per decision: person, moment, options, outcome, annual volume, owner, objective served."
- name: Value at Stake
definition: "Annual volume x value per improved decision x achievable lift over current treatment, minus cost to run, minus risk; every term a range."
- name: Value per Improved Decision
definition: "The value of the outcome that changes because the decision changed; not the value of the customer."
- name: Time to Evidence
definition: "Time until a test can detect the target lift, set by sample size per arm and outcome delay."
- name: First-Three Portfolio
definition: "One fast win, one memory-building bet, one measurement foundation."
- name: Personalization Mode
definition: "The declared depth of a decision (company-level, person-level, or rep-only brief and score), against which each record is judged."
decision_rules:
- if: "an idea has no nameable, measurable outcome"
then: "drop it from the inventory until one is named"
- if: "a decision is attractive only at the top of its value range"
then: "treat it as a deliberate bet, not a first decision"
- if: "lift is stated in relative terms"
then: "convert to absolute points over current treatment before ranking"
- if: "an action is irreversible or states prices, terms, or commitments"
then: "require approval or a fixed template; do not choose it as a first unattended decision"
- if: "a high-value decision has low data readiness"
then: "run it human-assisted as the memory-building bet"
- if: "a record's evidence is too thin for the declared depth"
then: "fall back to account, then segment, then approved static copy; do not guess"
- if: "copy states the customer's relationship with you (renewal, product owned, contract)"
then: "allow it only when memory confirms it; otherwise frame it as an opportunity"
- if: "the generic experience is broken for everyone"
then: "fix the default before personalizing it"
assessment_questions:
- "List the repeated decisions you make about individual customers. Which include doing nothing as an option?"
- "For your top three, what is annual volume, what happens today, and what absolute lift would matter?"
- "Which decisions produce actions you cannot undo?"
- "How many people per test arm could you assign to each decision in a quarter, and when does the outcome arrive?"
- "Which data sources would the highest-value decision need, and are they joined today?"
patterns: [Decision Inventory, Low-End Ranking, First-Three Portfolio]
anti_patterns: [The Streetlight Pilot, Feasibility Blending, Customer Value Mistaken for Decision Value]
maturity_dimension: decisioningReferences
- Ascarza, E. (2018). "Retention Futility: Targeting High-Risk Customers Might Be Ineffective." Journal of Marketing Research 55(1). https://journals.sagepub.com/doi/10.1509/jmr.16.0163
- Radcliffe, N. J., and Simpson, R. (2008). "Identifying who can be saved and who will be driven away by retention activity." Journal of Telecommunications Management 1(2). https://stochasticsolutions.com/pdf/SavedAndDrivenAway.pdf
- Kohavi, R., et al. (2014). "Seven Rules of Thumb for Web Site Experimenters," KDD 2014 (slides). https://exp-platform.com/Documents/2014-08-27ExperimentersRulesOfthumbKDD.pdf ; see also Kohavi, R., Tang, D., and Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press.
- Challapally, A., et al. (2025). The GenAI Divide: State of AI in Business 2025. MIT Project NANDA. Copy read: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- Gartner. "Toolkit: Discover and Prioritize Your Best AI Use Cases With a Gartner Prism." https://www.gartner.com/en/documents/4000998
- Arora, N., et al. (2021-11-12). "The value of getting personalization right, or wrong, is multiplying." McKinsey & Company. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying
- Gartner (2019-12-02). "Gartner Predicts 80% of Marketers Will Abandon Personalization Efforts by 2025." https://www.gartner.com/en/newsroom/press-releases/2019-12-02-gartner-predicts-80--of-marketers-will-abandon-person
- Litmus. "Apple's Mail Privacy Protection Is Here: What It Actually Means for Email Marketers." https://www.litmus.com/blog/apple-mail-privacy-protection-for-marketers
This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised