# 2. The First Era

*The question: How did we personalize before language models, and what did it teach us?*

> **Questions this chapter answers**
> - How did companies personalize before large language models?
> - Which of those methods actually worked, and where did each one stop working?
> - Are rules, segments, recommenders, and A/B tests obsolete now?
> - Why did a model that predicted our customers well still lead to bad decisions?
> - How do I choose between a rule, a model, and an LLM for a given decision?

## The short answer

Before language models, companies personalized with four tools: merge fields, machine learning, experiments, and people. Each one worked, measurably, and each one hit a ceiling that the others could not lift.

Merge fields were cheap and wide, and said almost nothing. Machine learning (segments, recommenders, propensity scores) predicted what tens of millions of people would do, but could only choose among options someone had already written. Experiments made personalization honest by measuring it against a control, but a classic A/B test picks one winner for everyone. People understood customers deeply and wrote for them individually, and there were never enough hours to do that for more than a few.

So the first era split the work. Machines knew a little about everyone. Humans knew a lot about a few. Nobody could know a lot about everyone and act on it in words.

Three lessons survive intact, and a leader should insist on them before funding anything new. **Simple first**: a rule or a well-tuned model beats an LLM when the decision is narrow and the data is structured. **Prediction is not decision**: the customer most likely to buy is often the one who needs no push. **Nothing counts without a control**: if you cannot say what would have happened without the personalization, you do not know that it worked.

Language models did not make the old tools stop working. They supplied the missing piece: turning a model of one person into words for that person, without a human writing every message.

---

On September 21, 2009, Netflix handed a check for one million dollars to a team of seven researchers calling themselves BellKor's Pragmatic Chaos. [1] Three years earlier the company had released 100,480,507 movie ratings from 480,189 customers and offered the prize to anyone who could beat its own recommendation algorithm, Cinematch, by ten percent on prediction error. The winners cleared the bar with a 10.06 percent improvement. A second team matched their score and lost because it submitted twenty minutes later.

The winning solution was never put into production.

Netflix had already adopted two simpler methods from an earlier stage of the contest. The final blend of more than a hundred models was different. Xavier Amatriain, who led the algorithms team, later wrote that it "would have taken a large engineering effort for a small gain in accuracy that was most likely not worth it." [2] By then Netflix had moved from mailing DVDs to streaming, and predicting what people would actually watch mattered more than predicting the stars they would give.

That story holds most of what the first era learned. The model got better. The objective moved. The accuracy that won the contest was not the accuracy the business needed. The most famous personalization competition in history was, in the end, a lesson in choosing the right problem.

## The first era had four tools, and each hit a different ceiling

Everything companies did to personalize before large language models falls into four families, taken here from shallowest to deepest. The running example is Larkspur Systems, a fictional mid-market company selling field-service and fleet software, with about 40,000 customer accounts, 250,000 contacts, and ten years of CRM history. Larkspur runs every tool in this chapter. It still says roughly the same thing to everyone.

![Four panels ordered from shallowest to deepest show the first era's personalization tools. Merge fields are cheap and wide and say who, but never why. Machine learning (segments, recommenders, propensity scores) predicts tens of millions of people, but can only choose among options, not write the message. Experiments measure lift against a control, but a classic test picks one winner for everyone; people understand and write for each customer, but are capped by hours. The result: machines knew a little about everyone, and humans knew a lot about a few.](/images/handbook/ch02-four-tools-four-ceilings.svg)

*Figure 2.1. Each first-era tool worked, and each hit a ceiling the others could not lift.*

## Merge fields: the cheapest personalization, and the easiest to get wrong

Mail merge arrived with the first personal-computer word processors around 1980; WordStar's companion program, MailMerge, was among the earliest. [3] The idea has not changed since. A template holds slots, a list holds values, and the software fills one from the other: `Dear {first_name}`, `your team at {company}`.

It is easy to dismiss, and the evidence says not to. In randomized field experiments, Navdeep Sahni, Christian Wheeler, and Pradeep Chintagunta added the recipient's name to email subject lines. Open rates rose from 9.05 percent to 10.80 percent, leads rose by 31 percent, and unsubscribes fell by 17 percent. [4] The name carried no new information. It still changed behavior, because it signaled that the message had been addressed rather than broadcast.

The ceiling is just as clear. A merge field says *who*, never *why*. It cannot tell a customer in the middle of a renewal from one who signed last week. And when the field is wrong it does worse than generic: the uncanny near-miss, the email that uses your name and gets your company or role wrong, leaving the sender worse off than if it had stayed general. Larkspur's quarterly newsletter opens with `Hi {first_name}` for 250,000 contacts, and some of those fields hold a company name, an initial, or the word "Accounts."

## Machine learning: the era that learned to predict people

The second family is where most of the first era's real value was created. It is a stack of techniques, each built to fix the weakness of the one before, and every layer still runs in production somewhere. For each, a practitioner needs three facts: what it needs, what it is good at, and where it breaks.

| Family | What it needs | Good at | Where it breaks |
|---|---|---|---|
| Segmentation and clustering | A few descriptive attributes | Planning, creative briefs, explaining the base | Treats everyone as the segment's average |
| Propensity and churn models | Labeled outcomes (who bought, who left) | Ranking a whole base cheaply | Predicts the outcome, not the effect of acting |
| Content-based filtering | Item attributes and a person's history | New items; explainable ("because you read...") | More of the same; only as good as the tagging |
| Collaborative filtering, matrix factorization | A large matrix of interactions | Tastes nobody labeled | Cold start; optimizes the target it was given |
| Item-to-item | Purchase or view histories | Real time at catalog scale | Similar is not the same as useful |
| Two-stage deep recommenders | Very large logs and serious infrastructure | Millions of items, many signals | Cost, opacity, offline wins that fail live |
| Learning to rank | Graded labels or clicks | Ordering a whole list, not scoring items alone | Clicks carry position bias |
| Contextual bandits | Live traffic and a reward | Learning while serving | Optimizes whatever reward you give it |

**Segmentation.** Group customers who look alike, then treat each group as one: first with hand-built rules (recency, frequency, and monetary value in direct marketing), later with clustering. Segments are cheap and easy to explain. Their weakness is built in: everyone is treated as the segment's average, and the average customer does not exist.

**Propensity, churn, and next-best-offer.** These predict an action. Who is likely to buy, upgrade, or leave in the next ninety days? Score every customer, rank, and send the top of the list to a campaign or a salesperson; next-best-offer systems do the same per offer. This is still the backbone of enterprise personalization, and it works. Its blind spot is this chapter's failure story: a score says who will act, not who will act *because* you did something.

**Recommender systems.** Here the first era reached individual scale. Content-based systems recommend items whose attributes resemble what the person already liked. Collaborative filtering ignores attributes and uses behavior: you work out how similar users are to one another, and how similar items are to one another, and you use the two together to match people to things. I have built features of this kind; it is one of the most ordinary techniques in software, and one of the most effective.

The Netflix Prize made the next step famous. Yehuda Koren, Robert Bell, and Chris Volinsky, three members of the winning team, summarized it in 2009: matrix factorization, which represents each person and each item as a short vector of learned "latent factors," beat the classic nearest-neighbor methods and could absorb implicit feedback, changes over time, and confidence levels. [17] Nobody names those factors in advance; the model finds the dimensions of taste that best explain the ratings.

The landmark industrial paper is Amazon's. In 2003, Greg Linden, Brent Smith, and Jeremy York described item-to-item collaborative filtering, built because existing methods could not scale to a store with "more than 29 million customers and several million catalog items." [5] Their key move was architectural, and it is worth a builder's attention: compute the expensive similar-items table offline, so that the online step depends only on how many items this customer has bought, not on the size of the catalog or the customer base. The result ran in real time and worked from as few as two or three purchases. The same paper observed that clustering customers into segments, the cheaper alternative, produced recommendations of "relatively poor" quality.

At Netflix, Carlos Gomez-Uribe and Neil Hunt reported in 2015 that about 80 percent of hours streamed came from recommendations, and estimated their retention value at more than one billion dollars a year, a company self-estimate rather than audited revenue. [6] A more careful 2025 Netflix study swapped the production system for simpler alternatives: engagement fell about 4 percent against basic matrix factorization and about 12 percent against popularity rankings, with the largest gains on titles that were neither blockbusters nor obscure. [7]

**Two stages: candidates, then ranking.** Deep learning did not replace this architecture; it formalized it. Paul Covington, Jay Adams, and Emre Sargin described YouTube's 2016 recommender as two networks. [18] The first, candidate generation, retrieves "a small subset (hundreds) of videos" from a corpus of millions using the person's history. The second ranks those hundreds with richer features, predicting expected watch time rather than click probability. They added a sentence every model team should pin to the wall: "live A/B results are not always correlated with offline experiments." A later YouTube system made retrieval a *two-tower* model: one network embeds the person and context, another embeds each item from its content features, and retrieval becomes a nearest-neighbor search over tens of millions of videos. [19] Ranking borrowed from search, where *learning to rank* trains on the order of a whole list rather than each item alone; an ensemble of LambdaMART rankers won Track 1 of the 2010 Yahoo! Learning to Rank Challenge. [20]

**Contextual bandits.** A bandit learns while it serves: it shifts traffic toward what is winning while still exploring, and a contextual bandit chooses per person from their features. On the Yahoo! front page, Lihong Li and colleagues reported a 12.5 percent click lift from a contextual bandit over a context-free one. [11] Chapter 14 treats it as a decision-maker.

**Features, served fresh.** Every family above is only as current as its inputs. When Uber described its machine learning platform in 2017, the centerpiece was a shared feature store, about 10,000 features that teams could reuse, built so that "the same data is used for training and serving." [21] That sentence names a common silent bug: a feature computed one way for training and another at serving time, so the model scores a customer it never learned. Chapter 17 covers when a feature must be fresh.

**Cold start.** Collaborative methods share one structural weakness. A model built from a population is a summary of what a population did, and with no history there is nothing to summarize. These systems have always struggled with the new arrival, the account that signed yesterday: the cold-start problem. There is nobody yet to be similar to. The standard patches are content features (the item tower in a two-tower model exists partly for this), popularity, stated preferences at signup, and a bandit's deliberate exploration.

The ceiling of all of these is the same. Machine learning could *choose* among things someone had already made (a segment's message, a catalog item, one of five offers), but it could not *write* the thing for the person. The model of the customer could be individual. The message could not.

What carries forward is the architecture, not just the models. The shape that scales is still candidate generation, then ranking, now followed by a language layer that writes for the one person the ranker chose. And the scores this era learned to produce (propensity, churn risk, similarity, predicted uplift) do not disappear. They become inputs to the decision layer in Chapter 14, which weighs them against constraints, cost, and the option of doing nothing.

## Experiments: the discipline that made personalization measurable

The third family is not a way to personalize. It is the way to know whether personalization worked, and it is the first era's most durable contribution.

In December 2007, Dan Siroker ran an experiment on the Obama campaign's splash page, testing combinations of images and button text. The winning variant raised sign-ups from 8.26 percent to 11.6 percent, a 40.6 percent relative lift. [8] (The widely repeated "$60 million" attached to that test is an extrapolation of donations from the extra sign-ups, not a measured figure.) At Bing in 2012, a change to how ad headlines were displayed sat in the backlog for more than six months before an engineer finally tested it. It raised revenue by 12 percent, worth more than $100 million a year in the US alone. [9]

The deeper lesson from experimentation is humility. Ronny Kohavi, who ran Microsoft's experimentation team, and his colleagues put the success rate of ideas at Bing at about 10 to 20 percent, and note that most progress comes from small improvements of 0.1 to 1 percent, "after a lot of work." [10] The confident plan in the meeting is usually wrong. The only defense is a control group.

**A/B and multivariate tests.** An A/B test randomizes people between a control and one change. A multivariate test, like Siroker's, crosses several factors at once. Kohavi, Randal Henne, and Dan Sommerfield's practical guide recommends single-factor tests for incremental changes and factorial designs only when factors are suspected to interact strongly, because every added combination splits the audience and cuts power; in their experience, interactions are less frequent than people assume. [22]

**Power, in plain terms.** Power is the chance that a test detects an effect that is really there. The required sample grows with the square of the precision you ask for: halve the effect you want to detect, and you need four times the people. The same guide works an example: on a site where 5 percent of visitors buy, detecting a 5 percent relative change in conversion takes under 500,000 users, while a 20 percent change takes about 30,400. [22] Personalization effects usually sit at the small end, so many tests are underpowered; Chapter 21 does this arithmetic for Larkspur.

**Peeking and sequential tests.** A classic test fixes its sample in advance; checking daily and stopping at the first significant result inflates false positives. Ramesh Johari and colleagues described "always valid" p-values that stay correct under continuous monitoring, deployed in a commercial testing platform since January 2015. [23] The rule for a team is simple: either fix the horizon and do not stop early, or use a sequential method built for looking.

**Variance reduction.** The cheapest power is the noise you remove. CUPED, from Alex Deng, Ya Xu, Ron Kohavi, and Toby Walker, adjusts each person's outcome using their own behavior before the experiment began. On Bing it reduced variance by about 50 percent, "effectively achieving the same statistical power with only half of the users, or half the duration," and the best covariate was usually the same metric from the pre-period. [24] Personalization programs have exactly this data: the customer's history is the covariate.

**Interleaving for rankers.** When the thing under test is a ranking, you can show one person both versions at once. Interleaving merges two rankers' lists and credits whichever one supplied what the person chose; Olivier Chapelle, Thorsten Joachims, Filip Radlinski, and Yisong Yue validated it at scale on two commercial search engines and a scientific-literature search system. [25] Netflix uses a variant of team-draft interleaving: the two rankers take turns picking titles like captains choosing sides, and viewing hours are credited to whichever supplied each title. Its team reported in 2017 that interleaving "requires >100× fewer users than our most sensitive A/B metric to achieve 95% power," and that its verdicts correlated strongly with that metric. [26] The catch is that interleaving measures relative preference, not retention, so Netflix uses it as a fast pruning stage and runs a conventional A/B test on the survivors.

**Bandits and uplift.** Bandits trade proof for speed; Chapter 21 explains when that trade is right and why a bandit cannot tell you whether a program was worth running. The second idea that pushed experiments toward the individual was **uplift modeling**, introduced by Nicholas Radcliffe and Patrick Surry in 1999 and named "true lift" by Victor Lo in 2002. [12] Uplift models do not predict who will buy. They predict who will buy *because* you contacted them, which is a different question and, for spending decisions, the right one. Holdouts, incrementality, and misleading proxies belong to Chapter 21.

The ceiling of classic experimentation is that it finds one winner for everyone: the best page on average, not the best page for this visitor. Bandits and uplift models narrowed that gap, but still chose among variants a person had written.

That ceiling is also why experiments are the first era's best gift to the next one. When a language model writes a different message for every person, there is no variant B to test against variant A, and the power arithmetic rules out testing message variants within any one context. So the unit of the experiment moves up a level. You test what stays fixed while the words change: the policy that decides who is contacted and when, the template and the zones a model may fill, the decision rules, the model and prompt version as a treatment arm, all against a holdout. You cannot test every message. You can test the system that writes them.

## Humans by hand: deep, and capped by hours

The fourth family is the oldest and still the deepest. An account manager who has handled a customer for years knows the politics inside the account, the failed project from two years ago, and that the new operations lead prefers a phone call. A good researcher reads a company's filings, job posts, and the champion's public writing, and comes back with the one fact that changes the conversation.

In 1993, Don Peppers and Martha Rogers argued in *The One to One Future: Building Relationships One Customer at a Time* that companies should compete for share of customer, not share of market. [13] The subtitle is close to this handbook's title, and the lineage is deliberate. They described the goal. What the first era lacked was a way to reach it without a person in every relationship.

That is the ceiling: hours. Salesforce's 2022 survey of 7,775 sales professionals found reps spending 28 percent of their time actually selling. [14] Larkspur has twelve account managers. They cover the top 400 accounts deeply. The other 39,600 accounts get segments, merge fields, and a propensity score.

## What the first era got right, and what the new era keeps

Here is the sentence this chapter turns on. **Every first-era method was good at one of three things (knowing the customer, deciding what to do, or saying it well), and none could do all three for the same person at scale.** Models knew, a little, about everyone. Experiments decided, on average. People said it well, for a few.

![A grid with three jobs as columns (knowing the customer, deciding what to do, saying it well in words) and three methods as rows. Models fill only the knowing cell, a little about everyone; experiments fill only the deciding cell, on average; people fill only the saying cell, for a few. A final row for one person at scale shows knowing built at scale, deciding only on average, and saying never built: nothing turned one customer's model into a message for that customer.](/images/handbook/ch02-knowing-deciding-saying.svg)

*Figure 2.2. Knowing, deciding, and saying: each first-era method covered one, and none covered all three for the same person at scale.*

Knowing a person and speaking to that person are separate capabilities. The first era built the first at scale and never the second: nothing could turn one customer's model into a message for that customer, and read the reply, without a person at each end. That is the subject of Chapter 5.

It does not follow that the old tools are obsolete. Three of them are more important now, not less.

**Rules still govern.** Contact limits, eligibility, consent, and suppression stay deterministic. A language model should never decide whether a customer who unsubscribed gets an email.

**Models still rank.** A propensity score for 250,000 contacts costs almost nothing. A language model reasoning about each contact is justified only where the reasoning changes the decision.

**Experiments still judge.** Every claim about lift should end in a controlled comparison (Chapter 21). They prove what moved, not why; Chapter 3 pairs them with qualitative research and models in one loop.

The split survived contact with production when I built a governed personalization engine for B2B campaigns. The rule it settled on is one line: use AI where judgment adds value, use code where certainty is available. A headline limit, a ban on non-HTTPS links, and the rule that copy may not say "your renewal is coming up" unless the relationship is confirmed are all enforced in code. Choosing the most relevant approved proof point is left to the model. Scoring followed the same line: a language model combines weak and strong signals into a score with a rationale, inside the campaign's criteria, but it does not replace the deterministic qualification rules, and code checks that its outputs agree with each other (a score of 88 labeled "low" is flagged, not shipped). Population questions, such as every contact scored above 80 or every record missing an artifact, are answered by a deterministic filter with no model call at all. And when someone thinks a message would perform better, that idea goes to a controlled test, not into another paragraph of prompt.

This is the **Simple First** principle: use the least expensive technique that makes the decision well, and add depth only where depth changes the outcome.

![A three-step ladder rising left to right along an axis of cost, speed, and difficulty to explain. Tier 1 is rules, which govern eligibility, consent, and contact limits; tier 2 is classic models, which rank with propensity scores and recommenders; tier 3 is language models, which read unstructured evidence and write for one person. Arrows show moving up a tier only when the cheaper tier demonstrably fails, and a band beneath all tiers shows experiments judging every program against a control group or holdout.](/images/handbook/ch02-simple-first-ladder.svg)

*Figure 2.3. Simple First: rules, then classic models, then language models, each tier earned by the failure of the one below.*

### Choosing a technique

Three questions sort most decisions: how much structured data you have, how costly a wrong call is, and how much someone needs to understand why the decision was made.

| Data available | Stakes of a wrong call | Need for explanation | Start with |
|---|---|---|---|
| Little or none | Low | Low | Rules, or popularity plus a bandit |
| Rich behavioral history | Low | Low | Recommender or propensity model |
| Rich history | High (money, contract, compliance) | High | Rules for eligibility, uplift model for targeting, human approval |
| Mostly unstructured (calls, emails, documents) | Medium | Medium | Language model reading the evidence, with a deterministic check after |
| Unstructured and deep, one account at a time | High | High | Language model research plus a human decision |

In words: structured data and low stakes favor classic models; high stakes demand rules and human judgment around any model; unstructured evidence is where language models earn their cost, because nothing in the first era could read it.

## What this does not do

The first era's record is distorted by numbers repeated past the evidence, some now used to sell the second era.

The most famous is that "35 percent of what customers buy on Amazon comes from recommendations." It is usually traced to a 2013 McKinsey article; Amazon has never published it. Linden, Smith, and York themselves wrote only that click-through and conversion rates "vastly exceed those of untargeted content." [5] Netflix's billion dollars is a retention estimate. Obama's $60 million is an extrapolation. None of these is invented, but none is the measured fact it is usually presented as.

A quieter overreach is the belief that a better model means a better business result. The Netflix Prize says otherwise, and so does Booking.com, which reported after deploying some 150 models that a better offline model did not necessarily bring more business value. [15] A recommender that looks better in offline testing routinely loses to a worse-looking one on real people. The same mistake is being made now with language models, scored on benchmarks and deployed on customers.

And the first era did not fail. In the companies I work with, recommenders and propensity models remain the working core of personalization (an observation, not a survey). What the era could not do was reach depth for most customers. That is a ceiling, not a failure.

## At scale

At Larkspur's size, each team owns one first-era tool: marketing owns segments and merge fields, product owns recommendations, sales owns the propensity score, customer success owns the account notes. Each holds a partial model of the same customer, and none reads the others. The churn model does not know about the open support escalation. The account manager's knowledge leaves when they do.

At millions of customers the pattern is the same, only more expensive. The first era's scale problem was never compute. It was that understanding stayed fragmented across tools that could not share it.

## Failure story: The Discount Mistake

Larkspur's data science team builds a good upgrade-propensity model. Marketing takes the top decile, about 4,000 accounts, and sends them a 20 percent discount on the next tier. Upgrades in that group are strong. The campaign is declared a success.

A year later someone asks what those accounts would have done without the discount. A small holdout, kept by accident, answers: they upgraded at almost the same rate. The model had correctly found the customers most likely to upgrade, and the campaign paid them to do what they were already going to do. (Illustrative numbers; the pattern is well documented.)

Uplift research calls these customers "Sure Things," alongside "Persuadables," "Lost Causes," and "Do-Not-Disturbs," the last being customers whom contact makes worse. [12] Studies of retention campaigns have found that targeting the highest-risk customers can be ineffective, and that the better target is sensitivity to the intervention, not risk. [16]

The error is reading high propensity as need for an incentive. Prediction is not decision.

## Patterns

**Simple First.** *Problem:* teams reach for the most capable technique by default. *Forces:* capable techniques are costlier, slower, harder to explain. *Solution:* start with rules, then classic models, then language models, moving up only when the cheaper tier demonstrably fails. *Tradeoff:* some early lift left on the table, in exchange for cost control and a baseline to beat.

**Keep a Holdout.** *Problem:* results are reported without a counterfactual. *Forces:* holdouts cost short-term revenue and feel wasteful. *Solution:* a random control group for every personalized program, plus a small permanent global holdout. *Tradeoff:* a small real cost, against knowing whether anything worked.

**Target Uplift, Not Propensity.** *Problem:* incentives go to customers who would act anyway. *Forces:* propensity models are easier to build and look accurate. *Solution:* for any action that costs money or attention, estimate the effect of the action, not the likelihood of the outcome. *Tradeoff:* uplift needs experimental data and is noisier.

## Leader questions

1. For our three largest personalization programs, what is the control group, and what did it do?
2. Which of our decisions are made by rules, which by models, and which by people? Is that split deliberate?
3. Where do we give incentives to customers most likely to buy, and have we measured whether they needed them?
4. Before we fund an LLM for this decision, what simpler method did we try, and how did it do?

## Build checklist

- [ ] Inventory every personalization in production and label its technique: rule, segment, recommender, propensity, experiment, human.
- [ ] Audit merge fields for bad values (company names in first-name fields, placeholders, stale titles) and define a fallback.
- [ ] Attach a control group to every personalized program; keep a small permanent global holdout.
- [ ] For every incentive-based campaign, compare treated and untreated outcomes before scaling.
- [ ] Define a cold-start path for new customers that does not depend on history.
- [ ] Keep eligibility, consent, and contact limits in deterministic code, separate from any model.
- [ ] Compute every model feature through one path for training and serving, and use pre-period behavior to reduce variance in tests.

## Metrics to watch

- **Incremental lift versus holdout**, per program, not raw conversion.
- **Share of decisions with a counterfactual**: the percentage of personalized actions covered by a control.
- **Merge-field error rate**: records whose personalization fields are missing or malformed.
- **Incentive efficiency**: incremental conversions per discount dollar, not conversions per discount.

## Reader Q&A

**Are A/B tests obsolete now that AI can personalize each message?**
No. When every message is different you cannot test messages against each other, but you can and must test the system against a holdout. The unit of the experiment moves from the variant to the policy.

**We have a propensity model that works. Do we need an LLM?**
Maybe not for ranking. A language model adds value where the decision depends on evidence your model cannot read (call notes, emails, documents) or where a message must be written for one person. Keep the propensity model; use it to decide where deeper work is worth paying for.

**Why do our recommendations work for long-time customers and fail for new ones?**
That is cold start: with no history, there is nobody to be similar to. Rules, popularity, stated preferences, and now a language model reading what little a new customer has said can cover the gap.

## For your AI

```yaml
chapter: 2
concepts:
  - name: First-era personalization
    definition: "Personalization before LLMs, using merge fields, machine learning (segments, recommenders, propensity models), experiments, and human effort."
  - name: Simple First
    definition: "Use the least expensive technique that makes the decision well; add depth only where depth changes the outcome."
  - name: Knowing, deciding, saying
    definition: "The three functions of personalization; each first-era method did one well, none did all three for one person at scale."
  - name: Cold start
    definition: "Failure of history-based models for new customers with no behavior to compare."
  - name: Uplift
    definition: "The change in outcome caused by an action for a person, as opposed to the probability of the outcome."
  - name: Candidate generation and ranking
    definition: "The two-stage recommender architecture: a cheap stage retrieves hundreds of plausible items from millions, a richer stage orders them; in the LLM era a language layer follows and ML scores feed the decision layer."
  - name: Training-serving consistency
    definition: "The requirement, served by a feature store, that a feature is computed the same way when a model is trained and when it is used."
  - name: Variance reduction (CUPED)
    definition: "Adjusting each person's experiment outcome by their own pre-experiment behavior to reach the same power with fewer people or less time."
  - name: Interleaving
    definition: "Comparing two rankers by merging their lists for the same person and crediting the ranker whose items were chosen; far more sensitive than A/B for rankers, but measures relative preference only."
decision_rules:
  - if: "the decision is narrow, the data is structured, and the stakes are low"
    then: "use rules or a classic model before a language model"
  - if: "an action costs money or attention (discount, call, retention offer)"
    then: "target on estimated uplift, not propensity, and keep a holdout"
  - if: "a personalized program has no control group"
    then: "treat its reported lift as unproven"
  - if: "a merge field value is missing or malformed"
    then: "fall back to an honestly general greeting"
  - if: "the decision depends on unstructured evidence (calls, emails, documents)"
    then: "consider a language model to read it, with deterministic checks after"
  - if: "two ranking algorithms must be compared"
    then: "prune with interleaving, then confirm business outcomes with an A/B test on the survivors"
  - if: "results will be checked while a test is running"
    then: "use a sequential method with always-valid inference, or fix the horizon and do not stop early"
  - if: "each person receives a differently generated message"
    then: "test the policy, template, zones, decision rules, or model version against a holdout, not individual messages"
  - if: "a requirement can be checked exactly (length, link scheme, confirmed relationship, score consistency)"
    then: "enforce it in code, not in the prompt; reserve the model for judgment within approved evidence"
assessment_questions:
  - "Which personalization programs run today, and which technique drives each?"
  - "Which programs have a control group or global holdout?"
  - "Where are incentives given, and has incremental effect been measured?"
  - "How are new customers with no history personalized?"
  - "What customer knowledge exists only in account managers' notes or heads?"
patterns: [Simple First, Keep a Holdout, Target Uplift Not Propensity]
anti_patterns: [The Discount Mistake, The Uncanny Near-Miss, Offline Accuracy Worship, Peeking, Training-Serving Skew]
maturity_dimension: decisioning
```

## References

1. Netflix Prize: launch October 2, 2006; 100,480,507 ratings from 480,189 users on 17,770 movies; Grand Prize awarded September 21, 2009, to BellKor's Pragmatic Chaos, 10.06% RMSE improvement. https://en.wikipedia.org/wiki/Netflix_Prize ; award announcement: https://www.netflixprize.com/community/topic_1537.html
2. Amatriain, X. "On the 'Usefulness' of the Netflix Prize." https://amatria.in/blog/on-the-usefulness-of-the-netflix-prize-403d360aaf2/ ; original: Amatriain & Basilico, "Netflix Recommendations: Beyond the 5 stars (Part 1)," Netflix Tech Blog, April 2012. https://netflixtechblog.com/netflix-recommendations-beyond-the-5-stars-part-1-55838468f429
3. "Mail merge" and "WordStar," Wikipedia (MicroPro's MailMerge, c. 1980). https://en.wikipedia.org/wiki/Mail_merge
4. Sahni, N. S., Wheeler, S. C., & Chintagunta, P. (2018). "Personalization in Email Marketing: The Role of Noninformative Advertising Content." *Marketing Science* 37(2). https://pubsonline.informs.org/doi/abs/10.1287/mksc.2017.1066
5. Linden, G., Smith, B., & York, J. (2003). "Amazon.com Recommendations: Item-to-Item Collaborative Filtering." *IEEE Internet Computing* 7(1), 76 to 80. https://dl.acm.org/doi/10.1109/MIC.2003.1167344
6. Gomez-Uribe, C. A., & Hunt, N. (2015). "The Netflix Recommender System: Algorithms, Business Value, and Innovation." *ACM TMIS* 6(4). https://dl.acm.org/doi/10.1145/2843948
7. Zielnicki et al. (2025). "The Value of Personalized Recommendations: Evidence from Netflix." arXiv:2511.07280. https://arxiv.org/abs/2511.07280
8. Siroker, D. (2010). "How Obama raised $60 million by running a simple experiment." Optimizely. https://www.optimizely.com/insights/blog/how-obama-raised-60-million-by-running-a-simple-experiment/
9. Kohavi, R., & Thomke, S. (2017). "The Surprising Power of Online Experiments." *Harvard Business Review*. https://hbr.org/2017/09/the-surprising-power-of-online-experiments
10. Kohavi, R., Tang, D., & Xu, Y. (2020). *Trustworthy Online Controlled Experiments.* Cambridge University Press. https://experimentguide.com/ ; Kohavi et al. (2014), "Seven Rules of Thumb for Web Site Experimenters." https://exp-platform.com/Documents/2014-08-27ExperimentersRulesOfthumbKDD.pdf
11. Li, L., Chu, W., Langford, J., & Schapire, R. E. (2010). "A Contextual-Bandit Approach to Personalized News Article Recommendation." WWW 2010. https://arxiv.org/abs/1003.0146
12. Radcliffe, N. J., & Surry, P. D. (1999). "Differential Response Analysis: Modeling True Response by Isolating the Effect of a Single Action." *Credit Scoring and Credit Control VI*; Lo, V. S. Y. (2002). "The True Lift Model." *SIGKDD Explorations* 4(2); Radcliffe & Surry (2011), "Real-World Uplift Modelling with Significance-Based Uplift Trees." https://www.research.ed.ac.uk/en/publications/real-world-uplift-modelling-with-significance-based-uplift-trees/
13. Peppers, D., & Rogers, M. (1993). *The One to One Future: Building Relationships One Customer at a Time.* Currency/Doubleday.
14. Salesforce (2022). *State of Sales*, 5th ed. (n = 7,775). https://www.salesforce.com/news/stories/sales-research-2023/
15. Bernardi, L., Mavridis, T., & Estevez, P. (2019). "150 Successful Machine Learning Models: 6 Lessons Learned at Booking.com." KDD 2019. https://dl.acm.org/doi/10.1145/3292500.3330744
16. Ascarza, E. (2018). "Retention Futility: Targeting High-Risk Customers Might Be Ineffective." *Journal of Marketing Research* 55(1). https://journals.sagepub.com/doi/10.1509/jmr.16.0163 ; Radcliffe, N. J., & Simpson, R. (2008). "Identifying who can be saved and who will be driven away by retention activity." https://stochasticsolutions.com/pdf/SavedAndDrivenAway.pdf
17. Koren, Y., Bell, R., & Volinsky, C. (2009). "Matrix Factorization Techniques for Recommender Systems." *IEEE Computer* 42(8), 30 to 37. https://dl.acm.org/doi/10.1109/mc.2009.263
18. Covington, P., Adams, J., & Sargin, E. (2016). "Deep Neural Networks for YouTube Recommendations." RecSys 2016. https://research.google/pubs/deep-neural-networks-for-youtube-recommendations/
19. Yi, X., et al. (2019). "Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations." RecSys 2019. https://dl.acm.org/doi/10.1145/3298689.3346996
20. Burges, C. J. C. (2010). "From RankNet to LambdaRank to LambdaMART: An Overview." Microsoft Research Technical Report MSR-TR-2010-82. https://www.microsoft.com/en-us/research/publication/from-ranknet-to-lambdarank-to-lambdamart-an-overview/
21. Hermann, J., & Del Balso, M. (2017-09-05). "Meet Michelangelo: Uber's Machine Learning Platform." Uber Engineering Blog. https://www.uber.com/us/en/blog/michelangelo-machine-learning-platform/
22. Kohavi, R., Henne, R. M., & Sommerfield, D. (2007). "Practical Guide to Controlled Experiments on the Web: Listen to Your Customers not to the HiPPO." KDD 2007. https://ai.stanford.edu/~ronnyk/2007GuideControlledExperiments.pdf
23. Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2017). "Peeking at A/B Tests: Why It Matters, and What to Do about It." KDD 2017, 1517 to 1525. https://www.semanticscholar.org/paper/b8b8c4627bf9bea8118a098f7dcd1612603f4795
24. Deng, A., Xu, Y., Kohavi, R., & Walker, T. (2013). "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data." WSDM 2013. https://exp-platform.com/Documents/2013-02-CUPED-ImprovingSensitivityOfControlledExperiments.pdf
25. Chapelle, O., Joachims, T., Radlinski, F., & Yue, Y. (2012). "Large-Scale Validation and Analysis of Interleaved Search Evaluation." *ACM TOIS* 30(1). https://dl.acm.org/doi/10.1145/2094072.2094078
26. Parks, J., Aurisset, J., & Ramm, M. (2017-11-29). "Innovating Faster on Personalization Algorithms at Netflix Using Interleaving." Netflix TechBlog. https://netflixtechblog.com/interleaving-in-online-experiments-at-netflix-a04ee392ec55
