# 1. What Personalization Is, and What It Is Worth

*The question: What is personalization, and why should a leader care now?*

> **Questions this chapter answers**
> - What do we actually mean by "personalization," and how is it different from segmentation or customization?
> - Is it worth the investment, and how much is it worth for us?
> - What should we be optimizing for: clicks, revenue, or something else?
> - Why does some personalization feel helpful and some feel creepy?
> - What is the minimum quality bar before we let a machine speak to a customer as an individual?

## The short answer

Personalization is not a feature, a merge field, or a data asset. It is a decision: choosing what to say, show, offer, or do for one specific person, based on what you know about them, instead of what you would do for everyone. Holding a lot of personal data is not the same as making personalized decisions. Most companies have the first and very little of the second.

The evidence says it is worth doing, within honest limits. In one randomized field experiment, adding a recipient's name to an email subject line raised opens by about 20% and leads by about 31% [4]. McKinsey reports typical revenue lifts of 10 to 15% for companies that do it well [1]. The evidence also shows the downside. In a 2024 Gartner survey, 53% of customers reported a negative personalization experience, and those customers were 3.2 times more likely to regret a purchase [3]. Personalization shifts the odds; it does not command outcomes. Done badly, it shifts them the wrong way.

So judge every personalized decision on four separate tests: relevance, usefulness, timing, and appropriateness. Measure its worth per decision, as lift over a real baseline, net of cost and risk. And hold one standard above the others: every personal claim is either **specific and true, or honestly general**. A message confidently wrong about a customer's own life costs credibility you had before you sent it.

Why now: the cost of understanding one person deeply has fallen far enough that depth no longer has to be rationed to your largest accounts. That changes which decisions are worth personalizing, and raises the price of getting them wrong at scale.

---

You have sent both kinds of message.

On a Tuesday morning at Larkspur Systems (a fictional mid-market software company that runs through this handbook), a customer success manager writes one email to one fleet operations director. She names the thing the director is actually dealing with this month: a new depot opening in two weeks, and the dispatch rules nobody has configured for it yet. She offers a thirty-minute setup call before the opening.

The same morning, the quarterly newsletter from the introduction goes out to 250,000 contacts. It produces the usual open rate and a few unsubscribes. The single email gets a reply within the hour that says thank you.

You already know the difference, because you enact it on both sides. As a customer, you delete the message written for everyone without quite deciding to, and you answer the one written for you. As an operator, you know which of the two your company can afford to send 250,000 times, and which it can afford only a few dozen times a week, by hand.

That gap is what this handbook is about. This chapter defines it, prices it, and sets the bar for crossing it.

## Personal data is an input. Personalization is a decision.

The word "personalization" is used for at least five different things, and the confusion is expensive, because teams buy tools for one and expect the results of another.

| Term | Who chooses | What varies | Example at Larkspur |
|---|---|---|---|
| **Segmentation** | The company, per group | Treatment per segment; everyone in a segment gets the same thing | Five segments by fleet size, one email version each |
| **Customization** | The person | What the person configures for themselves | A dispatcher sets their own dashboard layout |
| **Personalization** | The company, per person | The treatment for one individual, based on what is known about them | The renewal email mentions the depot this customer is opening |
| **Adaptive** | The system, over time | The treatment adjusts as the person responds | Onboarding steps reorder based on which features a new customer actually uses |
| **Agentic** | An AI system, within bounds | Research, decision, action, and follow-up, over several steps, for one person | An agent researches the account, drafts the note, schedules the follow-up, and waits for the reply before deciding what comes next |

These are not rungs you must climb in order; a mature program uses all five. But only the last three are decisions made about an individual. Segmentation decides for a group. Customization lets the customer decide.

A company can hold ten years of CRM history, product telemetry, support tickets, and call transcripts and still make no personalized decisions at all. Larkspur is exactly this company. It knows which accounts are renewing, which are angry, which just hired a new operations lead. It says the same thing to all of them. Data in a warehouse is potential. Personalization happens only at the moment a decision about one person uses it.

That tells you where to look for value: not in "our data," but in the repeated decisions your company makes about individual customers every day. What to send, what to show, what to offer, whom to call, when to wait. Chapter 6 turns that into an inventory. For now, hold the definition: **personalization is choosing the next action for one person, using what you know about them, where the choice would have been different for someone else.** If the choice would have been the same for everyone, it was not personalized, whatever the template says.

## A good personalized decision passes four separate tests

Teams tend to ask one question: "Is it relevant?" That question hides four different failures, with different causes and different fixes.

1. **Relevance.** Does it concern something this person has, does, or cares about? Fleet electrification pitched to a customer with no electric vehicles fails here. Usual cause: bad or missing data.
2. **Usefulness.** Does it help them do, decide, or avoid something? "We noticed you use our dispatch module!" is relevant and tells the customer nothing. Usual cause: no point of view about what the person needs.
3. **Timing.** Does it arrive when it can still matter? The setup offer two weeks before the depot opens is service; two weeks after, it is proof you noticed too late. Usual cause: batch schedules that ignore the customer's clock.
4. **Appropriateness.** Would this person expect, and accept, you knowing and using this here? Usual cause: a system that treats "we have it" as "we may use it."

![Four side-by-side panels show the four tests a personalized decision must pass, each with its question and its usual cause of failure: relevance (is it about something they have, do, or care about; bad or missing data), usefulness (does it help them do, decide, or avoid something; no point of view on need), timing (does it arrive when it can still matter; batch schedules that ignore their clock), and appropriateness (would they expect and accept you using this here; treating "we have it" as "we may use it"). All four feed one gate: the decision ships only when all four pass, scored separately with one owner per test.](/images/handbook/ch01-four-tests.svg)

*Figure 1.1. "Is it relevant?" hides four independent tests; a decision ships only when it passes all four.*

The tests are independent, and the failed one is often the one the customer remembers. Gartner's 2025 research found that personalization relevant at one stage of a purchase became irrelevant when the buyer moved to a harder task, and that customers receiving personalization were twice as likely to feel overwhelmed by information [3]. Relevance without usefulness and timing produces more messages, not better ones.

## What the evidence actually shows

Personalization statistics are among the most recycled numbers in business writing, and many of them cannot be traced to a method. So sort the evidence by what kind of claim it is before deciding how much weight it carries.

**Surveys of what customers say.** McKinsey's 2021 research reported that 71% of consumers expect personalized interactions and 76% get frustrated when they do not get them [1]. Salesforce's 2022 survey of more than 13,000 consumers and nearly 4,000 business buyers found 73% expect companies to understand their unique needs [5]; its 2023 edition found 61% feel treated as a number [6]. These are real findings about stated expectations, not measurements of behavior. A widely quoted 2018 figure, "80% of consumers are more likely to purchase" when brands personalize, came from a vendor survey of 1,000 US consumers asked about their intentions [7]. It is a fair example of how industry numbers are made, and a poor basis for a budget.

**Reports of what companies experienced.** McKinsey also reported that personalization most often drives a 10 to 15% revenue lift, with company-specific results from 5 to 25%, and that faster-growing companies drive 40% more of their revenue from personalization than slower-growing peers [1]. These come from a consultancy's client experience, not controlled experiments, and the second is a correlation: growing companies may personalize more because they are growing. Treat the range, not the midpoint, as the finding.

**Field experiments.** This is the strongest evidence, and it is narrower. Sahni, Wheeler, and Chintagunta ran randomized email experiments with real companies: adding the recipient's name to the subject line raised the open rate from 9.05% to 10.80%, raised sales leads by 31%, and cut unsubscribes by 17% [4]. The name told the reader nothing new. It shows how much attention the first signal of "this is for you" can buy, and also the ceiling of shallow personalization: the absolute rates stay low.

**Experiments on deeper tailoring.** Matz and colleagues reported in 2017 that ads matched to inferred personality, reaching about 3.5 million people, produced up to 40% more clicks and 50% more purchases [8]. Critics pointed out that ad-platform delivery was not fully randomized, which can inflate the effect [9]. A 2026 meta-analysis of 41 studies found that digital footprints predict personality only weakly and that personality-tailored messages have negligible average effects on behavior [10]. The honest reading: tailoring on a thin, inferred trait is a weak lever, and effects depend on what is inferred, how accurately, and what outcome is measured.

**Evidence of the downside.** Gartner's late-2024 survey of 1,464 B2B buyers and consumers found that 53% had experienced negative personalization, which made them 3.2 times more likely to regret a purchase and 44% less likely to buy again [3]. Twilio Segment's 2024 report found 61% of consumers worried that inaccurate data would undermine AI personalization [11].

Here is what it adds up to. Personalization does not pull a lever inside a specific person. It shifts the odds, a little, across a great many people, and a small shift across many decisions is worth a great deal. It is a game of rates, not of certainties. And the shift can go either way: the same machinery that earns a reply can earn a regret.

## The worth of personalization is measured per decision

No company buys personalization in aggregate. It buys better decisions, one decision type at a time. So price it that way.

**Value per decision = (outcome value) × (lift over baseline) − (cost to make the decision) − (expected cost of getting it wrong)**

- **Outcome value**: what one improved result is worth (a renewal saved, a meeting booked, a support contact avoided), in revenue, retention, cost to serve, or team time.
- **Lift over baseline**: the incremental change against your current generic treatment, not against doing nothing. Chapter 21 covers how to measure it.
- **Cost to decide**: data, compute, review time, delivery. It has fallen sharply for research and drafting (Chapters 4 and 5).
- **Cost of being wrong**: the term most business cases leave out. The Gartner numbers above are the reason to put it back.

Work one example with the assumptions visible (all figures illustrative). Larkspur has roughly 6,000 at-risk renewals a year. Suppose a saved renewal is worth $6,000 to $15,000 in annual contract value, and a researched, account-specific renewal plan improves the save rate by 1 to 4 percentage points over the generic sequence. That yields 60 to 240 additional saved renewals, worth roughly $0.36 million to $3.6 million a year. The range is wide, and it should be. The assumption carrying the weight is the lift, and only a test narrows it. Subtract research, drafting, and the expected cost of a damaging message, and you have a number you can defend, because every term can be checked.

![The formula for value per decision: outcome value times lift over the generic baseline, minus cost to decide, minus expected cost of being wrong, with the last term highlighted as the one most business cases leave out. Below it, the illustrative Larkspur (fictional) renewal example: 6,000 at-risk renewals a year times a 1 to 4 point save-rate lift gives 60 to 240 more saved renewals, and times $6,000 to $15,000 per saved renewal gives $0.36 million to $3.6 million a year before costs. A to-scale bar on a $0 to $4 million axis shows how wide that range is; the lift carries the weight, and only a test narrows it.](/images/handbook/ch01-value-per-decision.svg)

*Figure 1.2. Price each decision type as a range; the lift assumption drives the width, and only a test narrows it.*

Run the same arithmetic on the quarterly newsletter. The volume is large, but the value per recipient is small and the achievable lift modest. This is the most common error in first programs: personalizing the highest-volume touchpoint because it is easiest to automate, not the highest-value decision.

What should the decision optimize for? Five objectives, held together: **customer value** (did the person get something useful?), **business value** (did the outcome move, incrementally?), **long-term trust** (will they be more willing to hear from you next time?), **cost**, and **risk** (the worst credible outcome if the system is wrong). Optimizing only business value, usually through a proxy like opens or clicks, is how programs produce rising short-term numbers and falling relationships. Personalization that serves the customer tends to serve the business. The reverse is not guaranteed.

## The standard: specific and true, or honestly general

Here is the sentence this chapter turns on. **Dependable, high-quality personalization depends on an accurate model of the individual person.** Not a first name dropped into a template. A working understanding of who this person is, what they are dealing with, what they are trying to do, and what a message has to look like to belong in their situation.

Shallow personalization can create measurable lift, as the subject-line experiment shows. Depth changes the ceiling and the reliability. A shallow model produces the uncanny near-miss: the email that uses the customer's name and gets their role wrong, leaving the sender worse off than if it had stayed general. When I built systems to do this across hundreds of thousands of people at once, the property that decided whether the whole thing worked was never how impressive one generated output looked. It was whether the average output was relevant enough to earn a real response, and how rarely it said something confidently wrong about the reader's own business. The model is the machine. The message is just its exhaust.

A plain, hand-written template that names a real need will often beat a fluent machine-written message built on a shallow understanding. The reader does not reward eloquence. They reward being understood.

Accuracy does not mean the largest possible profile. More detail makes a message worse when it is irrelevant, stale, or included to prove the sender found it. What matters is decision-level fit: the real constraint, priority, or unfinished problem that makes the message useful now. A fact earns its place in a message by changing what the message should say; facts that only decorate it add length, cost, and risk.

Hence the standard this handbook holds every system to: every personal claim is either specific and true, or honestly general. When evidence is strong and recent, speak to the specifics. When it is thin or old, drop to what is true about the company or the industry, and say nothing you cannot support. A well-written general message is a success. A specific-sounding message built on an invented or outdated detail is a failure even when it reads better. Accuracy is not the goal; it is the floor. Usefulness is the goal, and you cannot be useful and wrong at the same time.

In the governed personalization engine I built for go-to-market teams, this standard is a ladder enforced by the system, not a line in a prompt. As evidence thins, output steps down from person-and-account copy to account level, then segment level, then approved static copy, and finally to sending nothing. The rule the design serves: specificity degrades before trust does. The near-miss it catches most is not an invented fact but an assumed relationship: "your renewal is coming up," "you already use our dispatch module," "your current contract." Unless the relationship is actually recorded, the claim is blocked and the copy is reframed as an opportunity to evaluate. A renewal reminder sent to someone who never bought tells them you do not know who they are.

![A three-tier ladder of specificity: person (speak to the specifics, when evidence is strong and recent), company (what is true about the company), and industry (what is true about the industry, when evidence is thin or old). A green downward arrow says thin or old evidence drops a tier; a crossed-out red upward arrow says never climb a tier by inventing a detail. Below, a well-written general message is marked a success, and invented or outdated specifics are marked a failure even when they read better.](/images/handbook/ch01-honestly-general.svg)

*Figure 1.3. Choose the specificity tier from the evidence; dropping a tier is a success, inventing to climb one is a failure.*

## Known, not watched

Appropriateness is where most personalization anxiety lives, and the research on it is clearer than the anxiety suggests.

Researchers call the tension the personalization paradox: people want to be recognized, and they feel exposed when they learn how much is known. It resolves once you look at *how* the knowledge was obtained. Aguirre and colleagues found in 2015 that personalized ads raised click-through when information was collected overtly and lowered it when collection was covert, because covert collection raised customers' sense of vulnerability [12]. Kim, Barasz, and John found in 2019 that explaining why someone saw an ad helped when the information flow matched their expectations, and hurt when it revealed flows they considered unacceptable, such as tracking across other sites or inferring what they had not shared [13].

The common thread is expectation. Customers do not object to being known. They object to being watched. A Larkspur customer expects Larkspur to know their fleet size and their open support tickets. They do not expect it to know about a personal detail found on another site, or an inference about their life drawn from usage patterns. The same fact can be service in one context and surveillance in another. Interface designers learned the same lesson about expectation decades ago: people accept adaptation they can predict, see the reason for, and correct, and they switch off what moves without warning (Chapter 3).

The most retold example is Target's 2012 pregnancy-prediction score, reported by the *New York Times* [14]. The score is reported fact; the famous anecdote about a father learning of his daughter's pregnancy from coupons is single-source and should not be repeated as confirmed. Either way, the lesson stands: accuracy and appropriateness are separate tests, and passing the first does not pass the second.

One honest line, once. Any system precise enough to serve people well is precise enough to pressure them. The difference is not in the model. It is in the objective and the governance, which is why later chapters treat both as engineering, not policy.

## What this does not do

This is where a chapter like this can lose the right to be believed by exaggerating, so here are the limits.

Personalization does not fix a weak offer. If the product does not solve the customer's problem, a message that describes the problem precisely only makes the gap clearer.

It does not improve every decision. Often the generic treatment is already near the best available, and personalizing it adds cost without lift. Part of a good program is deciding where not to personalize.

It does not scale with data. More data, especially inferred data, raises the risk of a stale, wrong, or inappropriate fact faster than it raises relevance.

The popular claims deserve pushback from both sides. In 2019, Gartner predicted that 80% of marketers who had invested in personalization would abandon it by 2025 for lack of ROI and data problems [2]. It is often quoted as a finding; it was a forecast. Its underlying point holds: programs built on data volume instead of decision value fail on ROI. At the other end, the "80% more likely to purchase" statistics measure intention, not purchases. I do not accept either as a planning number. The planning number is your own lift, on your own decisions, against your own baseline.

## At scale

At a few dozen accounts, a skilled person passes the four tests by judgment. At 250,000 contacts, judgment has to become architecture.

Errors become rates. A system wrong about a customer 2% of the time is wrong 5,000 times in one Larkspur send, so quality is measured on the population, not the demo account. Consistency becomes visible: when marketing, sales, and support each personalize from their own copy of the customer, the customer receives contradictions, which read as not being known at all. And a creepy or false message at scale is not a bad email; it is a brand event. "Specific and true, or honestly general" has to be enforced in code, before and after generation, not left to a prompt. Generative does not mean uncontrolled: use the model where judgment adds value, and code where certainty is available.

Success also splits into more states than a demo shows. In that engine, one record can succeed at generation (an output exists) and still fail on content (incomplete, unsafe, or inappropriate), on persistence (the next process cannot yet read it), or on delivery (the customer never received the right artifact). Above those sits a fifth: across hundreds of records, is the program behaving as agreed? A personalized decision has been made only when all five hold. Chapters 17 and 20 build that machinery; a demo shows only the first.

## Failure story: The Creepy Reveal

Larkspur's AI SDR pilot starts from thin list data (name, title, industry, company size) and adds whatever an ungoverned web search returns for the prospect's name and company. One prospect, a fleet manager at a regional utility, receives an opening line that congratulates her on "your recent move from Denver" and mentions a conference talk she gave three years earlier under her previous employer's name.

Every fact is true. None came from a relationship with Larkspur. She did not expect a vendor she had never spoken to to know where she lived, and the old conference detail felt like digging. She replies with one line asking to be removed and forwards the email to a colleague: "this is why I don't take vendor calls."

The system passed relevance and failed appropriateness. The fix was not better data. It was a rule: personal facts from outside the relationship may inform whether and when to reach out, but are not quoted back unless they are professional, recent, and something the person would expect a vendor to know.

## Patterns

**Four-Test Gate.** *Problem:* teams judge on relevance alone. *Forces:* each test has a different owner (data, strategy, operations, governance). *Solution:* score each decision type on all four tests separately; ship only when all pass. *Tradeoff:* slower first launches, far fewer brand events.

**Value per Decision.** *Problem:* business cases rest on aggregate industry statistics. *Forces:* volume, value, and lift vary enormously by decision. *Solution:* price each decision type as a range with the load-bearing assumption named; fund the highest-value decisions first. *Tradeoff:* needs a baseline and a test before the number is trusted.

**Honestly General.** *Problem:* systems invent specificity when evidence is thin. *Forces:* specific copy demos better. *Solution:* choose the specificity tier (person, company, industry) from the evidence; never climb a tier by inventing. *Tradeoff:* some outputs are plainer; none are false.

## Leader questions

1. Which three repeated decisions about individual customers are worth the most, and are they personalized today?
2. For each personalized program, what is the measured lift over our generic treatment?
3. What would a customer be surprised to learn we know and use? Who decides whether we use it?
4. When our system is unsure about a customer, what does it do?
5. Are we optimizing for a proxy (opens, clicks) or for the outcome and the relationship?

## Build checklist

- [ ] Inventory repeated per-customer decisions before choosing tools.
- [ ] Define the generic baseline for each decision so lift can be measured.
- [ ] Score each decision type on the four tests, one owner per test.
- [ ] Price decisions as value-per-decision ranges, including cost of error.
- [ ] Tag every customer fact with its source and whether the customer would expect you to hold it.
- [ ] Implement evidence-chosen specificity tiers (person, company, industry).
- [ ] Block quoting of out-of-relationship personal facts by default.
- [ ] Block relationship claims (renewal, current product, contract) unless the relationship is recorded; reframe as an opportunity instead.
- [ ] Sample real outputs across the population, not showcase accounts.

## Metrics to watch

- **Incremental lift** per decision type, against a holdout on the generic treatment.
- **Negative-signal rate**: unsubscribes, complaints, "remove me" replies, and spam reports per thousand personalized messages.
- **Factual error rate** in personal claims, from sampled review.
- **Specificity mix**: share of outputs at person, company, and industry level, watched for sudden shifts.

## Reader Q&A

**Is a first name still worth using?** It can lift opens [4], and it is the shallowest personalization there is. Use it where name data is clean; do not mistake it for a program.

**We have little data on most prospects. Should we wait?** No. Start where you hold the most first-party evidence, usually existing customers, and stay honestly general elsewhere.

**How do we know if something is creepy before sending it?** Ask whether the customer would expect you to know the fact, given how they have dealt with you. If the honest answer is "only if they knew what our enrichment vendor does," do not quote it.

## For your AI

```yaml
chapter: 1
title: "What Personalization Is, and What It Is Worth"
concepts:
  - name: Personalization
    definition: "Choosing the next action for one person, using what is known about them, where the choice would differ for someone else."
  - name: Personal data vs. personalized decision
    definition: "Holding customer data is potential; personalization happens only when a decision about one person uses it."
  - name: Four tests
    definition: "A personalized decision must pass relevance, usefulness, timing, and appropriateness, evaluated separately."
  - name: Value per decision
    definition: "Outcome value x lift over the generic baseline, minus cost to decide, minus expected cost of being wrong; expressed as a range."
  - name: Specific and true, or honestly general
    definition: "Every personal claim is either supported by evidence or replaced by a true, less specific statement."
  - name: Known, not watched
    definition: "Customers accept use of facts they expect the company to hold; facts from unexpected flows reduce trust and response."
decision_rules:
  - if: "a personalized decision fails any one of the four tests"
    then: "do not ship it; fix the failing dimension or use the generic treatment"
  - if: "evidence about a person is thin, stale, or single-source"
    then: "drop to company-level or industry-level specificity; never invent to stay specific"
  - if: "a fact was obtained outside the customer relationship and is personal rather than professional"
    then: "may inform whether or when to act; do not quote it to the person"
  - if: "copy asserts a customer relationship (renewal, product in use, current contract) that is not recorded"
    then: "block the claim and reframe as an opportunity or evaluation; specificity degrades before trust does"
  - if: "an output was generated"
    then: "do not count it as done until content, persistence, delivery, and program-level checks also pass"
  - if: "a business case cites aggregate industry statistics"
    then: "replace with value-per-decision ranges and a plan to measure lift against a baseline"
  - if: "a program optimizes opens or clicks only"
    then: "add incremental outcome, negative-signal rate, and error rate before scaling"
assessment_questions:
  - "Which repeated per-customer decisions carry the most value, and which are personalized today?"
  - "What is the generic baseline for each, and has lift been measured against a holdout?"
  - "What customer facts do you use that a customer would not expect you to have?"
  - "What happens in your system when evidence about a customer is thin?"
  - "Which objectives (customer value, business value, trust, cost, risk) does each program track?"
patterns: [Four-Test Gate, Value per Decision, Honestly General]
anti_patterns: [The Creepy Reveal, Volume Over Value, Fake Specificity, Proxy Optimization]
maturity_dimension: understanding
```

## References

1. Arora, N., et al. "The value of getting personalization right, or wrong, is multiplying." McKinsey & Company, 2021-11-12. https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying
2. Gartner. "Gartner Predicts 80% of Marketers Will Abandon Personalization Efforts by 2025." Press release, 2019-12-02. https://www.gartner.com/en/newsroom/press-releases/2019-12-02-gartner-predicts-80--of-marketers-will-abandon-person
3. Gartner. "Gartner Survey Reveals Personalization Can Triple the Likelihood of Customer Regret at Key Journey Points." Press release, 2025-06-03 (survey of 1,464 B2B buyers and consumers, November to December 2024). https://www.gartner.com/en/newsroom/press-releases/2025-06-03-gartner-survey-reveals-personalization-can-triple-the-likelihood-of-customer-regret-at-key-journey-points
4. Sahni, N. S., Wheeler, S. C., and Chintagunta, P. "Personalization in Email Marketing: The Role of Noninformative Advertising Content." *Marketing Science* 37(2), 2018. https://pubsonline.informs.org/doi/10.1287/mksc.2017.1066
5. Salesforce. *State of the Connected Customer*, 5th edition, 2022. https://www.salesforce.com/content/dam/web/en_ie/www/PDF/state-of-connected-customer-fifth-ed-comp.pdf
6. Salesforce. *State of the Connected Customer*, 6th edition, 2023.
7. Epsilon. "New Epsilon research indicates 80% of consumers are more likely to make a purchase when brands offer personalized experiences." Press release, 2018-01-09. https://www.epsilon.com/us/about-us/pressroom/new-epsilon-research-indicates-80-of-consumers-are-more-likely-to-make-a-purchase-when-brands-offer-personalized-experiences
8. Matz, S. C., Kosinski, M., Nave, G., and Stillwell, D. J. "Psychological targeting as an effective approach to digital mass persuasion." *PNAS* 114(48), 2017. https://www.pnas.org/doi/10.1073/pnas.1710966114
9. Eckles, D., Gordon, B. R., and Johnson, G. A. "Field studies of psychologically targeted ads face threats to internal validity." *PNAS* 115(23), 2018. https://www.pnas.org/doi/10.1073/pnas.1805363115
10. Perla, et al.. "The (In)Effectiveness of Psychological Targeting: A Meta-Analytic Review." *Psychology & Marketing*, 2026. https://onlinelibrary.wiley.com/doi/10.1002/mar.70073
11. Twilio Segment. *State of Personalization Report 2024*. https://www.twilio.com/en-us/report/state-of-personalization-report
12. Aguirre, E., Mahr, D., Grewal, D., de Ruyter, K., and Wetzels, M. "Unraveling the Personalization Paradox: The Effect of Information Collection and Trust-Building Strategies on Online Advertisement Effectiveness." *Journal of Retailing* 91(1), 2015. https://www.sciencedirect.com/science/article/abs/pii/S0022435914000669
13. Kim, T., Barasz, K., and John, L. K. "Why Am I Seeing This Ad? The Effect of Ad Transparency on Ad Effectiveness." *Journal of Consumer Research* 45(5), 2019. https://academic.oup.com/jcr/article/45/5/906/4985191
14. Duhigg, C. "How Companies Learn Your Secrets." *The New York Times Magazine*, 2012-02-16. https://www.nytimes.com/2012/02/19/magazine/shopping-habits.html
