Shallow to Deep

Why was deep personalization always limited to a few people, and what does "deep" actually mean?

Working draft, revised September 2026 · 18 min read · Markdown for your AI

Questions this chapter answers

  • Why could we only ever afford to really understand a handful of customers?
  • What is the difference between shallow and deep personalization, in terms my team can act on?
  • Is more depth always better, or can knowing more make things worse?
  • What does it cost to understand one customer now, compared with a person doing it by hand?

The short answer

For most of commercial history, personalization forced a choice. You could understand a few customers deeply, through account managers and advisers, or treat millions shallowly, through merge fields, segments, and scores. Not both, because understanding a person took a skilled person's time, the most expensive input in the building.

Picture the choice as a grid. One axis is depth: how much of a customer's situation a decision reflects. The other is scale: how many customers receive that treatment. Human service lives deep and few; first-era machine personalization, wide and shallow. Deep and wide stayed empty. Shallow and few is where failed pilots tend to live: expensive effort, a small audience, nothing learned.

AI changes the price of the input that kept the deep-and-wide corner empty. Reading an account, interpreting its situation, and drafting something specific now costs cents to about a dollar instead of tens of dollars of skilled time. That moves the frontier; it does not make every decision deep. Depth is a dial: build to the depth each decision requires, grounded in facts that are true, and nothing decorative beyond that.

The line to remember: a fact earns its place only if it changes what should be said.


Two prices, and the distance between them is the subject of this chapter.

The first is the postage on a presorted commercial marketing letter in the United States: between 37 and 47 cents a piece as of July 2026, depending on how finely it is sorted [1]. Add printing and it is still one of the cheapest ways ever devised to reach a stranger, cheap for a simple reason: it is substantially the same letter to everyone.

The second is the price of a skilled person paid to understand one customer. The median US wholesale and manufacturing sales representative earned between $72,080 and $104,920 in May 2025, depending on whether they sell technical products [2]. Wages are about 70% of what a private employer spends on a worker [3], so an hour of that person costs roughly $50 to $72 before tools, management, or unsold hours. Senior account managers cost more.

Attention that was cheap was not personal. Attention that was personal was not cheap.

Larkspur Systems, the running example in this handbook, is a fictional composite. It sells field-service and fleet software to about 40,000 customer accounts. It knows which accounts are renewing, which have an open escalation, which just hired a new operations lead. Its quarterly email says the same thing to all of them. Not laziness: arithmetic.

For most of history, understanding had to be rationed

To write a message that fits one person, someone had to understand that person: their situation, their language, their goal. That took time and attention, both finite. Deep understanding at large scale was not forbidden by any law or policy. It was ruled out by arithmetic.

Peppers and Rogers argued for building relationships "one customer at a time" in 1993 [4], and they were right about the destination. They could not change the price of the trip, so the industry built what it could afford: a few deeply served accounts, everyone else handled by rules. The price of understanding was the policy.

The Shallow-to-Deep grid

The vertical axis is depth of understanding: how far up the depth ladder (next section) a decision reaches. Low: a name or a segment label. High: the customer's situation, goals, and likely next need.

The horizontal axis is scale: how many people receive decisions at that depth. Low: dozens or hundreds, what a human team can hold. High: tens of thousands to millions.

The four quadrants:

Few peopleMany people
DeepThe Craft. Account managers, private bankers, concierge desks, the seller who knows all thirty accounts. Excellent, capped by hours.Deep at Scale. Situation-level understanding for every customer. Historically empty. Reachable now for the first time.
ShallowThe Empty Quadrant. Expensive effort applied to shallow data for a small audience. Neither the economics of scale nor the payoff of depth. Where failed pilots tend to live.The Broadcast. Merge fields, segments, recommenders, propensity scores. Cheap per person, measurable, and blind to situation.

For most of history the achievable options formed a line from top-left to bottom-right. You could slide along it, trading depth for reach, but you could not get above it, because every step up in depth multiplied the human hours per customer. Think of that line as the cost frontier.

A two-by-two grid with depth of understanding on the vertical axis and scale on the horizontal axis. The Craft (deep, few people) is excellent but capped by hours; The Broadcast (shallow, many people) is cheap per person but blind to situation; The Empty Quadrant (shallow, few people) pays Craft prices for Broadcast depth; Deep at Scale (deep, many people) was historically empty. A dashed old cost frontier runs from The Craft down to The Broadcast, and cheap reading moves a new frontier up and to the right so that it passes through Deep at Scale.

Figure 4.1. The Shallow-to-Deep grid: AI moves the cost frontier, not the axes, and Deep at Scale becomes reachable.

Two clarifications. The Broadcast is not stupid: a good recommender reads real behavior and matches items to people at a scale no human could. But it reaches rung three of the ladder below at most; it knows what a customer did, not what they are dealing with. And nobody chooses the Empty Quadrant. Teams arrive by pointing an expensive tool at a pilot group and feeding it the Broadcast's thin data. They pay Craft prices for Broadcast depth.

AI changes the frontier, not the axes. When reading a customer's record costs cents, the line moves up and to the right, and Deep at Scale stops being a contradiction. It becomes an engineering problem: memory, decisions, and governance, which is what the rest of this handbook is about.

What "deep" means: the depth ladder

Here "deep" means how much of the customer's actual situation shapes the decision. Six rungs, on one Larkspur account.

  1. Name. "Hi Jordan." Decides nothing.
  2. Firmographic or segment. Mid-market logistics company, 120 vehicles, Pro plan. Decides which of five templates to send.
  3. Behavior. Dispatch-module usage fell 40% over two months; two support tickets about route sync. Decides timing and topic.
  4. Situation and goals. A new operations lead started six weeks ago; the fleet is expanding into a second region; renewal is in 90 days. Decides what is worth saying.
  5. Reasoning about this customer's problem. The usage drop and the tickets point at one cause: the new region's depots were never configured, so dispatchers route by hand. Decides the recommendation.
  6. Anticipating the next need. A second region usually means a second dispatcher team to onboard before peak season. Decides whether to act now, later, or not at all.

The first era computed rungs one to three from structured data. Rungs four to six lived in call notes, emails, tickets, and an account manager's head. Reaching them meant a person reading: the rationed input.

Two warnings. The ladder measures understanding, not data volume; an account with ten thousand events can sit on rung three. Events record what people did, not what they were trying to do; Chapter 3 shows how field research reaches the rungs telemetry cannot. And confidence falls as you climb: rung six holds the most value and the most confident mistakes.

The depth ladder drawn as six ascending rungs, each with the decision it drives: 1 Name decides nothing; 2 Firmographic or segment decides which of five templates; 3 Behavior decides timing and topic; 4 Situation and goals decides what is worth saying; 5 Reasoning about this customer's problem decides the recommendation; 6 Anticipating the next need decides whether to act now, later, or not at all. Rungs one to three were computed from structured data in the first era; rungs four to six lived in call notes, emails, tickets, and an account manager's head, so they needed a person reading. Confidence falls as you climb, and rung six holds the most value and the most confident mistakes.

Figure 4.2. The depth ladder: each rung decides more, and the top three once required a person reading.

Depth is a dial, not a goal

The grid does not say "move everything to the top-right." Here is the sentence this chapter turns on. One correct fact that changes what should be said is worth more than a hundred facts that only decorate it.

Depth costs money, time, and risk. Detail makes a message worse when it is irrelevant, stale, or included to prove the sender found it. What matters is decision-level fit: the real constraint, priority, timing, or unfinished problem that makes the communication useful now. A model can be wrong in several places and right at the one place where the decision turns.

So set the dial per decision. A password reset needs rung one at most. A strategic renewal needs rung five. A reactivation offer might need rung three and nothing more.

The dial also turns by reader. In the governed personalization engine I built, one account's research compiles into an analytical seller brief and a short customer email. The seller may see a rung-five hypothesis about the account's problem that is never allowed into the buyer's email. A guess shown to a colleague is context; the same guess stated to a customer is a claim.

Customers report the cost of getting it wrong. In a Gartner survey of 1,464 B2B buyers and consumers, fielded in late 2024, 53% reported negative experiences with personalization, and they were 3.2 times more likely to regret a purchase [5]. Gartner locates the misfire where buyers switch tasks and their problem is more complex than the offer. In this frame, that is a depth failure: a rung-three offer to someone living at rung four. (Self-reports; read the direction, not the magnitude.)

A shallow model is not a failed model. It is a model built to the specification the decision required. When the facts to go deeper are missing or unverified, do not invent specificity. Become honestly general.

In the engine I built, honestly general is a ladder the system walks down on purpose: person and account, then account only, then segment, then approved static copy, then fail closed. A lead with a freemail address does not inherit the email provider's firmographics; it gets lower-specificity copy. A headcount that contradicts other evidence is hidden, because a missing number costs less than a confident wrong one. And "your renewal is coming up", a rung-four fact, is blocked unless the relationship is actually known; otherwise the copy is reframed as an evaluation. Specificity degrades before reliability does.

The binding constraint was the cost of understanding

Put numbers on it, as ranges with the assumptions visible.

The human pass. Assume a competent person needs 20 to 60 minutes to read an account's history, tickets, usage, and public news, and draft a first touch that reflects the situation. That range is my working assumption, not a measured figure. At $50 to $72 an hour loaded, one pass costs about $17 to $72 per account. For Larkspur's 40,000 accounts that is 13,000 to 40,000 hours: six to nineteen people doing nothing else for a year, before a single reply. And sellers never had the whole year: in Salesforce's 2022 survey of 7,775 sales professionals, reps reported spending 28% of their time actually selling [6].

The machine pass. As of September 2026, one major provider lists a mid-tier model at $2 per million input tokens and $10 per million output tokens, halves both for batch jobs, and charges $10 per 1,000 web searches [7]. Assume a research run reads 30,000 to 300,000 tokens (record, tickets, notes, a few web pages), writes 3,000 to 15,000, and runs 5 to 20 searches. That is roughly $0.15 to $1 per account at standard prices; a top-tier model costs several times more. For 40,000 accounts: about $6,000 to $40,000 per pass.

A log-scale comparison of the cost to research one account. By hand, 20 to 60 minutes at $50 to $72 an hour costs $17 to $72 per account, or 13,000 to 40,000 hours for 40,000 accounts. By machine, one research run with a mid-tier model at standard prices costs about $0.15 to $1 per account, or about $6,000 to $40,000 per pass over 40,000 accounts. The gap between them is one to two orders of magnitude.

Figure 4.3. Understanding one account costs tens of dollars by hand and cents to a dollar by machine.

The model call is not the whole system; identity, memory, and review cost real money (Chapter 12). But on the rationed input, the gap is one to two orders of magnitude and widening. The price of reaching a fixed capability level has fallen between 9x and 900x per year depending on the task [8]; GPT-3.5-level performance fell from $20 to $0.07 per million tokens between November 2022 and October 2024 [9].

The bigger change is where the money lands. The old model spent the expensive hour first, on a guess. The new model spends cents on every account, watches for a signal, and brings in a skilled person or a stronger model only when a customer engages or an exception needs judgment. The expensive hour is no longer spent on the guess. It is reserved for the signal.

The calculation underneath is one every growth team already runs: what a better decision is worth, times how likely the approach is to produce it, minus what it costs. Spend where that is positive. When the cost term falls fiftyfold, the set of customers worth deep treatment grows quietly. Accounts that never cleared the bar for an hour of human attention now clear it easily.

What this does not do

The easiest way to oversell this chapter is to treat a falling price as a solved problem. Here is what it does not claim.

Cheap reading is not accurate understanding. A one-dollar pass can produce a fluent, wrong account summary. The price of depth fell; the price of verified depth fell less. Chapters 11 and 20 close that gap.

Shallow still works, a little. In randomized field experiments, adding the recipient's name to an email subject line raised open rates from 9.05% to 10.80% and reduced unsubscribes [10]. Rung one is not worthless. It is a low ceiling.

"Deep for everyone" is overreach. AI does not make every customer a strategic account. The Craft still wins where the relationship is the product. The frontier moved; the top-left corner did not disappear.

Failure was never only about depth. In 2019, Gartner predicted that 80% of marketers invested in personalization would abandon it by 2025, citing lack of ROI, the perils of customer data management, or both [11]. A prediction, not a finding; I suspect much of what got abandoned lived in the Empty Quadrant. But the data half is real, and cheaper models do not fix it. Bad identity makes deep personalization confidently wrong about the wrong person.

Falling token prices do not mean falling bills. Research agents consume far more tokens than a chat reply. Budget per decision.

A large design space is not a large experiment. Cheap depth multiplies what a system could say. In an illustrative campaign from the engine I built, 4 personas, 6 industries, 5 regions, 8 message themes, and 3 offers yield 2,880 distinguishable contexts, and a page with 10 zones of 5 eligible options each has 5^10, nearly 9.8 million, possible compositions per context. That is expressive capacity, not experiment volume. Nobody generates or tests those pages, and no cohort is large enough to. The practical move is a few strong hypotheses and learning where the good regions are (Chapter 21).

At scale

At a few hundred accounts, depth is a research problem; at a few hundred thousand, a systems problem. Understanding must be stored, not recomputed (Chapters 10 and 11); in the engine I built, one account's research feeds the page, the email, and the seller brief, so it is paid for once and the surfaces are less likely to contradict each other. Depth must be chosen per decision by policy, not by whoever wrote the prompt (Chapter 14). And errors become rates: a 2% rate of confidently wrong specifics is invisible in a demo and is 800 bad messages across Larkspur's 40,000 accounts. Deep at Scale is not the Craft multiplied; it fails differently.

Failure story: The Empty Quadrant pilot

Larkspur's first AI move was an AI SDR pilot aimed at 300 prospect accounts. The tool was not cheap. The data it received was a name, a title, an industry, and a company size: rung two, plus whatever an ungoverned web search returned for the name and company. The model did what fluent models do with thin data: it wrote warm, specific-sounding paragraphs about challenges "companies like yours" face, and occasionally about challenges the prospect did not have. Reply rates fell below the plain template the team had used before.

The post-mortem blamed the model. The pilot paid for depth, supplied none, and reached too few accounts to learn from. Craft prices, Broadcast depth.

Patterns

Depth to Decision. Problem: teams inject every available fact. Forces: more context feels safer; each fact adds cost and error risk. Solution: per decision, name the required rung and the few facts that would change the output; retrieve and verify those, ignore the rest. Tradeoff: you must define the decision first (Chapter 6).

Wide, Then Deep. Problem: top-tier depth on every account is still too expensive. Forces: signals are uneven. Solution: a cheap pass over everyone to rung three or four; escalate to deeper reasoning or a human only on signal or exception. Put it in configuration, not discipline: in a self-hosted memory system I built, each function and tier routes to its own model (a small one for extraction on every record, a strong one only for the top tier), retrieval has a no-model mode that answers "do we know anything?" before a synthesized brief is paid for, and bulk loads can go through a provider's batch interface at about half price. Tradeoff: the cheap pass sets the ceiling; a missed signal is invisible downstream.

Honestly General. Problem: no verified fact exists at the depth the decision wants. Forces: generation fills gaps fluently. Solution: a defined fallback to the rung the evidence supports, treated as correct behavior. Tradeoff: a lower peak on some messages for far fewer confident errors.

Leader questions

  1. For our five most valuable repeated decisions, what rung of the depth ladder does each one actually need?
  2. Where are we paying Craft prices for Broadcast depth today?
  3. What does one researched account cost us today, by hand and by machine?
  4. When our system lacks a fact, does it fall back to honestly general, or does it guess?

Build checklist

  • Every personalized decision declares its target rung and the facts it depends on.
  • Rung four and above facts carry a source and a date.
  • A cheap wide pass runs on everyone; escalation rules define when deeper work or a human is triggered.
  • Research is stored and reused, and cost is tracked per account and per decision.
  • An honestly general fallback exists for every generated field.
  • Pilots are sized large enough to measure, with a plain-template control.

Metrics to watch

  • Depth reached per decision: share of decisions resting on a verified rung-four-or-higher fact.
  • Unsupported specificity rate: share of personal claims that cannot be traced to a source.
  • Signal-first ratio: share of expensive hours spent after a customer signal, not before.
  • Lift over honestly general: the deep version against the best generic version, not against nothing.

Reader Q&A

Should we skip segments? No. Segments are cheap, explainable, and often right. Add depth where a segment gives the wrong answer for enough customers to matter.

Our AI pilot underperformed the old template. Is this overhyped? Check the quadrant first. Rung-two data for a few thousand people tests the Empty Quadrant, not deep personalization.

How deep is too deep? When a fact would surprise the customer that you know it, or does not change the decision, leave it out. Chapter 19 covers the first case; this chapter the second.

For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 4
title: Shallow to Deep
concepts:
  - name: Shallow-to-Deep grid
    definition: "A 2x2 of depth of understanding (vertical) against scale (horizontal). Quadrants: The Craft (deep, few), The Broadcast (shallow, many), The Empty Quadrant (shallow, few), Deep at Scale (deep, many)."
  - name: Cost frontier
    definition: "The historical line from deep-and-few to shallow-and-many; companies could trade along it but not exceed it because depth required human hours. AI moves it outward."
  - name: Depth ladder
    definition: "Six rungs of understanding a decision can reach: 1 name, 2 firmographic or segment, 3 behavior, 4 situation and goals, 5 reasoning about this customer's problem, 6 anticipating the next need."
  - name: Depth is a dial
    definition: "Set depth per decision to what the decision requires; decision-level fit beats volume of detail."
  - name: Honestly general
    definition: "The correct fallback when facts at the required depth are missing or unverified: a message at the rung the evidence supports. Implemented as a specificity ladder: person and account, account, segment, approved static copy, fail closed."
  - name: Expressive capacity is not experiment volume
    definition: "The combinatorial space of contexts and zone compositions a governed system could produce (for example 2,880 contexts x 5^10 compositions) describes what it could say, not what has been tested; learn through a few strong hypotheses."
  - name: Signal-first spending
    definition: "Spend cheap machine attention on everyone and reserve expensive human or model attention for customers who show a signal or exception."
decision_rules:
  - if: "a decision's output would not change given a fact"
    then: "do not retrieve or include that fact"
  - if: "the decision needs rung 4 or higher and the supporting fact has no source or date"
    then: "fall back to honestly general at the highest verified rung"
  - if: "identity is weak (for example a freemail address) or a firmographic value contradicts other evidence"
    then: "suppress company-specific facts, hide the value, and step down the specificity ladder"
  - if: "an output goes to the customer rather than an internal seller"
    then: "exclude hypotheses and unconfirmed relationship statements; keep them in internal surfaces only"
  - if: "a pilot uses rung 1 to 2 data on fewer accounts than needed to measure lift"
    then: "flag it as an Empty Quadrant pilot; add data depth or scale before judging the approach"
  - if: "expected value of deep treatment for an account (value x likelihood minus cost) is positive at current machine cost"
    then: "include the account in the wide pass and escalate on signal"
  - if: "an action is high-stakes or relationship-defining"
    then: "keep a human (The Craft) in the loop regardless of machine depth"
assessment_questions:
  - "What are your five most valuable repeated customer decisions, and which depth rung does each need?"
  - "Which rungs can your current data support for most customers?"
  - "What does one researched account cost today by hand, and by machine?"
  - "Where do your pilots sit on the grid?"
  - "What happens when a generated message lacks a verified fact?"
patterns: [Depth to Decision, "Wide, Then Deep", Honestly General]
anti_patterns: [The Empty Quadrant Pilot, Decorative Depth, Paying Craft Prices for Broadcast Depth]
maturity_dimension: understanding

References

  1. United States Postal Service, Price List Notice 123, effective 2026-07-12. USPS Marketing Mail commercial automation letters: $0.374 (5-digit, DSCF entry) to $0.467 (mixed AADC). https://pe.usps.com/text/dmm300/notice123.htm
  2. US Bureau of Labor Statistics, Occupational Outlook Handbook, "Wholesale and Manufacturing Sales Representatives" (median wages, May 2025). https://www.bls.gov/ooh/sales/wholesale-and-manufacturing-sales-representatives.htm
  3. US Bureau of Labor Statistics, "Compensation costs for private industry workers averaged $46.89 per hour worked in June 2026," The Economics Daily. https://www.bls.gov/opub/ted/2026/compensation-costs-for-private-industry-workers-averaged-46-89-per-hour-worked-in-june-2026.htm
  4. Peppers, D., and Rogers, M. (1993). The One to One Future: Building Relationships One Customer at a Time. Currency/Doubleday.
  5. Gartner, "Gartner Survey Reveals Personalization Can Triple the Likelihood of Customer Regret at Key Journey Points," 2025-06-03. https://www.gartner.com/en/newsroom/press-releases/2025-06-03-gartner-survey-reveals-personalization-can-triple-the-likelihood-of-customer-regret-at-key-journey-points
  6. Salesforce, State of Sales, 5th edition (2022 survey, n = 7,775). https://www.salesforce.com/news/stories/sales-research-2023/
  7. Anthropic, API pricing (list prices, batch discount, web search), accessed 2026-09-26. https://platform.claude.com/docs/en/about-claude/pricing
  8. Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks," 2025-03-12. https://epoch.ai/data-insights/llm-inference-price-trends
  9. Stanford HAI, AI Index Report 2025. https://hai.stanford.edu/ai-index/2025-ai-index-report
  10. Sahni, N. S., Wheeler, S. C., and Chintagunta, P. (2018). "Personalization in Email Marketing: The Role of Noninformative Advertising Content." Marketing Science 37(2). https://pubsonline.informs.org/doi/abs/10.1287/mksc.2017.1066
  11. Gartner, "Gartner Predicts 80% of Marketers Will Abandon Personalization Efforts by 2025," 2019-12-02. https://www.gartner.com/en/newsroom/press-releases/2019-12-02-gartner-predicts-80--of-marketers-will-abandon-person

This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.

Get chapters by email as they are revised

Prefer LinkedIn? Subscribe to the newsletter there instead.