# 14. The Decision Layer

*The question: What should happen next for this person, including nothing?*

> **Questions this chapter answers**
> - Our models score every customer. Why do the actions they trigger still feel wrong so often?
> - When is "do nothing" the right answer, and how does a system choose it on purpose?
> - Where should an LLM make the call, and where should plain code or a rule make it?
> - How much confidence do we need before the system acts on its own?
> - How do we plan across many steps, many channels, and many teams without burying people in messages?

## The short answer

Most personalization programs are built to answer the wrong question. They ask "how likely is this person to buy, churn, or click?" and then act on the answer. That is a prediction. The question that creates value is different: "what should we do next for this person, and what changes because we did it, compared with doing nothing?" That is a decision, and it needs its own layer in the architecture.

The decision layer sits between memory (what we know about the person) and generation (what we say). It does four things. It lists the actions that are actually available, and that list always includes waiting, asking, routing to a human, and doing nothing. It applies hard constraints (consent, contact limits, open escalations, legal rules) before any model has an opinion. It ranks what remains against objectives that include the customer's long-term trust, not only this quarter's conversion. And it matches the level of autonomy to the consequence of being wrong: cheap, reversible actions can run automatically; irreversible ones get a gate that can say no.

Three findings should shape how a leader funds this. The customers a model scores as most at risk are not reliably the ones an intervention will save, and some retention campaigns measurably increase churn. Sending less can be the better decision: LinkedIn removed four of every ten emails and reported member complaints cut in half. And most decisions do not need a language model; rules and scores are cheaper, faster, and auditable.

If you fund one thing from this chapter, fund a decision layer that can choose silence on purpose, log why it chose it, and learn whether it was right.

## The upsell that should not go out

Larkspur Systems (a fictional company, a composite of many real ones) has a propensity model it is proud of. On a Monday morning it scores Dana, the fleet operations lead at Harlow Fleet Services (also fictional), a regional fleet customer, at 0.81 for the route-optimization add-on: the highest score in her segment. Her usage fits the pattern. Her fleet grew by a third this year. The model is not wrong. She probably would buy it, eventually.

The campaign system reads the score and queues a personalized email. It is well written. It mentions her fleet size and a feature that would save her dispatchers an hour a day.

What the campaign system cannot see is in the support desk. Nine days ago, a sync failure started dropping routes from Dana's dispatch board. Her ticket is marked priority one. She has escalated twice, and her last message ends with "we are evaluating whether this platform is reliable enough for us." Her renewal is in eleven weeks.

Send the upsell and Larkspur tells Dana, in its own voice, that it knows exactly who she is and has chosen to sell to her while her dispatch board is broken. The correct decision for Dana this week is not a better email. It is no email: suppress the campaign, tell her account manager, and put the offer back in the queue for thirty days after the ticket closes, conditional on the fix holding.

The score was accurate. The decision was wrong. That gap is the subject of this chapter.

## A score predicts; it does not decide

A propensity score answers "how likely is this outcome?" A decision needs "how much does this outcome change if we act?" Those are different quantities, and confusing them is the most expensive habit in personalization.

The uplift-modeling literature, which Nicholas Radcliffe helped shape and Gutierrez and Gérardy review, makes the distinction concrete with four groups. **Persuadables** do the desired thing only if you act. **Sure Things** do it either way, so acting wastes money. **Lost Causes** do not do it either way. **Sleeping Dogs** (also called Do-Not-Disturbs) respond badly because you acted. A response model ranks all four groups together, and the customers at the top of a churn score are often the most dissatisfied, which means the list is rich in Sleeping Dogs.

![A two by two grid of what a person does if we act against what they do if we do nothing. Sure Things do it either way, so acting wastes money; Lost Causes do not do it either way; Sleeping Dogs, also called Do-Not-Disturbs, respond badly because we acted; only Persuadables do it only if we act, and they are the people worth targeting. A response model ranks all four groups together, and the top of a churn score, often the most dissatisfied customers, is rich in Sleeping Dogs.](/images/handbook/ch13-uplift-four-groups.svg)

*Figure 14.1. A score finds who is likely to act; only Persuadables are moved by the action.*

This is not a thought experiment. In a paper on mobile operators, Radcliffe and Simpson describe two real retention campaigns. One cut churn in its target group from 30% to 25%. The other *increased* churn from 9% to 10%. Their explanation is plain: a dissatisfied customer is often held in place by inertia, and an intrusive retention call gives the dissatisfaction a moment to crystallize. Their first management implication reads: "Retention programmes can increase churn as well as reduce it."

Eva Ascarza tested the idea with two field experiments, at a wireless provider and a membership organization (*Journal of Marketing Research*, 2018). The customers at highest risk were not the best targets; firms should target on sensitivity to the intervention, not on risk.

Here is the sentence this chapter turns on. **The question is never how likely this person is to act. It is what changes because we acted, compared with doing nothing.** Every other idea in the decision layer follows from taking that sentence seriously.

It has an uncomfortable corollary, the one Chapter 2's Discount Mistake showed: the customer with the highest purchase propensity is often the one who least needs an incentive.

A score tells you who is likely to move. It does not tell you who you will move.

## The action space always includes nothing

Before you can choose well, you have to list what you can choose between. Most systems quietly assume one action (send the message) and optimize only its content and timing. A real decision layer starts with a wider menu:

| Action | What it means | Typical use |
|---|---|---|
| Show | Change what the person sees in a surface they already visit | In-product tips, site content |
| Recommend | Rank options for the person to choose from | Content, products, next steps |
| Send | Push a message into a channel | Email, SMS, direct mail |
| Wait | Hold the action until a condition or date | After a ticket closes; after a trial milestone |
| Ask | Request missing information from the person | Preference, intent, correct contact |
| Route | Hand the situation to a person or a specialist agent | Account manager, support lead |
| Escalate | Route with urgency and authority to override | Risk to the relationship or to compliance |
| Suppress | Block a category of action for this person | Open escalation, legal hold, opt-out |
| Do nothing | Take no action and record that choice | Nothing clears the bar |

Two items deserve emphasis.

**Do nothing is an option with a score, not an absence.** Pega's documentation for its next-best-action engine says the system weighs customer needs against business needs and may "do nothing at all (for example, if no offer is relevant enough to warrant the customer's attention)." The design is simple: give "no action" a value (the expected outcome without intervention, plus the attention you did not spend) and make every other action beat it. When nothing does, the system stays silent and logs why. HCI reached the same rule decades earlier, in Horvitz's mixed-initiative principles and the Office Assistant's failure to follow them (Chapter 3).

The strongest public evidence that silence is a decision comes from LinkedIn. In 2015 the company rebuilt email sending as an optimization problem and announced: "For every 10 emails we used to send, we've removed 4 of them," and "member complaints have been cut in half." The engineering work (Gupta and colleagues, KDD 2016) treats each candidate email as a cost-benefit decision and drops those whose value does not justify the attention. An earlier experiment found 45% more negative responses (unsubscribes and spam reports) among members who received every email than among members from whom a random share was withheld.

**Ask is the most underused action in AI systems.** When evidence is thin, one good question often beats acting on a guess. Language models do not do this by default. Su and Cardie (2026) found that models recognize an ambiguous request when asked to judge it, but rarely ask a clarifying question in normal use, and that retrieved context makes them *less* likely to ask. A model with rich customer memory will be more confident, not more careful. If "ask" matters, make it an explicit option the system can select.

Silence is a decision. Make it on purpose.

## Hard constraints first, objectives second, time horizon always

Every decision mixes three kinds of input, and most failures come from treating them as one.

**Hard constraints** are not traded off. Consent withdrawn, an open priority-one escalation, a legal hold, a contact cap reached: none is weighed against expected revenue. They remove options before ranking begins. "Open P1 escalation suppresses promotional sends" should live in code, run on every path, and be impossible for a model to reason around.

**Objectives** are traded off, and there are always several: value to the customer, value to the business, the cost of the action, the risk of being wrong, and trust. An objective that counts only conversion will find the most efficient way to annoy people, because annoyance is not in the function. Write the tradeoff down, with weights someone signed off on.

**Time horizon** is where good-looking decisions go bad. Rewards arrive in days (an open, a meeting); costs arrive in months (unsubscribes, ignored channels, a renewal that starts from irritation). Hohnhold, O'Brien and Tang (2015) showed Google that users learned to ignore search ads as ad load rose, and the company cut mobile search ad load by half. Chapter 21 covers measuring this. The decision layer's job is to carry a cost for attention even when this week's metric cannot see it.

## Six ways to decide, and when each earns its cost

"AI decisioning" hides several different machines. They have different strengths, different costs, and different ways of failing. Choosing among them is the first design decision, and it is made per decision, not per company.

| Mode | What drives the choice | Strength | Main risk | Use it for |
|---|---|---|---|---|
| Rule | Explicit logic and policy | Transparent, fast, deterministic | Brittle across nuanced cases | Constraints, eligibility, compliance, known branches |
| Data-driven | Scores, probabilities, rankings | Learns patterns from history | Prediction mistaken for intervention value | Ranking candidates at volume |
| Optimized | Objectives plus constraints | Balances competing goals | Bad objectives pursued efficiently | Contact budgets, offer allocation, channel mix |
| Reasoning | Evidence, interpretation, strategy | Handles ambiguity and novel situations | Cost, nondeterminism, false confidence | Ambiguous cases, account strategy, conflicting signals |
| Agentic | Reasoning plus tools, plans, permissions | Executes multi-step work | Errors become actions, not suggestions | Research, multi-step follow-up under guardrails |
| Reflective | Outcome evaluation and feedback | Improves over time | Feedback loops, reward hacking, drift | Tuning policies and weights from observed results |

The table reads left to right as an escalation in capability and in risk. The rule I follow: **use the least powerful mode that handles the decision well, and keep deterministic policy deterministic.**

In the prospecting agents I build, the pattern that survived production is code-orchestrated. Code assembles the context. One model call returns a typed classification (for a reply: positive, negative, question, out-of-office, referral). Code executes the branch. The model never writes to the record or sends anything on its own authority, and opt-out, rate-limit, and sequence-status checks run on every path regardless of what it returned. When someone asks why the system did something, the answer is two artifacts: the JSON the model produced and the branch the code took.

Scoring follows the same split. In a governed personalization engine I built, a model combines weak and strong signals into a score, a band, and a rationale; code checks that they agree, so a score of 88 labeled LOW is flagged rather than shipped.

Reasoning earns its cost in a narrower band than vendors suggest. Dana's case is in it: a high score, a live escalation, a renewal approaching, a thread whose tone matters. A rule catches the escalation; reasoning helps decide what the account manager should do about it. Whether a person who unsubscribed may be emailed has one answer, and it should come from code every time.

Let the model reason; let the code decide what is allowed.

## Candidate, rank, decide

At scale, the decision layer is a pipeline with three stages, each with a different job.

**Candidate generation** assembles every action that could apply to this person now: open campaigns they qualify for, follow-ups owed, service actions triggered by events, and the standing options (wait, ask, route, do nothing). Eligibility rules and hard constraints run here, so ineligible options never reach the ranker.

**Ranking** estimates the value of each remaining candidate for this person, ideally as incremental value over doing nothing, net of cost. Where you have control-group data, rank on predicted uplift, not predicted response. Where you do not, say so, and treat the ranking as a hypothesis.

**Deciding** applies the things a ranker should not own: the per-person contact budget, cross-team arbitration (sales and marketing wanting the same person the same week), confidence thresholds, and the consequence gate described in the next section. The output is one action (or none), a reason, and a trace.

![The decision layer as a three-stage pipeline. Candidate generation lists every action that could apply now (campaigns the person qualifies for, follow-ups owed, event-triggered service actions) plus the standing options wait, ask, route, and do nothing, and hard constraints run at this stage. Ranking estimates each remaining option's incremental value over doing nothing, net of cost, on predicted uplift rather than predicted response where control data exists, and as a labeled hypothesis where it does not. Deciding applies what a ranker should not own (per-person contact budget, cross-team arbitration, confidence thresholds, the consequence gate) and outputs one action or none, a reason, and a trace.](/images/handbook/ch13-candidate-rank-decide.svg)

*Figure 14.2. Candidate, rank, decide: constraints filter, the ranker values, the decision stage governs.*

The cleanest instance I have shipped is duplicate-record resolution in a self-hosted memory system I built. A deterministic finder proposes candidate pairs from hard evidence (a shared strong identifier, or the same name at the same employer; a name alone never qualifies). A model judges one pair and may call exactly one of two tools: merge, or keep distinct. Code owns the rest: a merge below the confidence floor (0.9 by default) is refused, a new judge only proposes until an operator promotes it, and every merge is journaled so it can be reversed for 90 days. The operator approves the judge, not each verdict.

Ranking needs to learn, and learning needs exploration: sometimes trying an option the model is less sure of. The contextual bandit is the classic tool. Li, Chu, Langford and Schapire (2010) framed news recommendation on the Yahoo! Front Page as a sequential decision problem; their LinUCB algorithm, evaluated on more than 33 million logged events, reported a 12.5% click lift over a context-free bandit. The lesson is not the algorithm. It is that exploration is a deliberate, budgeted part of the decision, best spent on low-consequence actions.

## Confidence is judged against consequence

A system that asks "am I confident enough?" without asking "for what?" will be reckless or useless. The same 70% confidence is fine for recommending an article and unacceptable for cancelling a contract.

Picture a grid. The horizontal axis is confidence, low to high. The vertical axis is consequence, from low and reversible (a recommendation, a draft, a flag for review) to high and irreversible (a message reaching a customer at a sensitive moment, a price commitment, a deleted record). The four cells call for four behaviors:

- **Low consequence, high confidence: automate.** Act without review, log the decision, and sample the logs.
- **Low consequence, low confidence: explore or ask.** Cheap mistakes are data. Try the option, or ask the person, and learn from the result. This is where exploration budget belongs.
- **High consequence, high confidence: act through a gate.** A deterministic check sits in the execution path and can refuse. For the most serious actions, a person approves.
- **High consequence, low confidence: do not act.** Gather evidence, ask, or route to a human. The default here is nothing, and it should be the easiest path in the system.

![A two by two grid with confidence on the horizontal axis, low to high, and consequence on the vertical axis, from low and reversible (a recommendation, a draft, a flag for review) to high and irreversible (a price commitment, a deleted record). Low consequence with high confidence: automate, log the decision, and sample the logs. Low consequence with low confidence: explore or ask, because cheap mistakes are data and the exploration budget belongs there. High consequence with high confidence: act through a deterministic gate that can refuse, with a person approving the most serious actions. High consequence with low confidence: do not act; gather evidence, ask, or route to a human, with nothing as the easiest default path.](/images/handbook/ch13-confidence-consequence-grid.svg)

*Figure 14.3. Autonomy is set by confidence relative to consequence, not by confidence alone.*

Reversibility is the axis most programs ignore. They apply uniform oversight, which becomes light oversight everywhere because heavy oversight everywhere is unaffordable. Sort actions by whether they can be undone, and put a real gate (a component in the path that can say no, and that the model cannot argue past) in front of the irreversible ones only. Fewer controls, carrying more weight.

Confidence also sets how specific an action may be. In the personalization engine, every record ends in one of three states: full quality, safe degraded, or fail closed. When identity is weak (a personal email address must not inherit the email provider's firmographics), deterministic policy suppresses company research and steps the copy down toward the segment, without waiting for the model to notice. Chapter 15 names the rungs; the bottom one, fail closed, is the do-nothing option reached because the evidence ran out.

The same logic applies when the system changes its own behavior. An agent that proposes is useful; one that applies its own proposals is a different capability, and the line must be structural. In the memory system, a background consolidation recipe starts propose-only, and an operator promotes it one stage at a time, only after a pass has completed at the current stage. The limit lives in the tool path, not the prompt, so the model cannot talk its way past it, and one call demotes a recipe mid-run. For campaign policy, the design I am building toward adds two human approvals: before a change is implemented and before it reaches production. Chapter 18 covers this in full.

## The decision architecture in ten steps

Every decision, from a one-line rule to a multi-step agent, fits the same ten steps. Writing them down per decision is the most useful hour a team can spend.

1. **Goal.** What outcome are we improving for this person, in this situation? (Dana: keep a strategic account through renewal.)
2. **Policies.** What is allowed, required, prohibited, or subject to approval? (Open P1 suppresses promotion.)
3. **Evidence.** What signals and memory are relevant? (Score 0.81, P1 ticket, two escalations, renewal in eleven weeks.)
4. **Interpretation.** What does the evidence imply, and how uncertain is it? (Buying intent is real; relationship risk is higher.)
5. **Options.** What actions are available, including nothing and asking? (Send, wait, route, do nothing.)
6. **Decision.** Which option best satisfies objectives within constraints? (Route to the account manager; hold the offer.)
7. **Plan.** If multi-step, what sequence follows? (Revisit thirty days after resolution if the fix holds.)
8. **Execution.** Carry it out with the right permissions and channel controls.
9. **Observation.** Capture the result and new evidence. (Ticket closed; tone of next reply.)
10. **Learning.** Judge the outcome and update future behavior. (Did holding the offer change renewal or later conversion?)

Steps 2 and 5 are where most programs cut corners, and they are where Dana's email should have died.

## Planning across steps, channels, and people

A single decision is the easy case. The hard case is a sequence (a six-week onboarding, a quarter-long renewal) or an account four teams all want to talk to.

Three rules keep planning sane. **Plan as a policy, not a script.** A sequence states its goal and exit conditions ("stop if they reply, if a ticket opens, if the champion leaves") and re-decides at each step with fresh evidence, rather than firing step four because step three happened. **Treat attention as a shared budget per person.** Every team and channel spends from the same balance, and the decision layer is the one place that can see it. **Arbitrate centrally.** When two teams want the same person the same week, one decision point chooses against the same objectives.

Contact policy and fatigue are not settings in an email tool. They are decisions the layer makes per person. Internal alerts too: the layer I am designing splits each seller or operator notification into four decisions (whether, what, when, and through which channel), with "not at all" scored like any other option. Chapter 16 works through that notification decision in full, including interruption levels and permission as a budget.

## What this does not do

Decisioning is a field that oversells itself easily, so here are the limits, stated plainly.

**Uplift modeling is harder than the four-group chart suggests.** You cannot observe one person both treated and untreated, so uplift is estimated from randomized control groups and needs far more data than a response model. Many companies lack clean control groups for most actions. Fernández-Loría and Provost offer a useful qualification: accurate causal-effect estimates are not required for good causal decisions, because what matters is ranking people correctly on whether to treat. That helps. It does not remove the need for experiments.

**A next-best-action engine is not a decision layer by itself.** Pega calls its hub an "always-on brain," a single centralized decision authority. Central arbitration is the right idea, and its do-nothing option is exactly right. But an authority is only as good as the evidence it sees. If support tickets never reach it, it will send Dana the upsell with full confidence. The hard part is wiring every signal and every team's actions into it.

**Bandits need volume and fast feedback.** LinUCB learned from tens of millions of events with clicks arriving in seconds. A renewal decision with a quarterly outcome and a few thousand accounts gives a bandit little to learn from. For slow, rare, high-value decisions, careful rules, human judgment, and periodic experiments beat online learning.

**Reasoning does not make decisions safer by default.** It makes them more flexible. Reasoning belongs inside constraints, never in place of them.

**Doing nothing has a cost too.** A customer in trouble who hears nothing may conclude nobody noticed. "Do nothing" must beat the alternatives on the same scale, not win because it avoids blame.

## At scale

With a few thousand customers, a thoughtful team can make many of these decisions by hand. With Larkspur's 250,000 contacts, six teams, and five channels, three things change.

First, **small error rates become large numbers.** A 1% error in a daily decision over 250,000 people is 2,500 wrong actions a day. Consequence tiers and log sampling become essential.

Second, **coordination becomes the main problem.** Each team's logic may be locally sound; the customer experiences the sum. Without one place that sees every pending action for a person, you get contradictions and fatigue no single team caused.

Third, **cost per decision becomes a design constraint.** A reasoning call on every person every day is rarely affordable or necessary. The economical shape is rules and scores for the bulk, reasoning for the cases they flag as ambiguous or high-stakes. If reasoning handles most decisions, the design is probably wrong.

## Failure story: The Fatigue Spiral

Before its decision layer, Larkspur (fictional) ran five programs that each made sense: a weekly nurture email, in-app tips and a monthly feature digest, customer-success check-ins and health alerts, a sales expansion sequence, and event invitations.

Every program had a good open rate. A mid-sized customer's operations lead, whom all five targeted, received eleven messages in two weeks from four senders, two announcing the same feature in different words. She set up a filter. Three weeks later the renewal notice, which actually mattered, went to the same folder.

No single message was bad. Nobody owned the sum. Each team optimized its own channel against its own metric, and the cost (a person who stopped reading Larkspur) landed on a metric none of them tracked. The fix was not better copy. It was a per-person contact budget, one arbitration point, and "do nothing" as a ranked option in every program.

Every message spends from an account whose balance nobody on the sending side can see.

## Patterns

**Do-Nothing Option.** *Problem:* the system acts whenever a candidate exists, so volume grows until people tune out. *Forces:* each team wants its action taken; silence is invisible in most metrics. *Solution:* model "no action" as a scored candidate every action must beat; log each silent decision with its reason. *Tradeoffs:* needs a counterfactual estimate, usually from holdouts; send volume drops before the benefit shows.

**Candidate, Rank, Decide.** *Problem:* one model filters, scores, and governs at once, and governance loses. *Forces:* rankers need freedom to learn; constraints must never bend. *Solution:* constraints in candidate generation; value (uplift where possible) in ranking; budgets, arbitration, and consequence gates in the decision stage. *Tradeoffs:* more components; far easier to debug and audit.

**Consequence-Gated Autonomy.** *Problem:* uniform oversight is unaffordable or too thin. *Forces:* cost pushes toward automation; irreversible mistakes cannot be recovered. *Solution:* tier actions by reversibility; automate the low tier, gate the high tier with a component that can refuse, require human approval for the irreversible, and separate "propose" from "change" with different credentials or operator-promoted stages. *Tradeoffs:* the tier list needs an owner, and approvals must be usable or they get rubber-stamped.

## Leader questions

1. For our three most important personalized actions, what is the "do nothing" option, and how often does the system choose it?
2. Are we targeting people who are likely to act, or people our action will change? How would we know the difference?
3. Which of our automated actions cannot be undone, and what component in the path can stop each one?
4. Who can see every message a single customer will receive from us next week, across all teams?
5. What share of our decisions go to rules, scores, and reasoning models, and what does each share cost?

## Build checklist

- [ ] Every personalized decision has a written ten-step spec, including policies and options.
- [ ] "Wait," "ask," "route," and "do nothing" are explicit, selectable actions, and each silent decision is logged with a reason.
- [ ] Hard constraints run in code before any ranking or model call, on every path.
- [ ] Rankers estimate incremental value where control data exists, and are labeled as hypotheses where it does not.
- [ ] A per-person contact budget is enforced across all teams and channels at one arbitration point.
- [ ] Every action is tiered by reversibility; irreversible actions pass a gate that can refuse.
- [ ] Model outputs are typed; code, not the model, executes actions.
- [ ] Weak identity lowers specificity by deterministic policy; "fail closed" is a logged outcome.
- [ ] Every decision writes a trace: evidence used, options considered, choice, reason, mode, and outcome when known.

## Metrics to watch

1. **Incremental lift over a do-nothing holdout**, per decision type (not response rate).
2. **Suppression and silence rate**: the share of eligible candidates the layer chose not to act on, with reasons. Near zero means the option is not real.
3. **Negative-response rate per person per window**: unsubscribes, spam reports, and channel disengagement, tracked against contact volume.
4. **Decision mix and cost per decision** by mode (rule, score, optimization, reasoning).

## Reader Q&A

**We already have a next-best-action tool. Do we need a decision layer?** Maybe not a new product. Check your tool against the build checklist; the gaps are usually in the wiring (support and sales signals that never reach it), not the engine.

**We have no control groups. Can we still do this?** Start with rules and scores, and start holding out a small random share of eligible people from each program today. Every week you delay a holdout is a week of decisions you cannot evaluate later.

**Should an LLM make our decisions?** Some of them: where interpreting messy, ambiguous evidence changes the answer. Keep eligibility, consent, and contact limits in code, and have the model return a typed choice that code executes.

**How do we set the confidence threshold?** Per consequence tier, not globally. Start strict on irreversible actions and loose on reversible ones, then adjust from the observed error rate in each tier.

## For your AI

```yaml
chapter: 14
title: The Decision Layer
concepts:
  - name: Decision Layer
    definition: "The architectural layer between memory and generation that chooses the next action for one person, including waiting, asking, routing, or doing nothing, under explicit constraints and objectives."
  - name: Prediction vs decision
    definition: "A score estimates how likely an outcome is; a decision requires estimating how much the outcome changes because of an action, compared with no action."
  - name: Uplift groups
    definition: "Persuadables (act only if treated), Sure Things (act anyway), Lost Causes (never act), Sleeping Dogs (react badly to treatment)."
  - name: Do-Nothing Option
    definition: "No action modeled as a scored candidate that every other action must beat; silent decisions are logged with reasons."
  - name: Confidence x Consequence grid
    definition: "Autonomy is set by confidence relative to consequence: automate low-consequence confident actions, explore or ask on low-consequence uncertain ones, gate high-consequence confident ones, do not act on high-consequence uncertain ones."
  - name: Decision modes
    definition: "Rule, data-driven, optimized, reasoning, agentic, reflective; use the least powerful mode that handles the decision well."
  - name: Decision architecture
    definition: "Goal, policies, evidence, interpretation, options, decision, plan, execution, observation, learning."
decision_rules:
  - if: "a candidate action conflicts with a hard constraint (consent, open escalation, legal hold, contact cap)"
    then: "remove it before ranking; never let a model weigh it against value"
  - if: "targeting is based on propensity or risk score alone"
    then: "treat it as a hypothesis; add a randomized holdout and move toward ranking on estimated uplift"
  - if: "no candidate beats the value of doing nothing"
    then: "take no action and log the reason"
  - if: "evidence is thin and the action is low consequence"
    then: "explore or ask the person rather than guess"
  - if: "the action is irreversible or high consequence and confidence is low"
    then: "do not act; gather evidence or route to a human"
  - if: "the decision has one correct answer given known facts (eligibility, opt-out)"
    then: "decide in deterministic code, not with a reasoning model"
  - if: "more than one team has a pending action for the same person in the same window"
    then: "arbitrate at a single decision point against a shared per-person contact budget"
  - if: "identity confidence is low or evidence is thin"
    then: "lower specificity by deterministic policy (full, degraded, fail closed); do not rely on the model to notice"
  - if: "a model returns a score and a band or label that contradict each other"
    then: "flag the record in code before any output ships"
assessment_questions:
  - "Which personalized actions run today, and what is the do-nothing option for each?"
  - "Do you have randomized holdouts for your main programs, and for how long have they run?"
  - "Which signals (support tickets, product usage, sales activity) reach the system that decides outbound actions?"
  - "Which automated actions are irreversible, and what component can stop each one?"
  - "Is there a per-person contact budget shared across teams and channels?"
  - "What share of decisions are made by rules, scores, optimization, and LLM reasoning?"
patterns: [Do-Nothing Option, "Candidate, Rank, Decide", Consequence-Gated Autonomy]
anti_patterns: [The Fatigue Spiral, The Reasoning Everywhere Problem, The Autonomous Error]
maturity_dimension: decisioning
```

## References

1. Radcliffe, N. J., and Simpson, R. (2008). "Identifying who can be saved and who will be driven away by retention activity." *Journal of Telecommunications Management* 1(2). Author PDF: https://stochasticsolutions.com/pdf/SavedAndDrivenAway.pdf
2. Radcliffe, N. J., and Surry, P. D. (2011). "Real-World Uplift Modelling with Significance-Based Uplift Trees." https://www.research.ed.ac.uk/en/publications/real-world-uplift-modelling-with-significance-based-uplift-trees/
3. Ascarza, E. (2018). "Retention Futility: Targeting High-Risk Customers Might Be Ineffective." *Journal of Marketing Research* 55(1). https://journals.sagepub.com/doi/10.1509/jmr.16.0163 ; preprint: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2759170
4. Fernández-Loría, C., and Provost, F. (2021). "Causal Decision Making and Causal Effect Estimation Are Not the Same... and Why It Matters." arXiv:2104.04103. https://arxiv.org/abs/2104.04103
5. Pega Academy. "One-to-one Customer Engagement paradigm." https://academy.pega.com/topic/one-one-customer-engagement-paradigm/v2 (accessed 2026-09-26); "Working with the always-on brain (outbound)." https://academy.pega.com/topic/working-always-brain-outbound/v1 (accessed 2026-09-26)
6. Awan, A. (2015-07-27). "Less is more: You're about to receive less email from LinkedIn." LinkedIn Blog. https://www.linkedin.com/blog/member/product/less-email-from-linkedin
7. Gupta, R., Liang, G., Tseng, H.-P., Holur Vijay, R. K., Chen, X., and Rosales, R. (2016). "Email Volume Optimization at LinkedIn." KDD 2016, 97 to 106. https://dl.acm.org/doi/10.1145/2939672.2939692
8. Gupta, R. "Less Is More: Optimizing Email Volume, Part 1." LinkedIn Engineering Blog. https://www.linkedin.com/blog/engineering/archive/less-is-more-optimizing-email-volume-part-1
9. Su, J., and Cardie, C. (2026). "Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions." arXiv:2605.25284. https://arxiv.org/abs/2605.25284
10. Hohnhold, H., O'Brien, D., and Tang, D. (2015). "Focusing on the Long-term: It's Good for Users and Business." KDD 2015. https://dl.acm.org/doi/10.1145/2783258.2788583
11. Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010). "A Contextual-Bandit Approach to Personalized News Article Recommendation." WWW 2010. https://arxiv.org/abs/1003.0146
12. Gutierrez, P., and Gérardy, J.-Y. (2017). "Causal Inference and Uplift Modelling: A Review of the Literature." PMLR 67. https://proceedings.mlr.press/v67/gutierrez17a.html
