Questions this playbook answers
- Our onboarding is the same kickoff deck for every new account. What should change by role, goal, and experience?
- Our health score is green on accounts that leave and red on accounts that stay. Is the score broken, or are we using it wrong?
- Should we reach out to at-risk accounts before they reach out to us, and when does outreach make things worse?
- How do we walk into a renewal or a QBR already knowing what happened this year?
- What can a customer success agent do on its own, and what must stay with a person?
The moment
Monday, 8:30 a.m. The head of customer success at Larkspur Systems (a fictional company, a composite of many real ones, that sells field-service and fleet software) opens the weekly health review. The dashboard colors every account green, yellow, or red from a blend of logins, seat use, open tickets, and the last survey score. Red accounts get an automated re-engagement sequence and a save call; green accounts get nothing. Six customer success managers (CSMs) carry about 1,200 mid-tier accounts; twelve account managers cover the top 400; everyone else is served by email and in-product messages. Two rows tell the story.
Brookline Landscape Group (also fictional) runs 45 crews. It turned red in December: logins down 60 percent since November. The sequence sends "We miss you! Here are five features you might not know about," and a CSM books a save call. Brookline's owner answers the call from a trade show: they are a landscaping company, it is January, and their crews come back in March. The login drop is the calendar, and it was in the account's own history for four winters running.
Alder Creek Utilities (also fictional) runs 220 vehicles and is deep green: 94 percent of seats active, no open tickets. Its renewal is 90 days out. The executive sponsor who signed the contract left in the summer. The new chief operating officer has asked for a list of every software vendor, with what each one delivered this year. That sentence sits in a CSM's call note. Nothing in the score can read it, and nobody at Larkspur can answer "what did we deliver this year?" without two days of digging.
The score sent a save call to an account that was fine and silence to an account that was deciding. It was asked to do a job a score cannot do.
The decision
Customer success makes a sequence of decisions per account and per user, from the sales handoff to the renewal, and each includes the option to do nothing (Chapter 14).
- Onboard. Which path does each account and user get, given who they are, why they bought, and what they already know?
- Adopt. Which play runs now, if any: training, an in-product guide, a CSM call, or nothing?
- Intervene. When risk appears, does contact help, hurt, or change nothing?
- Renew and review. What must the renewal and the quarterly business review (QBR) show, and who prepares it?
- Hand off. When does this belong to sales or to support, and what travels with it?
Figure CS.1. One account, five decisions, one memory: every stage reads what the last one learned.
Onboarding: personalize by experience first, then by role and goal
The strongest controlled evidence in this function is about onboarding. In a 2011 field experiment at a public cloud infrastructure provider, 366 of 2,673 new customers received proactive guidance on how to use basic features. The treatment halved the number of customers who churned in the first week, cut questions in that week by 19.55 percent, and raised accumulated usage over the next eight months by 46.57 percent. The effect was strongest for customers with less experience with the provider. And the direct effect decayed within one week [1].
The intervention was one kind of guidance, not a personalized program, so the study does not prove that tailoring works. What it shows is heterogeneity: the same help mattered most to people who knew least. That makes prior experience the first personalization variable, ahead of job title. A dispatcher moving from whiteboards needs a different first week from an administrator migrating from a competitor, who needs import mapping and the places Larkspur differs from what she knows. And because the direct effect fades in days while the usage gain compounds for months, the first week is the window.
Role and goal come next: the users who do the daily work, the administrator who configures it, and the sponsor who signed because of a goal ("cut overtime on emergency calls"). Each needs a different path, and the sponsor's goal should shape all of them. In-product slots belong to the Product playbook and self-serve activation to the Software Companies playbook; the CS side is the account plan: kickoff, training, milestones, and the definition of success the customer agrees to.
That definition is the most valuable record customer success writes, and it starts at the sales handoff. If the reason they bought is not captured as a typed goal with its source, the CSM re-asks it at kickoff and the renewal team guesses it a year later.
Adoption plays: a play is a decision with an exit
An adoption play is a decision with five parts: a trigger (a milestone missed, a new user, a goal-linked feature untouched), a target (the user or role who can move it), an action, a success condition, and an exit (including "the account got there without us").
Tie plays to the customer's goal, not to feature breadth. Pendo's often-quoted finding that 80 percent of features are "rarely or never used" comes from its own customers' usage over three months [2]; it says nothing about which features this customer needs.
Usage is still the best signal CS has: in 3,959 subscriptions at a European software provider, adding usage data improved churn prediction over firmographic, transactional, and support features alone [3]. But usage has a shape (seasonal, project-based), and a play that ignores the shape is Brookline's save call.
Health scores describe; they do not decide
Here is the sentence this playbook turns on: a health score tells you how an account looks; it does not tell you whether anything you do will change how it ends.
Start with the evidence. Most health scores are weighted blends designed in a spreadsheet. Gainsight's DEAR framework (deployment, engagement, adoption, ROI) was described by the company's chief customer officer as producing scores with which renewals "aligned almost perfectly," but the published material gives no numbers; its template uses placeholder percentages [4]. I found no peer-reviewed validation of a vendor health-score framework in B2B software. "A model can predict churn" is not "your score does."
Even an accurate score answers the wrong question. A review by twelve retention researchers notes that the customers at highest risk "do not necessarily overlap 100%" with those who should be targeted, that a campaign launched too early can start a customer thinking about churning, and that in a well-known churn-modeling tournament the top decile of predicted risk still contained roughly 96 percent customers who did not churn [5]. Ascarza's field experiments found targeting the highest-risk customers can be ineffective; target on sensitivity to the intervention [6], or, with Lemmens and Gupta, on the incremental profit of intervening, net of its cost [7].
So split what the dashboard blends into three separate things:
- State, with receipts. What is true now: seats active, milestones reached against the success plan, open cases, stakeholder changes, each with a source and a date. No weighting, no color.
- Risk. A calibrated probability of non-renewal, validated against past renewals by cohort, with its reasons.
- Expected response. For each play, including doing nothing, the estimated effect on this account. This is what should choose the action; early on it is judgment backed by experiments, not a model.
Notion's CS team, in a case told on its vendor's blog, redesigned its score per segment, served the long tail with digital CS instead of a CSM, and warned against automated signals without a manual override [8]. The score is an instrument for people, not a trigger for machines.
Figure CS.2. The health score sorts accounts by risk; only the response to contact tells you what to do.
Proactive outreach helps, until it prompts a review
Proactive CS has evidence on both sides. It helped in the cloud onboarding experiment [1]. It hurt at a wireless carrier: Ascarza, Iyengar and Schleicher proactively recommended cost-saving plans to customers, genuinely in their interest, and churn over the next three months rose from 6 percent in the control group to 10 percent among those contacted. The outreach lowered the inertia that kept customers from reconsidering and made them more sensitive to their own usage [9]. Radcliffe and Simpson report a retention campaign that raised churn from 9 to 10 percent for the same reason [10]. And in a telecom field experiment, Becker, Spann and Barrot found proactive post-sales service reduced churn and service calls, until a cross-sell rode along with it and customers began to doubt the motive [11].
The pattern I take from the three: outreach that helps someone do the job they bought you for tends to help; outreach that invites them to re-evaluate the purchase can wake the dog. Onboarding guidance, a fix, training tied to their goal: help. "Let's review your plan," "you may be paying for seats you don't use," an empty check-in to a frustrated account: invitations to review. For Sleeping Dogs, the play is not a better email. It is fixing what made them unhappy, and saying so when it is fixed.
What you need to know
CS needs the relationship over time, not a snapshot. By the Boundary Rule (Chapter 10), what you filter on is a property; what you would quote is evidence:
| What | Form | Why it matters |
|---|---|---|
| Success plan: goal, success criteria, agreed milestones, source (sales call, kickoff) | Typed properties plus the quoted evidence | Onboarding path, adoption plays, QBR, and renewal all read it |
| Onboarding path and prior experience (new to this kind of software, migrating, returning) | Properties | Chooses the first-week path; the heterogeneity in [1] |
| Adoption against the goal, with its seasonal or project shape | Properties computed from telemetry, timestamped | Separates Brookline's winter from a real decline |
| Relationship map: sponsor, admin, champion, users; when each changed | Properties plus graph edges, with validity dates | Alder Creek's sponsor left; the new COO is a new buyer (the Stale Self, Chapter 11) |
| Promises and constraints: "no discount talk," "training before renewal" | Typed constraints (Chapter 18) | Kestrel Haulage's lost promise in Chapter 13 |
| Open and recent cases, linked defects | Properties plus graph edges | Suppresses plays; routes to support |
| Offer history and value delivered, each with a receipt | Append-only entries | Renewal and QBR evidence; stops repeat offers (Playbook S2) |
| Commercial: renewal date, entitlements, tier | Properties | Timing of renewal prep; never changes what the customer is owed |
Stakeholder facts decay fastest: supersede the sponsor when they leave, keeping the old value with its dates (Chapter 11). And a note is not a signal until something reads it: the COO's vendor review in Alder Creek's call note should have become a typed event that moved the account into renewal preparation early.
A person cannot re-read 1,200 accounts' notes every week. A scheduled job can: in the self-hosted memory system I built, recurring prompts run beside the memory, so "each Monday, read last week's notes for every account and propose stakeholder and risk events, citing the sentence behind each" is a configuration, not a project. Proposals stay proposals until a person or a rule accepts them.
The action
Most CS work is asynchronous (Chapter 17): none of it needs to happen in under a second; all of it needs to be right.
Renewal preparation starts 120 days out, from memory
The renewal brief is Larkspur's memory-building bet from Chapter 6, where six CSMs were rebuilding account history by hand before every save call. At 120 days, a job assembles a brief per account: the goal from the success plan and the evidence of progress against it; value delivered, each item linked to its source; stakeholder changes; open risks and promises; what was offered before and what the customer said. The account owner reads it and decides.
Three rules from the governed personalization engine I built carry over. Internal is not customer-facing: the engine's rep-facing playbooks hold hypotheses that must never be copied into a buyer message; a renewal brief can say "sponsor change, possible vendor consolidation, unconfirmed," and the renewal email cannot. Technical completion is not semantic completion: the engine does not accept a playbook because the job reported success; if load-bearing sections are empty, it retries within bounds. A brief with a blank "value delivered" section has failed. Thin data gets a thin brief, labeled: when enrichment misses its deadline, the engine produces a deterministic, lower-specificity brief instead of inventing detail. "No success plan on record" tells the CSM what to ask.
For the forecast, start simple: Fader and Hardie's shifted-beta-geometric model projects renewal curves from a few periods of cohort data, in a spreadsheet [12]. Report retention by cohort, or survivor bias will flatter you [5].
For Alder Creek, the brief would have surfaced the sponsor change in the summer. The right action was not a discount. It was a one-page account of what Larkspur delivered this year, in the new COO's terms, and a meeting before her vendor list was drawn up.
A QBR is the success plan, read back
A good QBR answers three questions: what did you want, what happened, and what next. All three live in memory if the success plan was written at onboarding. Numbers in the deck should be computed by code from telemetry and cited, never generated. In the engine I built, statements like "your renewal is coming up" or "your unused module" are blocked unless that relationship is on the record, by guideline and in code, and otherwise reframed. A QBR draft needs the same guard; it is the document that sounds most like knowing the customer.
The QBR is also where the "ask" action (Chapter 14) belongs: is the goal on record still the goal? Write the answer back to the success plan.
Handoffs carry state, and keep service clean
To sales, for expansion. Hand off when the customer's goal outgrows what they bought (the fleet doubled, a new region), carrying the success plan and evidence, not a lead score (Chapter 13). Keep the conversations separate: proactive service loses its effect when a cross-sell rides along [11]. Expansion decisions belong to Playbook S2.
To support, and back. An open severe case pauses every success play for that account until the fix holds (Playbook SP). When it closes, support hands back what happened and what was promised, so the next CS touch does not apologize for a solved problem. Across functions, the coordinator in Playbook P9 decides what reaches the account this week.
CS agents act inside scoped authority
Agents suit most of the work above: reading notes, assembling briefs, drafting QBR decks, proposing stakeholder events. They should hold no authority over Chapter 18's irreversible list: credits and discounts, renewal pricing and terms, roadmap commitments, contract or consent records, first sends in a new program.
Enforce that in credentials, not prompts. In the memory system I built, each agent can get its own key with a role and a data scope: a reader key's write tools are simply absent, record visibility is enforced inside the database and fails closed, and one account's negotiated exception can be scoped so only those who see that account can read it. Long investigations run as hosted jobs on a credential scoped to the job and revoked when it ends. Governance documents are retrieved on every action, so the pricing rule the agent reads is the one the renewal team signed.
A practical ladder: read and propose from day one; write notes and internal tasks once proposals are mostly accepted; send low-stakes messages (a training reminder, a milestone recap) with sampled review; never commit money or terms. The engine I built treats its own improvement the same way: AI observes, diagnoses, proposes, and validates outside production; production authority stays with people.
What can go wrong
Failure story: The Wake-Up Call
A composite of a pattern that is easy to fall into (illustrative, not a specific deployment). Larkspur's first automated success play targeted every yellow and red account with a friendly, well-written email: "We noticed your team is using 60 percent of its seats. Let's make sure you're on the right plan," with a link to book a review. It was meant as service. Over the next quarter, contacted accounts downgraded seats noticeably more than a small group the team had forgotten to include, and some brought a procurement lead to the review.
The email did what Ascarza's plan recommendations did: it made customers look at what they paid for against what they used, at the moment Larkspur chose, not the moment they would have [9]. The play measured bookings. It should have measured seats and renewals against a holdout. An email that invites a customer to audit you should be sent on purpose, or not at all.
Governance risks
- The Color Trigger. A health color wired straight to an automated sequence. Scores inform people; actions come from the decision layer, with a do-nothing option and a holdout.
- The Seasonal Alarm. Usage compared to last month instead of to the account's own pattern. Brookline's save call.
- The Lost Why. The reason they bought is never captured at handoff, so onboarding, the QBR, and the renewal each guess it.
- Internal in external. A brief's hypothesis, or a relationship claim with no record behind it, reaches customer copy. Declare internal fields prohibited context and guard relationship claims in code (Chapters 15, 18).
- Money in an agent's hands. Discounts come from a catalog, behind an approval, never from an agent's judgment.
What this does not do
Memory and agents do not make a product worth renewing. If Alder Creek's COO finds Larkspur delivered little, a perfect brief makes that visible sooner; it does not change it. What this playbook can do is stop CS from spending scarce attention on the wrong accounts, at the wrong moment.
The evidence has gaps. Beyond Retana and colleagues' one experiment [1], I found no controlled public evidence that personalized onboarding or in-app guides lift retention. "86 percent stay loyal to brands that invest in onboarding" comes from 216 consumers stating an intention, published by a video vendor [13]. And the discipline's origin story, Salesforce's 8 percent monthly churn in 2005, is told by Gainsight executives in their own book [14]: a useful story, not an audit.
How you will know
Every play gets a randomized holdout at the account level (Chapter 21), because users in one account influence each other. The outcomes are renewal, cohort gross revenue retention, and adoption against the success plan, not bookings, replies, or health color.
- For onboarding: randomize new accounts between the tailored path and the best static one; measure first-month milestones, early support questions, and renewal. Retana and colleagues' design is a usable template [1]; check that the effect is largest among the least experienced.
- For risk interventions: three arms where volume allows (the play, a lighter touch, nothing), with downgrades, cancellations, and discount requests among contacted accounts as guardrails. That is where Sleeping Dogs show up.
- For the health score: backtest it against past renewals by cohort. Of accounts scored at 20 percent risk, did about 20 percent leave? What share of the top risk decile renewed anyway? If you cannot answer, the score is a mood, not a measure.
- For renewal preparation: hours per brief, and the share of briefs that surfaced a risk the owner did not know.
Be honest about power: Larkspur's 6,000 at-risk renewals a year, split across plays and arms, leave small cells for an outcome that arrives once a year. Pool across quarters, validate leading indicators against renewal, and label underpowered calls as judgment. And name the population behind any benchmark: the median private SaaS company in KeyBanc and Sapphire's 2024 survey reported net retention around 101 percent and gross around 90 percent [15], while investor frameworks call 120 percent net "best" [16].
Reader Q&A
We have one CSM per 200 accounts. Where do we start? With the handoff and the renewal brief. Capture why each account bought, as a typed goal with its source, from today. Build the 120-day brief from what you already have (tickets, calls, telemetry) for next quarter's renewals. Every other play reads from both.
Should we throw our health score away? No. Split it. Keep the state, with receipts; validate the risk part against past renewals; and stop letting either one trigger outreach on its own. The action comes from the expected response to contact, which you learn from holdouts.
Can the CS agent send the renewal quote? It can draft one from the approved catalog for a person to send. Price and terms are commitments (Chapter 18). As drafts prove reliable, lighten the review of the drafting, not the authority over the price.
For your AIThis playbook's concepts, patterns and checklists as structured data. Paste it into your assistant.
playbook: CS
title: "Customer Success"
question: "How does customer success personalize onboarding, adoption and renewal for each account and each user?"
concepts:
- name: Success plan
definition: "Typed record of why the account bought (goal), what success means (criteria, milestones), and the source of each, captured at the sales handoff and read by onboarding, adoption plays, QBRs, and renewal."
- name: Adoption play
definition: "A decision with a trigger, a target user or role, an action, a success condition, and an exit, tied to the account's goal rather than to feature breadth."
- name: Split health
definition: "Replace one blended health score with three separate things: state with receipts, calibrated risk with reasons, and expected response to each possible action including doing nothing."
- name: Review invitation
definition: "Outreach that prompts a customer to re-evaluate the purchase (plan reviews, seat right-sizing, empty check-ins); the class of contact most likely to wake Sleeping Dogs."
- name: Renewal brief
definition: "Internal, sourced account summary assembled from memory 120 days before renewal: goal and progress, value delivered with receipts, stakeholder changes, risks, promises, and offer history; labeled thin when data is thin."
- name: Scoped CS agent
definition: "An agent whose credentials allow reading and proposing, then writing internal notes and low-stakes sends as trust is earned, and never money, terms, roadmap commitments, or record changes on the irreversible list."
decision_rules:
- if: "a new account is onboarding"
then: "choose the path by prior experience first, then by role and goal; deliver it in the first week"
- if: "the reason an account bought is not recorded at handoff"
then: "capture it as a typed goal with its source before kickoff; do not re-ask the customer what sales already heard"
- if: "an account's usage drops"
then: "compare to the account's own seasonal or project pattern before treating it as risk"
- if: "a health score or color changes"
then: "inform the account owner; do not trigger automated outreach from the score alone"
- if: "an account is high risk and contact is expected to prompt re-evaluation"
then: "do not send a check-in or plan review; fix the cause and report the fix"
- if: "a stakeholder (sponsor, admin, champion) changes"
then: "supersede the relationship record with dates and move the account into renewal preparation"
- if: "an account has an open severe support case"
then: "pause success plays until the fix is confirmed"
- if: "the customer's goal outgrows its entitlement"
then: "hand off to sales with the success plan and evidence; keep the service conversation free of cross-sell"
- if: "a renewal brief or QBR draft has empty load-bearing sections or thin data"
then: "retry within bounds or produce a labeled thin brief; never invent value delivered"
- if: "customer-facing copy states a relationship fact (your renewal, your plan, your unused seats)"
then: "allow only when the fact is on the record; otherwise reframe"
- if: "a CS agent's action involves price, credit, terms, roadmap, or contract records"
then: "the agent drafts from the approved catalog; a person approves and sends"
assessment_questions:
- "Where is the reason each customer bought recorded, and who reads it after the sale?"
- "Has your health score ever been backtested against actual renewals by cohort? What share of the top risk decile renewed anyway?"
- "Which success plays run automatically from a health color, and do any have a holdout?"
- "How long does it take to answer 'what did we deliver to this account this year?'"
- "What credentials do your CS agents hold, and which irreversible actions can they take?"
- "How does an open support case change what CS sends?"
patterns: [Success Plan at Handoff, Experience-First Onboarding, Adoption Play with Exit, Split Health, Renewal Brief at 120 Days, QBR from the Success Plan, State-Carrying Handoff, Scoped Agent Ladder, Customer-Status Guard]
anti_patterns: [The Wake-Up Call, The Color Trigger, The Seasonal Alarm, The Lost Why, The Stale Self, Bookings as Success]
metrics: [renewal_rate_vs_holdout, gross_revenue_retention_by_cohort, milestone_attainment_vs_success_plan, downgrades_among_contacted_vs_holdout, health_score_calibration, renewal_brief_hours_and_new_risk_surfaced]
links: {memory: [10, 11], long_running: 13, decision: 14, timing: 17, generative_experiences: 16, governance: 18, measurement: 21, playbooks: [PR, SW, S2, SP, P9]}
maturity_dimension: decisioningReferences
- Retana, G. F., Forman, C., Wu, D. J. (2016). "Proactive Customer Education, Customer Retention, and Demand for Technology Support: Evidence from a Field Experiment." Manufacturing & Service Operations Management 18(1): 34-50. DOI 10.1287/msom.2015.0547. https://ideas.repec.org/a/inm/ormsom/v18y2016i1p34-50.html
- Pendo (2019). "The 2019 Feature Adoption Report." Vendor data from 615 subscriptions. https://www.pendo.io/resources/the-2019-feature-adoption-report/
- Sanchez Ramirez, J., Coussement, K., De Caigny, A., et al. (2024). "Incorporating usage data for B2B churn prediction modeling." Industrial Marketing Management. https://www.sciencedirect.com/science/article/abs/pii/S0019850124000865
- Capote, K. (2022-12-06). "Use customer health data to grow and forecast NRR." TechCrunch. https://techcrunch.com/2022/12/06/use-customer-health-data-to-grow-and-forecast-nrr/ ; Gainsight DEAR ebook https://www.gainsight.com/wp-content/uploads/2023/04/DEAR-Ebook.pdf
- Ascarza, E., Neslin, S., Netzer, O., et al. (2018). "In Pursuit of Enhanced Customer Retention Management: Review, Key Issues, and Future Directions." Customer Needs and Solutions 5(1-2): 65-81. DOI 10.1007/s40547-017-0080-0. https://business.columbia.edu/sites/default/files-efs/pubfiles/25780/ascarza_et_al_choice_symposium.pdf
- Ascarza, E. (2018). "Retention Futility: Targeting High-Risk Customers Might Be Ineffective." Journal of Marketing Research 55(1). https://journals.sagepub.com/doi/10.1509/jmr.16.0163
- Lemmens, A., Gupta, S. (2020). "Managing Churn to Maximize Profits." Marketing Science 39(5): 956-973. DOI 10.1287/mksc.2020.1229
- Van Lew, T. (2023-06-13). "How Notion Approaches Customer Health Scoring for All Segments." Gainsight blog (vendor-told case). https://www.gainsight.com/blog/how-notion-approaches-customer-health-scoring-for-all-segments/
- Ascarza, E., Iyengar, R., Schleicher, M. (2016). "The Perils of Proactive Churn Prevention Using Plan Recommendations: Evidence from a Field Experiment." Journal of Marketing Research 53(1): 46-60. DOI 10.1509/jmr.13.0483. https://journals.sagepub.com/doi/abs/10.1509/jmr.13.0483
- Radcliffe, N. J., Simpson, R. (2008). "Identifying who can be saved and who will be driven away by retention activity." Journal of Telecommunications Management 1(2). https://stochasticsolutions.com/pdf/SavedAndDrivenAway.pdf
- Becker, J. U., Spann, M., Barrot, C. (2020). "Impact of Proactive Postsales Service and Cross-Selling Activities on Customer Churn and Service Calls." Journal of Service Research 23(1): 53-69. DOI 10.1177/1094670519883347
- Fader, P. S., Hardie, B. G. S. (2007). "How to Project Customer Retention." Journal of Interactive Marketing 21(1): 76-90. DOI 10.1002/dir.20074
- Wyzowl (2020). "Customer Onboarding Statistics." Survey of 216 consumers. https://wyzowl.com/customer-onboarding-statistics/
- Mehta, N., Steinman, D., Murphy, L. (2016). Customer Success: How Innovative Companies Are Reducing Churn and Growing Recurring Revenue. Wiley. Chapter 1 excerpt: https://catalogimages.wiley.com/images/db/pdf/9781119167969.excerpt.pdf
- KeyBanc Capital Markets and Sapphire Ventures (2024-10-23). 15th annual Private SaaS Company Survey, press release. https://sapphireventures.com/press/keybanc-capital-markets-and-sapphire-ventures-private-saas-company-survey-reveals-a-continued-focus-on-operational-efficiency-and-profitability/
- Bessemer Venture Partners (2023-04-11). "State of the Cloud 2023." https://www.bvp.com/atlas/state-of-the-cloud-2023
This playbook is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised