Questions this playbook answers
- Our onboarding is one checklist for everyone. Should it be personalized, and how far?
- Which parts of the interface should change per user, and which should never move?
- How do we personalize inside the product without adding one more stream of interruptions?
- How do we prove an adaptive product works better than a well-designed static one?
- We hold thousands of calls, tickets, and usage traces. How does that become evidence for the roadmap instead of the loudest request?
The moment
On the same Monday last winter, two new users log in to Larkspur Systems for the first time. (Larkspur is the fictional field-service and fleet software company that runs through this handbook: about 40,000 customer accounts, product telemetry, a support desk, ten years of CRM history.)
Marisol is the only dispatcher at a twelve-van plumbing company that signed on Friday. She has never used dispatch software, and she has eleven jobs to route before 8 a.m.
Tom is a fleet administrator at a 400-seat account migrating from a competitor. His team has already imported 380 routes, and in the sales call his director said, in passing, that the old system's weak point was offline mode at rural job sites.
Both see the same nine-step checklist. Step one is "Create your first route," which assumes Marisol knows what a route template is. Tom needs none of it. He dismisses the checklist, then the tooltip tour, then the "Did you know?" panel, and goes looking for offline settings, which live three menus deep. By Wednesday he has opened a ticket. By Friday, the product team is in its winter planning meeting, months before the spring release that would rebuild offline mode, ranking feature requests by votes on the ideas board. Offline mode has eleven votes. The phrase "offline" appears in 214 call transcripts and support tickets from the last two quarters, across 96 accounts, and nobody in the meeting knows it. (All Larkspur figures here are illustrative.)
Larkspur knew enough to treat the two users differently, and more, in aggregate, about what to build next. It used neither.
The decision
In marketing and sales, the unit of personalization is a message. In the product, it is a slot: a place in the interface where the right content depends on who is looking. The onboarding step, the empty state, the default view, the suggested next action, the help panel, the order of items in a "recommended" row. The decision, per user and per slot, is what fills it now.
The action space is wider than "show a tip," and it follows the menu in Chapter 14:
- Show or reorder the contents of a designated slot (the onboarding path, a recommended row).
- Prefill a default (a route template sized for a twelve-van fleet; offline sync on for accounts that work rural sites).
- Defer a feature until it is useful (route optimization after the first week of routes, not on day one).
- Ask one question when the evidence is thin ("How many vehicles do you dispatch?").
- Route to a person (the onboarding specialist, when a 400-seat migration stalls).
- Stay out of the way. No tip, no panel, no tour. For Tom, most of the week.
Here is the sentence this playbook turns on: in the product, personalization is judged by the work the user finished, not by the suggestion they clicked. A product is the one channel where the customer is already doing the work. Every other channel has to earn attention; the product has to avoid spending it. A tooltip is a message too, and it draws on the same per-person attention budget as the nurture email and the account manager's check-in.
The second half of the playbook's question is a different decision, made by the team rather than the system: which problems to solve for everyone. The same memory that tells the product that Tom needs offline settings tells the roadmap that 96 accounts do.
What you need to know
Product personalization is unusual in one way: its best signal is generated by the product itself. It is also usually B2B-shaped, so memory has two levels.
The account. What the customer bought and configured: plan and entitlements, integrations, fleet size, imported data, whether this is a migration. What they have said: the sales call that mentioned rural sites, open tickets, the renewal date. This comes from the CRM, billing, support, and call notes, resolved to one account (Chapter 9).
The user. Role (dispatcher, technician, administrator, manager), what they have done, and what they have declared. Telemetry becomes typed properties: first_route_dispatched_at, routes_created_last_7d, offline_mode_enabled, tour_dismissed. Following the Boundary Rule in Chapter 10, anything you would filter or count on is a property. "Said on the kickoff call that her technicians lose signal in the valley" is an evidence-level memory, with a source and a date.
Three signals carry most of the weight, in this order.
- What the user declared. For a new user there is no history to model, which is the cold-start problem Chapter 2 describes. The cheapest fix is to ask one or two questions whose answers change the path: role, and fleet size or job type. Declared data is also the most defensible data to use (Chapter 7). Spotify offers a telling example from the other end of a product's life. After ten years of Discover Weekly, a playlist built entirely on inferred taste that had passed 100 billion streams, it added a control letting listeners refresh the playlist by choosing from up to five genres, themselves personalized from listening history [1]. Mature inference did not remove the need to let people say what they want.
- What the user did. Telemetry is the strongest behavioral signal most companies own, and it is first-party and fresh. It is also narrow: it records the click, not the reason.
- What the account said. Calls and tickets hold the reason. "Offline mode fails at rural job sites" explains why Tom goes looking for settings, and warns the product not to promote a feature that is broken for this account (the product-side version of Chapter 14's upsell that should not have gone out).
Freshness matters more here than anywhere, because product state changes by the minute. A checklist that asks a user to do what they finished yesterday reads as a system that is not paying attention. Give each property a freshness policy (Chapter 11), and let the event stream, not a nightly batch, update the ones that drive onboarding.
The action
Adapt the contents, never the map
The oldest evidence on adaptive interfaces is a warning. In controlled lab studies at the University of British Columbia, Findlater and McGrenere compared pull-down menus whose top section was static, adaptable (the user chose what went there), or adaptive (the system reordered by frequency and recency). The static menu was significantly faster than the adaptive one, and most participants preferred the adaptable version [2]. People build a map of an interface, and a system that rearranges the map makes them slower even when its predictions are good.
The result is two decades old and about menus, so treat it as a caution, not a law. It points to the design the closing chapter reaches from the other direction: a stable frame with variable parts (Closing). Navigation, menu order, and the location of controls stay fixed. Designated slots vary. Users get a visible way to change what the system chose, which is Spotify's lesson again. Chapter 3 gives the fuller research record, including where accurate adaptation wins, and Chapter 16 extends the frame to generated pages, screens, and notifications.
A governed personalization engine I built uses the same shape for pages: a static shell with roughly 8 to 12 text zones, never navigation or legal copy. Each zone has a contract (purpose, allowed and forbidden context, length limit, checks, approved fallback). A product slot needs one too.
Adapt the contents, never the map.
Figure PR.1. Stable frame, variable parts: the same map for everyone, with only designated slots adapting to the person.
Personalize the core loop, not the chrome around it
The strongest public examples of product personalization change the work itself, not the tips about it.
Duolingo's Birdbrain model estimates both how much each learner knows and how hard each exercise is, and the lesson generator uses those estimates to build sessions at the right difficulty for that learner. By October 2020, more than 20% of lessons were personalized this way, and the company reported that its A/B tests showed learners who got Birdbrain-built lessons learned more, did more lessons, and were more likely to return day after day [3]. Too hard and the learner quits in frustration; too easy and the lesson teaches nothing. The personalization lives exactly where the product delivers its value.
Netflix personalizes the artwork for each title per member: the same film can show its romantic leads or its comedian, depending on what that member watches. The team treated it as a contextual bandit, evaluated policies offline by replay before exposing members, and looked at "the quality of engagement to avoid learning a model that recommends 'clickbait' images: ones that entice a member to start playing but ultimately result in low-quality engagement" [4]. The artwork is evidence, offered at the point of choice, of why this title might fit this person, and it is held to what happens after the click.
Netflix's 2025 paper on the value of its recommender adds a useful fact about where product personalization pays. Replacing the production system with simpler baselines reduced engagement by roughly 4% (matrix factorization) to 12% (popularity), and most of the gain came from targeting, with the largest effects for titles of middling popularity [5]. Everyone already finds the hits. Personalization earns its keep on the long middle of a product, the features and content that are right for some users and invisible to them.
For Larkspur, the core loop is dispatch. Marisol's first week is not improved by a better tour; it is improved by a route template prefilled for twelve vans and plumbing job types, and by route optimization appearing after she has dispatched a week of routes by hand and has something to optimize.
Where generation fits, and where it does not
Language models make two product moves newly cheap. Conversational onboarding (the product asks what the user is trying to do and answers in their terms) ships now (Chapter 5). Generated interface does not, at product speed: Google's own generative UI researchers report generation can take more than a minute, with occasional inaccuracies, and raters still preferred pages built by human experts [6]. Against a budget of about 0.1 second to feel instant (Chapter 17), fully generated screens are a lab demo as of September 2026 (Chapter 16's Generative Experience Ladder). An in-app notification is a decision too: whether, what, when, and how loudly, drawn from the same per-person budget as email (Chapter 16).
The workable shape today: precompute each user's slot contents on events (signup, role declared, first route dispatched, ticket opened) and let the live request only select among them (Precompute-then-serve, Chapter 17). Generated text (a help answer, an empty-state explanation) goes into typed zones, and every claim about the user's own data ("your team dispatched 40 routes last week") is grounded in telemetry or left out (Chapter 15). Render only components from a governed catalog. Specific and true, or honestly general.
Generated text is untrusted input to the renderer. In that engine every value is coerced to its zone's type and escaped, and a javascript: link is refused however confidently a model returned it. A failed check renders the fallback; thin evidence steps the slot down (user, account, segment, static default), so specificity degrades before trust does.
Expect memory to say no, too. The self-hosted memory system I built refuses work it cannot run, with a retry-after signal, rather than falling behind, and budgets reads apart from writes so a burst of saves does not stop lookups. On the interactive path a refusal is ordinary: render the precomputed default and never block the task.
Memory to roadmap
The second output of this playbook is not shown to users. It is a quarterly (or continuous) roll-up of what the memory knows about problems, delivered to the product team.
Votes on an ideas board measure who visits the ideas board. Calls, tickets, onboarding answers, and usage traces measure the customer base. The roll-up has four rules:
- Count accounts, not mentions. One angry administrator who writes nine tickets is one account.
- Separate the job from the request. "Add a sync button" is a request; "technicians lose work when they lose signal" is the job. Cluster by job.
- Keep the provenance. Every theme links to the memories behind it (account, date, source, quote), so a product manager can read the twenty best examples before believing the number. Background curation (the dreaming cycle in Chapter 11) can maintain these theme summaries while the system is idle. In my memory system, a new consolidation recipe only proposes until an operator promotes it, every write is stamped and ledgered, corrections overwrite in place with history kept, and curation queues behind live requests, so tidying memory never slows the product.
- Read the personalization logs as product data. If the same adaptation fires for most users (say, 80% of dispatchers hide the analytics widget in week one), it is not personalization. It is a bug report about your default. Chapter 6 makes the same point from the other side: personalization is not a fix for a bad default.
Figure PR.2. The memory that adapts the product for one user also counts, by account, what the roadmap should fix for everyone.
What can go wrong
Failure story: The Shifting Map
Larkspur's first adaptive feature was a dispatch-board toolbar that reordered itself by each dispatcher's frequency of use. Offline tests showed it predicted the next click well. In production, experienced dispatchers began mis-clicking during the 7 a.m. rush: "Reassign" had moved to where "Hold" used to be. Tickets about "jobs disappearing" rose, and the support team could not reproduce them, because each dispatcher's toolbar was different and nobody had logged what the user saw. The fix was to freeze the toolbar, move the adaptation into a "suggested for this job" slot beside it, and log every rendered variant with the user and time. Findlater and McGrenere could have predicted the first half. The logging gap was Larkspur's own. (Fictional composite.)
A personalized screen the support team cannot see is a bug nobody can reproduce.
Governance risks specific to the product
Upsell inside the workflow. An in-app prompt for a paid add-on is a sales action in a work surface. It obeys the same hard constraints as an email (suppressed during an open escalation, counted against the contact budget) and never blocks a task (Chapter 14).
Dark patterns. The product is where manipulative design is easiest to build and hardest to see. The FTC's 2022 staff report catalogued the common forms: disguised ads, hard-to-cancel subscriptions, buried terms, and tricks that extract data [7]. A personalized product must not learn them. Keep "cancel," "downgrade," and "turn off" out of the adaptive layer entirely: fixed location, fixed wording, for everyone.
Leakage between users in one account. B2B memory is shared at the account level; product surfaces are seen by individuals. An administrator may see team activity; a technician should not see a colleague's performance reflected in "people like you" suggestions. Scope what each role's surfaces may read, and test it (Chapter 18). Enforce it below the application where you can. In my memory system, each person and agent gets a scoped key, row-level security in the database fails closed, and an in-product assistant connected over MCP is offered only its role's tools: a read-only assistant has no write tools to misuse.
Employee data used on employees. Fleet telemetry includes technician location and timing. Routing jobs with it is the product's purpose; commenting on a technician's pace on their own screen is a different purpose, and it will feel like monitoring. Run the appropriateness check in Chapter 19 before any telemetry field drives a user-facing message. The same purpose test applies to the roadmap roll-up: check contracts and data-processing terms before customer conversations feed product analysis, and aggregate before sharing. (This is design guidance, not legal advice.)
What this does not do
Evidence on onboarding personalization is thin, and I want to be plain about it. Vendor posts claim large activation lifts from personalized onboarding; in this research pass I could not trace one to a controlled study with a published method. The academic work I found is small: one randomized online experiment with 150 participants found personalization cues amplified the effect of commitment cues on intention to use a mobile app, which is a stated intention, not retention [8]. The public examples with real experiments (Duolingo, Netflix) personalize the core product, not the onboarding tour. The Software Companies and Customer Success playbooks reach the same verdict; the closest controlled evidence they cite, a cloud provider's proactive guidance for new customers, tested help, not personalization.
Personalization also cannot rescue a product that is hard for everyone. If most new users stall at the same step, redesign the step.
How you will know
Pick an outcome that is the job, not the click. For Larkspur: time to first dispatched route; seats active in week four; jobs completed offline at rural accounts. Guardrails: tickets per new account, dismiss and undo rates on adaptive slots, and task time for experienced users, who pay first for a moving interface.
Randomize at the unit of influence. Users in one account share configuration and train each other, so account-level assignment is cleaner for onboarding and defaults (Chapter 21). Individual slot content (which recommended item, which help answer) can be tested per user.
Do the power arithmetic before launch. Suppose 40% of new accounts dispatch a route in week one and you want to detect a rise to 44%. With 80% power at 5% significance, that needs roughly 2,400 accounts per group. If Larkspur signs about 500 new accounts a quarter (illustrative), a 50/50 account-level test takes more than two years. The honest options from Chapter 21 apply: test a more frequent outcome, test per user where influence does not travel, pool quarters, or label the decision as judgment.
Do not confuse the design space with the test plan. Ten slots with five eligible options each allow 5^10, about 9.8 million screens per user context. That number measures what the system can express, not what anyone has tested. The practical program is a few strong hypotheses, tested where the arithmetic above allows.
Use bandits inside slots, holdouts around the program. A contextual bandit is well suited to choosing slot contents, as Netflix did with artwork, provided its reward is engagement quality, not the click, and it keeps an exploration floor so it does not become the Feedback Echo of Chapter 21. Whether adaptive onboarding beats the static version is a holdout question: keep a small share of new accounts on the best static flow, permanently.
Log what each user saw. Every rendered variant and the decision that chose it goes in the decision trace: support reproduces problems from it, analysts reweight biased data with it, and the roll-up learns from it which defaults are wrong. Review the rendered screen, not only the stored values: a stale template or cache can render correct data wrong.
Measure the roll-up by what it changed: how often a theme's evidence changed a priority, and whether the shipped fix moved the job-level outcome for the accounts that raised it.
Reader Q&A
Should we ask users questions at signup, or infer everything? Ask one or two whose answers change the path, then infer. A question that changes nothing is friction; a question that saves a wrong week is service. Store the answer as declared data with a date, and let behavior revise it.
Should an LLM generate our interface? Not the frame, and not yet at interaction speed. Let models write into typed zones and choose among governed components; precompute what a waiting user needs.
We are sales-led with few users per account. Does this apply? Yes, with the weight shifted to the account: migration status, integrations, and what was said in the sales process are richer signals than a thin usage history.
Should product memory run in our own infrastructure? It is a choice, not a requirement. My memory system runs as one container on your own Postgres with pgvector (any cloud, on-prem, or air-gapped), with a model per function, local ones included. You then own backups, scaling, and a database that is a single point of failure unless made highly available. Self-hosting alone is not compliance (Chapter 19).
Who owns the roadmap roll-up: product, research, or customer success? Product owns the decision; the roll-up is a shared service over the same memory the other functions read. The Research and Strategy playbook covers rolling individual insight into market and ICP strategy; this one covers product priorities.
For your AIThis playbook's concepts, patterns and checklists as structured data. Paste it into your assistant.
playbook: PR
title: Product
question: "How does the product adapt to each user, and what does the team learn?"
concepts:
- name: Slot
definition: "A designated place in the product interface whose contents depend on who is looking; the unit of in-product personalization."
- name: Stable Frame, Variable Parts
definition: "Navigation, menu order, and control locations stay fixed; only designated slots adapt, and users can override."
- name: Core-Loop Personalization
definition: "Adapting the work the product delivers (difficulty, defaults, templates, ranking) rather than the tips and tours around it."
- name: Memory-to-Roadmap Roll-up
definition: "Aggregating account and user memories into job-level themes, counted by account, with provenance to source evidence."
- name: Slot Contract
definition: "Per-slot purpose, allowed and prohibited context, length limit, checks, and an approved fallback; generated values are coerced, escaped, and link-checked before render."
- name: Default Signal
definition: "When one adaptation fires for most users, it indicates a wrong default for everyone, not a personalization opportunity."
decision_rules:
- if: "an adaptation would move navigation, menu order, or the location of a control"
then: "do not adapt it; move the adaptation into a designated slot beside the stable control"
- if: "a new user has no usage history"
then: "ask one or two questions whose answers change the path; store answers as declared data with a date"
- if: "an in-product prompt promotes a paid feature"
then: "apply the same hard constraints and contact budget as outbound messages; suppress during open high-priority escalations"
- if: "a feature has an open defect or ticket for this account"
then: "do not promote it to that account's users"
- if: "the same adaptation is applied to most users of a role"
then: "change the default for everyone and report it to the product team"
- if: "a surface is on the interactive path (about 0.1 s budget)"
then: "precompute slot contents on events; select live; never generate the frame at request time"
- if: "a memory read is refused or times out on the interactive path, or a generated value fails its slot checks"
then: "render the precomputed or approved static default; never block the user's task"
- if: "generated text makes a claim about the user's own data"
then: "ground it in telemetry or memory with provenance, or render the honestly general version"
- if: "a control concerns cancel, downgrade, consent, or turning a feature off"
then: "exclude it from the adaptive layer; fixed location and wording for all users"
- if: "customer conversations feed product analysis"
then: "check contracts and purpose, aggregate by account, keep links to evidence"
assessment_questions:
- "Which slots in your product could vary per user, and which controls must never move?"
- "What does your onboarding ask, and does each answer change what the user sees?"
- "Which telemetry events are stored as typed, fresh properties that surfaces can read?"
- "Can support see exactly what a given user was shown, and why?"
- "Is there a static holdout for adaptive onboarding, randomized by account?"
- "How does evidence from calls, tickets, and usage reach roadmap prioritization today, and is it counted by account?"
- "Which in-product prompts are commercial, and do they share a contact budget with outbound channels?"
patterns: [Stable Frame Variable Parts, Core-Loop Personalization, Precompute-then-serve, Memory-to-Roadmap Roll-up, Governed Component Catalog, Specificity Ladder, Do-Nothing Option]
anti_patterns: [The Shifting Map, The Tour Nobody Needed, Upsell in the Workflow, Votes as Evidence, Design Space as Test Volume, The Proxy Trap, The Feedback Echo]
metrics: [time to first value by role, week-4 active seats, dismiss and undo rate per slot, task time for experienced users, tickets per new account, share of roadmap priorities changed by evidence]
links: {cold_start: 2, data_types: 7, identity: 9, memory: 10, freshness: 11, decision_layer: 14, generation: 15, latency: 17, governance: 18, privacy: 19, measurement: 21, generative_experiences: 16, forecast: closing}
maturity_dimension: experience_generationReferences
- Spotify Newsroom (2025-06-30). "Discover Weekly Turns 10: Celebrating 100 Billion Tracks Streamed and a Decade of Personalized Discovery." https://newsroom.spotify.com/2025-06-30/discover-weekly-turns-10-celebrating-100-billion-tracks-streamed-and-a-decade-of-personalized-discovery/
- Findlater, L., and McGrenere, J. (2004). "A Comparison of Static, Adaptive, and Adaptable Menus." CHI 2004. https://dl.acm.org/doi/10.1145/985692.985704 ; thesis version: Findlater, L. (2004), "Comparing Static, Adaptable, and Adaptive Menus," University of British Columbia. https://www.cs.ubc.ca/labs/imager/th/2004/Findlater2004/Findlater2004.pdf
- Duolingo Blog (2020-10-07). "Learning how to help you learn: Introducing Birdbrain!" https://blog.duolingo.com/learning-how-to-help-you-learn-introducing-birdbrain
- Chandrashekar, A., Amat, F., Basilico, J., and Jebara, T. (2017-12). "Artwork Personalization at Netflix." Netflix Technology Blog. https://netflixtechblog.com/artwork-personalization-c589f074ad76 ; RecSys 2018 talk: https://dl.acm.org/doi/10.1145/3240323.3241729
- Zielnicki, K. et al. (2025, rev. 2026). "The Value of Personalized Recommendations: Evidence from Netflix." arXiv:2511.07280. https://arxiv.org/abs/2511.07280
- Leviathan, Y., Valevski, D., Natchu, V., and Matias, Y. (2025-11-18). "Generative UI: A rich, custom, visual interactive user experience for any prompt." Google Research. https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/
- Federal Trade Commission (2022-09). "Bringing Dark Patterns to Light." Staff report. https://www.ftc.gov/system/files/ftc_gov/pdf/P214800+Dark+Patterns+Report+9.14.2022+-+FINAL.pdf ; press release: https://www.ftc.gov/news-events/news/press-releases/2022/09/ftc-report-shows-rise-sophisticated-dark-patterns-designed-trick-trap-consumers
- Terres, P., Klumpe, J., Jung, D., and Koch, O. (2019). "Digital Nudges for User Onboarding: Turning Visitors into Users." ECIS 2019 Research Papers, 125. https://aisel.aisnet.org/ecis2019_rp/125/
This playbook is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised