What Designers Knew

What did UX research and HCI learn about adapting interfaces to people, and what does it teach AI personalization?

Working draft, revised September 2026 · 25 min read · Markdown for your AI

Questions this chapter answers

  • Why did adaptive interfaces like Clippy and Office's personalized menus frustrate the people they were built to help?
  • Adaptive or adaptable: who should control the change, the system or the person?
  • What do personas, jobs-to-be-done, and field research add that behavioral data cannot?
  • How do qualitative research, machine learning, and A/B testing work together?
  • Can synthetic users replace testing with real people?

The short answer

Long before language models, designers and HCI researchers ran the experiment AI personalization is running now: software that learns a person and changes itself to fit. They learned three things that still hold.

First, where a system adapts matters as much as how well. Interfaces that moved things around (menus that reordered themselves, toolbars that promoted and demoted buttons) made people slower and less sure of themselves, even when the predictions were decent. Interfaces that kept a stable frame and adapted inside a clearly marked area did better. Accuracy matters too, and a lab study found it mattered more than predictability; but inaccurate adaptation in the places people rely on is the combination that fails.

Second, when to act is its own decision. Microsoft's Office Assistant rested on a respectable probabilistic model of what users were trying to do. It failed on a cruder rule about when to interrupt them.

Third, control has to be real. People value seeing and steering what a system believes about them, and they stop trusting controls that do not visibly work.

UX research also brings a method the data team usually lacks. Qualitative research finds why people behave as they do; machine learning finds who behaves that way at scale; experiments prove what worked. A personalization program that skips any of the three is guessing somewhere.

For AI personalization, the translation is five rules: a stable frame, adaptive slots, predictable change, visible control, and explanation on request. Synthetic users can help draft research. They cannot validate it.


In 1999, a person opening the Format menu in Word 2000 saw a short menu. The commands they used most were there; the rest were hidden behind a chevron at the bottom, or appeared after a few seconds. As they worked, the menu learned: items they used were promoted to the short menu, items they ignored were demoted. The idea, as Jensen Harris of the Office user-experience team later put it, was "that eventually you'd have a fully-tuned, auto-customized UI that would show you only what you need and use." [1]

In 2006, explaining why Office 2007 would abandon menus altogether, Harris wrote the post-mortem. "Adaptive Menus were not successful." Scanning took two passes, because pressing the chevron reflowed the menu and reset the eye. And the prediction was never quite right: "Auto-customization, unless it does a perfect job, is usually worse than no customization at all." What people experienced was "a sense [of] randomness and unpredictability: one time, a menu item would be in a certain place, and then two days later it wasn't there anymore." His verdict on the feature and its toolbar twin: "most customers, especially those in corporate environments, turn both of these features off." [1]

The same suite had shipped an even more ambitious learner two years earlier. The Office Assistant, the paperclip most people remember as Clippy, drew on Microsoft Research's Lumière project, which used Bayesian models to infer a user's goals from their actions and queries. [2] Eric Horvitz, who led Lumière, recorded what happened on the way to the product: the Office team put "a relatively simple rule-based system on top of the Bayesian query analysis system to bring the agent to the foreground," and his group "had been concerned upon hearing this plan that this system would be distracting to users." [2] Microsoft made the Assistant off by default in Office XP in 2001 and removed it in Office 2007. [3]

Two lessons, and neither is "the AI was bad." The menus adapted the one thing people needed to stay put. The Assistant's inference was reasonable; its decision about when to speak was not.

Designers modeled need before anyone modeled data

Personalization teams tend to start from the data they hold. Designers started from the person.

Personas. Alan Cooper popularized personas in The Inmates Are Running the Asylum (1999): design for a specific, archetypal user, not an elastic "user" who can be stretched to justify any feature. [4] The critique arrived quickly. Christopher Chapman and Russell Milham argued in 2006 that nobody can tell how many real users a persona represents, and that personas cannot be verified or falsified. [5] A persona is a communication device that keeps a team arguing about a person rather than an abstraction. It is not a segment, and it is not data.

Jobs to be done. Clayton Christensen and colleagues argued that people "hire" a product to make progress in a circumstance. In their best-known example, milkshake research stalled on demographics and moved once researchers watched who bought and when: nearly half were bought before 9 a.m. by lone commuters, hiring the shake to make a long, boring drive more tolerable. [6] Nielsen Norman Group's reading is the useful one: jobs and personas are complements, because two groups can share a job and still want different things from it. [7]

Field research. Contextual inquiry (watching and interviewing people where they actually work), journey maps (the steps, thoughts, and emotions of someone pursuing a goal), and diary studies (behavior logged over weeks) are how designers find the moments that matter. [8][9][10]

Here is why a personalization team should care. These methods do not produce predictions. They produce decisions worth personalizing, which is the input Chapter 6 asks every leader to find. A Larkspur Systems researcher (Larkspur is the fictional field-service software company used throughout this handbook) shadows three dispatchers for a morning and sees each of them print the day's route sheet at 6:40. The tablets lose signal at rural job sites, and paper is the backup. Telemetry records only a print event. The personalization decision that follows is not "which tip to show a dispatcher." It is "which jobs to preload for offline use, per technician, before the van leaves." Telemetry would have ranked that decision nowhere. One morning of fieldwork put it first.

The data knows what people did. Only watching them tells you what they were trying to do.

Five users is a method, not a sample size

Usability testing gave the industry its most quoted rule of thumb, and its most misquoted one. Jakob Nielsen's 2000 argument that you only need to test with five users rests on a model in which each user reveals a share of the problems, about 31 percent on average in the projects he and Tom Landauer studied, so five users find roughly 85 percent. [11] Read the rest of the article and the rule narrows. Nielsen recommends three rounds of five, with a redesign between each, rather than one round of fifteen; three or four users per group when user groups differ; and about twenty users for quantitative measures. [11]

The empirical record narrows it further. Across 49 users on four production websites, Jared Spool and Will Schroeder found the first five revealed about 35 percent of the problems. [12] Laura Faulkner resampled a 60-user study and found that some sets of five found 99 percent of known problems while others found 55 percent. [13] The 85 percent is an average over the obvious problems for one population doing one set of tasks.

That caveat is the personalization problem in miniature. A personalized surface is a family of interfaces, one per combination of person, evidence, and fallback, and the version the design team tests is rarely the one a thin-data customer sees. So test small rounds per variant class (rich record, thin record, fallback fired, each role), and test what was rendered, not what was stored. In the governed personalization engine I built for go-to-market teams, quality review inspects the live page, because stored fields can be correct while the experience is wrong: a stale template, a hidden zone, the wrong page identity, a stale cache, a broken link. [14]

Five users find the obvious problems in one experience. Personalization ships thousands.

Adaptive or adaptable: who controls the change?

Nielsen Norman Group draws the distinction that organizes the rest of this chapter: personalization is done by the system; customization is done by the user. [15] HCI research uses adaptive and adaptable for the same pair, and it has tested both for three decades.

The early result favored stability. In 1994, Andrew Sears and Ben Shneiderman proposed split menus: the few most frequently used items move to a top section, and the rest keep their order. In two field studies, selection times fell 17 to 58 percent. [16] The split was set from usage data, but it did not keep shifting under the user. A decade later, Leah Findlater and Joanna McGrenere compared static, adaptive, and adaptable split menus in the lab: the static menu was faster than the adaptive one, and the abstract reports that most users preferred the control of the adaptable version. [17] The Product playbook applies this to in-app surfaces; I will not repeat it here.

Then the story gets more interesting, and it is where I part company with Harris's rule that imperfect auto-customization is "usually worse than no customization at all." In 2008, Krzysztof Gajos and colleagues at the University of Washington and Microsoft Research ran 23 participants through a modified Word interface with an adaptive toolbar, crossing two accuracy levels (50 and 70 percent) with two strategies (a predictable most-recently-used rule, and random). Both predictability and accuracy raised satisfaction. Accuracy also raised performance and use of the adaptive toolbar, and, "contrary to our expectations," it had the stronger effect. [18] Notice the design of that toolbar: an adaptive area, the authors wrote, "clearly designated as hosting changing content," sitting beside the familiar toolbars, which did not move. The user could always take the familiar route.

The strongest evidence for adaptation comes from the same research line. Gajos's SUPPLE system generates interfaces automatically; its ability-based version, SUPPLE++, built interfaces from short motor tests of people with motor impairments. They were 26.4 percent faster and made 73 percent fewer errors than with the manufacturers' default interfaces, and strongly preferred the generated ones. [19] In an in-vehicle system, Tal Lavie and Joachim Meyer found adaptivity helped most in routine situations. [20]

Put the evidence together and the defensible claim is narrower than either camp's slogan. Adaptation helps when it is accurate, when it rests on something measured rather than guessed, and when it happens somewhere the person expects change. It hurts when it is inaccurate and rearranges what the person has already learned. Harris was describing the worst corner of that space, and he described it accurately.

That gives a spectrum, and the useful question at each point is who controls the change.

A horizontal spectrum of five interface designs, ordered from the person controlling more to the system controlling more. 1, static: the designer decides once and nothing learns. 2, adaptable: the person changes it with settings, pins, and favorites. 3, suggested: the system proposes and the person accepts, as in a split menu or a recommended row. 4, adaptive in a slot: the system changes a marked area while the frame stays fixed and the person can override. 5, adaptive everywhere: the system changes the frame itself, as Office 2000's personalized menus did, which is where trust collapses and customers turned it off. Designs 3 and 4 are outlined as the working zone (stable frame, accurate change in marked slots), and a closing line says accuracy matters most, but only inside a place the person expects to change.

Figure 3.1. The adaptive-to-adaptable spectrum: the working zone is a stable frame, accurate change in marked slots, and a person who can override.

A system that learns should know when to stay quiet

In 1999, Horvitz published "Principles of Mixed-Initiative User Interfaces," twelve factors for systems that share initiative with a person. [21] Among them: consider uncertainty about the user's goals; time the service to the user's attention; weigh the expected costs and benefits of acting; minimize the cost of poor guesses; and scope the precision of the service to match uncertainty, doing less but doing it right. His LookOut prototype chose among three actions for each incoming email: do nothing, ask, or act, according to how confident it was about the user's intent. [21]

Clippy, as shipped, appeared on a rule rather than on the model's uncertainty, and interrupted rather than waited.

Twenty years later, Saleema Amershi and colleagues at Microsoft consolidated more than 150 design recommendations into 18 Guidelines for Human-AI Interaction, tested by 49 design practitioners against 20 AI products. [22] Eight of the 18 map to Horvitz's principles. For personalization, the essential ones are G10 (scope services when in doubt), G14 (update and adapt cautiously), G15 and G17 (granular feedback and global controls), G18 (notify users about changes), and G11 (make clear why the system did what it did). G11 drew one of the highest counts of violations. One participant, testing an e-commerce recommender: "I have no idea why this is being shown to me. Is it trying to sell me stuff I do not need?"

The decision layer in Chapter 14 includes "do nothing" as a first-class action for the same reason Horvitz's LookOut did. Personalization is two decisions, not one: what fits this person, and whether acting now is worth their attention. Clippy got the first roughly right and the second wrong, and the second is the one people remember.

Control that does not work is worse than no control

Recommender research reached the same conclusion from the other side. Nava Tintarev and Judith Masthoff listed seven aims an explanation can serve, among them transparency, trust, and scrutability: letting a person tell the system it is wrong. [23]

What people want from control is specific. In focus groups on news recommenders, Jaron Harambam and colleagues found that participants valued an intelligible profile (what the system thinks they read and like) and the ability to steer the algorithm toward their own goals. Filtering results after the fact was "not sufficient to make users feel that they have control," and participants openly doubted whether controls did anything at all. [24]

That doubt was measured in 2022. Mozilla analyzed data from more than 20,000 volunteers and over 500 million YouTube videos and found that "Don't recommend channel" prevented 43 percent of unwanted recommendations, and "Not interested" only 11 percent. In a survey of 2,758 participants, 62.3 percent said the controls were ineffective or gave mixed results. [25] The button existed. It barely worked. People noticed.

Shipped products have since added controls that change the model rather than the list. Spotify lets listeners exclude playlists (sleep sounds, a child's music) from their taste profile, so context stops masquerading as preference. [26] TikTok lets people reset their For You feed. [27] Those are the right shape: they act on what the system believes, and the effect is visible.

A control is a promise. If it does not visibly change what happens next, it is a broken one.

Qual finds the why, ML finds the who, experiments prove the what

Each discipline has a blind spot. The research team knows why a dispatcher prints the route sheet and cannot say how many dispatchers do. The data team can count the print events across 40,000 accounts and cannot say why they happen. The experimentation team can prove a change moved a metric and cannot say for whom or why. Chapter 2 covers the models and the experiment design in depth; the point here is how the three disciplines hand work to one another.

The experimentation canon says this itself. Ronny Kohavi and colleagues popularized "HiPPO" (the highest-paid person's opinion) in arguing for testing over conviction [28], and the same authors' textbook devotes a chapter to the complementary methods (user experience research, surveys, focus groups, log analysis) that generate hypotheses and explain why metrics moved. [29] Experiments tell you what moved. They do not tell you why.

So run the three as a loop, which I call the Triangulation Loop:

  1. Qualitative research finds the why. Fieldwork and interviews surface a job and the moment it is at stake: dispatchers need tomorrow's jobs on the tablet before the van leaves the yard.
  2. Machine learning finds the who. Telemetry and memory size it: which accounts work rural sites, which technicians lose signal, how often.
  3. Experiments prove the what. Preloading jobs for those technicians is tested against a holdout (Chapter 21), and the measured effect decides whether it ships.
  4. Anomalies go back to the start. When the experiment surprises you, or a segment behaves differently, the next round of interviews goes to those people.

Three stations joined by arrows, with a return path. Qualitative research (fieldwork, interviews, diary studies) finds the why; in the Larkspur (fictional) example, dispatchers print route sheets at 6:40. Machine learning over telemetry and memory finds the who by sizing the need: which accounts work rural sites, and how often. Experiments prove the what: preloaded offline jobs are tested against a holdout, which decides whether it ships. Surprises and segments that behave differently return to qualitative research, and a closing line says that skipping a station leaves the program guessing the why, the who, or the what.

Figure 3.2. The Triangulation Loop: each discipline answers the question the other two cannot.

The loop also works inside a running AI system, and this is where designers' habits pay off most. In the engine I built, a campaign's first outputs are read the way a researcher reads interview transcripts: a person, helped by AI, reviews the first ten to twenty-five real outputs, then the first fifty to a hundred, then a few hundred, and asks questions across the population. What errors repeat? Where is the content too generic? Which fallbacks fire too often? What belongs in code rather than in another guideline? Every finding is classified by root cause before anyone proposes a fix, and a hunch that some content would perform better is routed to a controlled test rather than into the prompt. [14] That is qualitative research on the machine's own output: small samples, read closely, with the numbers held for the experiment.

Deception is the boundary, not a setting

Designers also named what personalization must not become. Harry Brignull coined "dark patterns" in 2010 for interface tricks that make people do things they did not mean to; he now prefers "deceptive patterns." [30] In 2019, Arunesh Mathur and colleagues crawled about 53,000 product pages across roughly 11,000 shopping sites and found 1,818 instances of such patterns, 183 sites using outright deceptive practices, and 22 third-party vendors selling them as ready-made components. Their definition is the one to keep: design that benefits a service "by coercing, steering, or deceiving users." [31]

Personalization makes this more dangerous, not less, because a pattern tuned per person is harder for anyone else to see. The boundary for a builder is structural. Adapt what helps someone decide; never adapt the exit, the price disclosure, the consent, or the cancel path. Those stay fixed, identical, and outside every adaptive slot. Chapters 18 and 19 make this enforceable.

Synthetic users draft; people validate

The newest temptation in UX research is to let language models play the participants.

In Nielsen Norman Group's comparisons, synthetic users produced shallow, unprioritized lists of needs, were sycophantic about concepts, and reported finishing every online course they started, where real participants admitted quitting after three of seven. [32] A 2025 review of three studies found that synthetic users capture the direction of effects but underestimate their size and show less variance than people, and that "digital twins" built from interviews with real individuals do better than personas built from demographics. [33] The strongest result points the same way. Joon Sung Park and colleagues built agents from two-hour interviews with 1,052 people; agents given the interviews and survey answers reproduced held-out survey answers at 86 percent of the consistency people showed with themselves two weeks later, against 74 percent for agents given demographics alone. [34] The interviews were the input. The research was not replaced; it was reused.

Synthetic users are a drafting tool (piloting an interview guide, pretesting wording, listing edge cases), not a validation tool. The Research and Strategy playbook applies the same rule to market research.

What it means for AI personalization

Everything above converges on one design for AI personalization. Here is the sentence this chapter turns on: people forgive a system that guesses wrong inside a place they know will change; they do not forgive one that moves the place.

That yields five rules, which I group as Stable Frame, Adaptive Slots:

  1. Stable frame. Navigation, core controls, legal copy, prices, and exits stay identical for everyone.
  2. Adaptive slots. Change happens only in declared, bounded places, each with a purpose, allowed evidence, limits, and an approved fallback.
  3. Predictable change. A slot changes on meaningful events, not on every visit, and follows rules a person could describe.
  4. Visible control. People can see and correct what the system believes, and the correction visibly changes what happens next.
  5. Explanation on request. "Why am I seeing this?" gets a true, short answer grounded in what the system actually used.

The engine I built follows the first two by construction. A page is a governed static shell with, in a typical first implementation, about 8 to 12 text zones; brand, navigation, and legal copy are never zones. Each zone carries a contract: its purpose, the context it may and may not use, a length limit, its validations, and an approved fallback. [14] Generative does not mean uncontrolled. Chapter 15 covers zone contracts in the build lane, and the chapter on generative experiences (Beyond the Message) extends the same rules to whole pages, apps, and notifications, where Nielsen Norman Group has already named the risk: generated interfaces can cost people the consistency and learned efficiency they rely on. [35]

The rules are old. What is new is that a language model can fill an adaptive slot with something written for one person. The frame is what lets them trust it.

What this does not do

The evidence here is older and smaller than the confidence people bring to it. The menu studies are lab studies with a few dozen participants, run on 2000s desktop software. Extending them to generated content in AI products is an inference, not a finding; I think it is a sound one, but I am labeling it.

The evidence also cuts against pure stability. Gajos found accuracy mattered more than predictability [18], and SUPPLE++ showed large gains from adapting the whole interface when the adaptation rested on measured ability [19]. "Never adapt the frame" is a default for general audiences, not a law. Accessibility may be the strongest case for adapting the frame itself.

Personas remain unfalsifiable [5], journey maps can become wall art, and qualitative research is slow and costly per person. That is its limit, and it is the gap that customer memory (Chapter 10) partly fills: memory read closely can show patterns across thousands of accounts. It cannot replace watching someone work.

Explanation is not a cure. Tintarev and Masthoff note that explanations can be built to persuade rather than to inform [23]; a "why am I seeing this" written to sell is a dark pattern with better manners.

At scale

At scale, controls must be global and shared. A "not interested" pressed in the app has to reach the email program and the account manager's brief, or the control is only local, and the person experiences it as broken, which is the Mozilla finding reproduced across channels. One memory, many surfaces (the Customer Memory chapters and Playbook P9) is what makes a control mean the same thing everywhere.

Proxies also show up at scale. Netflix, which personalizes title artwork per member from viewing history, faced complaints in 2018 that some Black subscribers were shown artwork foregrounding minor Black cast members; the company said it did not use race, gender, or ethnicity, only viewing history. [36] A proxy can reproduce a sensitive attribute nobody collected, and people judge the output, not the input policy. Only a review of rendered outputs across groups catches it.

Failure story: The Helpful Interruption

Larkspur (fictional) ships an in-app assistant for dispatchers. When the product sees three reassignments of the same job within five minutes, a rule opens a panel over the dispatch board offering to optimize the route. The trigger is accurate in its way: repeated reassignment does signal trouble. It fires most at 7 a.m., when dispatchers are reassigning jobs as technicians call in sick, and the panel covers the board they are working. Within a month most dispatchers have found the setting that turns the assistant off, including for its useful features. (Illustrative composite.)

The inference was fine. The interruption policy was the failure, as it was for the Office Assistant. The fix follows Horvitz and Amershi: move the suggestion into a non-modal slot beside the board, suppress it during the morning peak, scope it to one clear action, add "why am I seeing this," and let a dismissal teach the system something.

A good guess at the wrong moment is still a bad decision.

Patterns

Stable Frame, Adaptive Slots. Problem: adaptation erodes people's learned map of the interface. Forces: personalization pays in specific places; people rely on everything else staying put. Solution: fix the frame (navigation, controls, legal, price, exit); adapt only in declared slots with contracts and fallbacks; change slots on events, not visits. Tradeoff: less expressive than whole-screen adaptation; much easier to test, support, and trust.

Working Controls. Problem: controls exist but do not visibly change outcomes, and people stop trusting them. Forces: model-level controls are harder to build than list filters. Solution: controls act on what the system believes (the profile, the evidence), apply across surfaces, say when they take effect, and are measured for effectiveness. Tradeoff: engineering cost; some users will exclude signals that were predictive.

Triangulation Loop. Problem: each discipline optimizes alone and misses what the others see. Forces: research is slow, models are opaque about why, experiments are narrow. Solution: qual finds the why, ML finds the who, experiments prove the what, and surprises return to qual. Tradeoff: slower first cycle; far fewer confident mistakes.

Leader questions

  1. Which parts of our product and pages are allowed to change per person, and who decided that?
  2. When a customer tells us "stop showing me this," how many surfaces honor it, and how do we know it worked?
  3. What was the last personalization decision that came from watching customers rather than from a dashboard?
  4. Can a customer, or our support team, see why the system showed something?
  5. Where are we using synthetic users, and is any of that output being treated as validation?

Build checklist

  • Map every surface into frame (fixed) and slots (adaptive); keep exits, prices, consent, and cancel paths in the frame.
  • Give each slot a contract: purpose, allowed and forbidden evidence, limits, validations, fallback.
  • Define the events that may change a slot; do not re-adapt on every visit.
  • Ship "why am I seeing this" for every adaptive slot, grounded in the evidence actually used.
  • Make controls act on the profile, propagate across channels, and state when they take effect.
  • Measure control effectiveness: after a "not interested," how often does similar content return?
  • Run small usability rounds per variant class (rich data, thin data, fallback, each role), and review rendered surfaces, not stored fields.
  • Put a qualitative step in every release: read a sample of real outputs closely and classify findings by root cause.
  • Tag synthetic research output as synthetic and keep it out of customer memory.

Metrics to watch

  • Control effectiveness: share of "not interested" or exclusion actions after which similar content does not return within a set window.
  • Adaptive-feature opt-out rate: how many people turn personalization off, per surface; a rising rate is the Office 2000 signal.
  • Task time on core flows, before and after adaptation, for the frame's key tasks (they should not get slower).
  • Explanation-request rate and follow-up: how often people ask why, and how often they then correct or dismiss.
  • Share of personalization decisions with a qualitative origin: how many shipped slots trace to research, not only to data.

Reader Q&A

Should we let users customize instead of personalizing for them? Offer both. Adaptable controls give people a sense of ownership and a way to correct the system; adaptive slots help people who do not know what they need. The evidence favors systems that adapt accurately inside a stable frame and let the person override.

We ran five usability tests on the new personalized page. Are we covered? For the version you tested, on its obvious problems. Test small rounds per variant class, especially thin-data and fallback versions, and review a sample of what real customers were actually shown.

Do personas still matter if we have customer memory? Yes, as a communication device and a source of hypotheses, not as data. Memory tells you what each customer did and said. A persona, grounded in research, keeps a team designing for a person rather than for a metric.

Can we use AI-generated users to test personalization? To draft and pilot, yes. To decide, no. Simulated users flatten variance and agree too easily; decisions need real people or a controlled experiment.

For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 3
title: "What Designers Knew"
concepts:
  - name: Adaptive vs adaptable
    definition: "Adaptive interfaces change themselves based on the system's model of the user; adaptable interfaces are changed by the user. Most working designs sit between: the system suggests or adapts in marked slots, and the person can override."
  - name: Stable Frame, Adaptive Slots
    definition: "Keep navigation, core controls, legal copy, prices, and exits fixed for everyone; allow change only in declared, bounded slots with contracts and fallbacks."
  - name: Predictability vs accuracy
    definition: "Two separate properties of adaptation; lab evidence shows both raise satisfaction and accuracy has the stronger effect, but inaccurate adaptation of what people rely on fails worst."
  - name: Mixed-initiative interaction
    definition: "Systems that share initiative with the person, choosing among doing nothing, asking, or acting according to uncertainty about the person's goal and the cost of interrupting."
  - name: Triangulation Loop
    definition: "Qualitative research finds the why, machine learning finds the who, experiments prove the what; surprises return to qualitative research."
  - name: Working Controls
    definition: "User controls that act on the system's model of the person, apply across surfaces, state when they take effect, and are measured for effectiveness."
  - name: Synthetic users
    definition: "LLM-simulated research participants; useful for drafting and piloting research, not for validating decisions."
decision_rules:
  - if: "an element is navigation, a core control, legal copy, a price, consent, or an exit path"
    then: "keep it in the stable frame; never personalize it"
  - if: "the system is uncertain about the person's goal or the cost of interrupting is high"
    then: "do less: stay quiet, or offer a non-modal suggestion in a slot rather than interrupting"
  - if: "a user control exists"
    then: "verify it changes the system's model and the next outputs across all surfaces; measure its effectiveness"
  - if: "a slot shows personalized content"
    then: "provide a short, true explanation on request grounded in the evidence actually used"
  - if: "a personalization idea comes only from data"
    then: "check it against qualitative research before building; if it comes only from research, size it with data; in both cases prove it with an experiment"
  - if: "testing a personalized surface"
    then: "run small usability rounds per variant class, including thin-data and fallback versions, and review rendered output"
  - if: "research output was produced by synthetic users"
    then: "treat it as a draft or hypothesis; validate with real people or an experiment; never write it into customer memory"
assessment_questions:
  - "Which surfaces change per person, and which elements are guaranteed not to change?"
  - "Which user controls exist for personalization, and has their effect been measured?"
  - "Can customers and support staff see why a personalized element was shown?"
  - "When did qualitative research last change a personalization decision?"
  - "How are personalized variants usability-tested, including fallback versions?"
  - "Where are synthetic users or simulated panels used, and for what decisions?"
patterns: [Stable Frame Adaptive Slots, Working Controls, Triangulation Loop]
anti_patterns: [The Helpful Interruption, The Shifting Map, Broken Controls, Synthetic Validation, Five-User Certainty]
maturity_dimension: experience_design

References

  1. Harris, J. (2006-03-31). "Combating the Perception of Bloat (Why the UI, Part 3)." Office User Interface Blog, Microsoft. https://learn.microsoft.com/en-us/archive/blogs/jensenh/combating-the-perception-of-bloat-why-the-ui-part-3
  2. Horvitz, E., Breese, J., Heckerman, D., Hovel, D., & Rommelse, K. (1998). "The Lumière Project: Bayesian User Modeling for Inferring the Goals and Needs of Software Users." UAI 1998. https://arxiv.org/abs/1301.7385 ; project page with the Office Assistant account: http://erichorvitz.com/lum.htm
  3. "Office Assistant." Wikipedia (off by default in Office XP, announced April 2001; removed in Office 2007). https://en.wikipedia.org/wiki/Office_Assistant ; Sinofsky, S. (2005-09-12). "More about learning from the past: Office Assistant." https://learn.microsoft.com/en-us/archive/blogs/techtalk/more-about-learning-from-the-past-office-assistant
  4. Cooper, A. (1999). The Inmates Are Running the Asylum. Sams.
  5. Chapman, C. N., & Milham, R. P. (2006). "The Personas' New Clothes: Methodological and Practical Arguments against a Popular Method." Proc. HFES Annual Meeting 50(5), 634 to 636. https://journals.sagepub.com/doi/10.1177/154193120605000503
  6. Christensen, C. M., Hall, T., Dillon, K., & Duncan, D. S. (2016). "Know Your Customers' 'Jobs to Be Done'." Harvard Business Review, September 2016. https://hbr.org/2016/09/know-your-customers-jobs-to-be-done
  7. Laubheimer, P. (2017-08-06). "Personas vs. Jobs-to-Be-Done." Nielsen Norman Group. https://www.nngroup.com/articles/personas-jobs-be-done/
  8. Beyer, H., & Holtzblatt, K. (1998). Contextual Design: Defining Customer-Centered Systems. Morgan Kaufmann. https://dl.acm.org/doi/10.5555/286067
  9. Gibbons, S. (2018-12-09). "Journey Mapping 101." Nielsen Norman Group. https://www.nngroup.com/articles/journey-mapping-101/
  10. Flaherty, K. (2024-03-29). "Diary Studies: Understanding Long-Term User Behavior and Experiences." Nielsen Norman Group. https://www.nngroup.com/articles/diary-studies/
  11. Nielsen, J. (2000-03-18). "Why You Only Need to Test with 5 Users." Nielsen Norman Group. https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/ ; history in Sauro, J., "A Brief History of the Magic Number 5 in Usability Testing," MeasuringU. https://measuringu.com/five-history/
  12. Spool, J., & Schroeder, W. (2001). "Testing Web Sites: Five Users Is Nowhere Near Enough." CHI '01 Extended Abstracts, 285 to 286. https://dl.acm.org/doi/10.1145/634067.634236
  13. Faulkner, L. (2003). "Beyond the Five-User Assumption: Benefits of Increased Sample Sizes in Usability Testing." Behavior Research Methods, Instruments, & Computers 35(3), 379 to 383. https://doi.org/10.3758/BF03195514
  14. Author's governed personalization engine: zone contracts (static shell, about 8 to 12 zones), rendered-surface review, and human-supervised review of early campaign cohorts with root-cause classification. Internal capabilities reference, 2026-08-25 revision, sections 5.1, 9.4, 10.1 to 10.3 (status LIVE; root-cause workflow partly productized). [PRODUCT-NAMING: name the engine here if the product decision allows]
  15. Schade, A. (2016-07-10). "Customization vs. Personalization in the User Experience." Nielsen Norman Group. https://www.nngroup.com/articles/customization-personalization/
  16. Sears, A., & Shneiderman, B. (1994). "Split Menus: Effectively Using Selection Frequency to Organize Menus." ACM TOCHI 1(1), 27 to 51. https://api.drum.lib.umd.edu/server/api/core/bitstreams/f53cafcd-2b8d-46d0-a8fe-87420f708f72/content
  17. Findlater, L., & McGrenere, J. (2004). "A Comparison of Static, Adaptive, and Adaptable Menus." CHI 2004, 89 to 96. https://dl.acm.org/doi/10.1145/985692.985704
  18. Gajos, K. Z., Everitt, K., Tan, D. S., Czerwinski, M., & Weld, D. S. (2008). "Predictability and Accuracy in Adaptive User Interfaces." CHI 2008. https://dx.doi.org/10.1145/1357054.1357252
  19. Gajos, K. Z., Wobbrock, J. O., & Weld, D. S. (2008). "Improving the Performance of Motor-Impaired Users with Automatically-Generated, Ability-Based Interfaces." CHI 2008. https://dl.acm.org/doi/10.1145/1357054.1357250 ; SUPPLE: Gajos, K., & Weld, D. S. (2004), IUI 2004. https://dl.acm.org/doi/10.1145/964442.964461
  20. Lavie, T., & Meyer, J. (2010). "Benefits and Costs of Adaptive User Interfaces." International Journal of Human-Computer Studies 68(8), 508 to 524. https://doi.org/10.1016/j.ijhcs.2010.01.004
  21. Horvitz, E. (1999). "Principles of Mixed-Initiative User Interfaces." CHI '99, 159 to 166. https://erichorvitz.com/chi99horvitz.pdf
  22. Amershi, S., Weld, D., Vorvoreanu, M., et al. (2019). "Guidelines for Human-AI Interaction." CHI 2019. https://www.microsoft.com/en-us/research/publication/guidelines-for-human-ai-interaction/
  23. Tintarev, N., & Masthoff, J. (2007). "A Survey of Explanations in Recommender Systems." ICDE Workshops 2007. https://dl.acm.org/doi/10.1109/ICDEW.2007.4401070 ; extended in UMUAI 22 (2012). https://link.springer.com/article/10.1007/s11257-011-9117-5
  24. Harambam, J., Bountouridis, D., Makhortykh, M., & van Hoboken, J. (2019). "Designing for the Better by Taking Users into Account: A Qualitative Evaluation of User Control Mechanisms in (News) Recommender Systems." RecSys '19, 69 to 77. https://doi.org/10.1145/3298689.3347014
  25. Mozilla Foundation (2022-09-20). Blog post on the "Does This Button Work?" investigation of YouTube user controls. https://www.mozillafoundation.org/en/blog/mozilla-investigation-youtubes-dislike-button-other-user-controls-largely-fail-to-stop-unwanted-recommendations/
  26. Spotify (2023-02-08). "'Exclude From Your Taste Profile' Will Make Your Personalized Recommendations Even Better." https://newsroom.spotify.com/2023-02-08/exclude-from-your-taste-profile-will-make-your-personalized-recommendations-even-better/
  27. TikTok (2023-03-16). "Introducing a Way to Refresh Your For You Feed on TikTok." https://newsroom.tiktok.com/en-us/introducing-a-way-to-refresh-your-for-you-feed-on-tiktok-us
  28. Kohavi, R., Henne, R. M., & Sommerfield, D. (2007). "Practical Guide to Controlled Experiments on the Web: Listen to Your Customers not to the HiPPO." KDD 2007. https://ai.stanford.edu/~ronnyk/2007GuideControlledExperiments.pdf ; HiPPO origin: https://exp-platform.com/hippo/
  29. Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press, ch. 10, "Complementary Techniques." https://www.cambridge.org/core/books/abs/trustworthy-online-controlled-experiments/complementary-techniques/D65F23223FD2B9F88929BA3FF579E48C
  30. Brignull, H. (2023). Deceptive Patterns: Exposing the Tricks Tech Companies Use to Control You. Testimonium. https://deceptive.design/about-us/
  31. Mathur, A., Acar, G., Friedman, M. J., Lucherini, E., Mayer, J., Chetty, M., & Narayanan, A. (2019). "Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites." Proc. ACM HCI 3 (CSCW), Article 81. https://arxiv.org/abs/1907.07032
  32. Rosala, M., & Moran, K. (2024-06-21). "Synthetic Users: If, When, and How to Use AI-Generated 'Research'." Nielsen Norman Group. https://www.nngroup.com/articles/synthetic-users/
  33. Budiu, R. (2025-08-15). "Evaluating AI-Simulated Behavior: Insights from Three Studies on Digital Twins and Synthetic Users." Nielsen Norman Group. https://www.nngroup.com/articles/ai-simulations-studies/
  34. Park, J. S., Zou, C. Q., Kamphorst, J., et al. (2024, rev. 2026). "LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals" (earlier title: "Generative Agent Simulations of 1,000 People"). arXiv:2411.10109 (current abstract). https://arxiv.org/abs/2411.10109
  35. Moran, K., & Gibbons, S. (2024-03-22). "Generative UI and Outcome-Oriented Design." Nielsen Norman Group. https://www.nngroup.com/articles/generative-ui/
  36. The FADER (2018-10-22), report on Netflix artwork and race, with Netflix's statement. https://www.thefader.com/2018/10/22/netflix-target-users-race ; Chandrashekar, A., Amat, F., Basilico, J., & Jebara, T. (2017-12-07). "Artwork Personalization at Netflix." Netflix TechBlog. https://netflixtechblog.com/artwork-personalization-c589f074ad76

This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.

Get chapters by email as they are revised

Prefer LinkedIn? Subscribe to the newsletter there instead.