# Playbook RS. Research and Strategy

*The question: How does understanding individuals roll up into strategy?*

> **Questions this playbook answers**
> - How deep should we research each account, and which accounts deserve it?
> - We hold thousands of account records, call notes, and research briefs. How do we turn them into an ICP and a market view we can defend?
> - Can AI-simulated customers (synthetic panels) replace our surveys and interviews?
> - Where do synthetic panels genuinely help, and where do they mislead?
> - How do we know a strategy built this way is right?

This playbook covers three connected jobs: deep research per account, rolling individual understanding up into an ideal customer profile (ICP) and strategy, and synthetic research panels with their limits. Research for a first message belongs to Sales I; here, account research is raw material for strategy.

## The moment

Larkspur Systems (a fictional composite company used throughout this handbook) is in annual planning, and two documents are on the table.

The first is the ICP slide, unchanged for three years: fleet operators with more than 100 vehicles, in logistics and utilities. Every targeting list is cut from it.

The second is an export nobody asked for. Six months earlier, Larkspur started writing account research into its shared customer memory (Chapter 10) with a fixed schema, including a typed field for what triggered each purchase, and backfilled it from two years of call notes and emails. An analyst in revenue operations counted that field across every won and lost opportunity. Fleet size and industry barely separated the two groups. A situation did: in 61% of wins, the account had opened a second depot or taken over dispatch for another site in the prior year, against 18% of losses (illustrative figures for a fictional company). Larkspur's best customers were not defined by how many vehicles they ran. They were defined by how many places they dispatched from, and by having just added one.

In the same meeting, the product lead proposes testing a new per-depot pricing tier. Recruiting real fleet managers for a survey would take three weeks. A synthetic panel of 2,000 simulated fleet managers could answer overnight.

The CFO asks the useful question: which of these two results would you bet the plan on?

The first came from counting real customers; the second would come from a model's idea of customers. Both have uses. Only one is evidence. Here is the sentence this playbook turns on: **research on one account is a brief; research on every account, written into memory with the same schema, is a market study you did not have to commission.**

## The decision

Three decisions repeat here, and each includes doing nothing.

**1. How deep to research this account, now.** The options are deep research, a standard first pass, a refresh of what is already known, or nothing. The rule: research depth follows a pending decision. A renewal in 120 days or a large expansion justifies depth; an account with no decision waiting gets nothing new. Depth has a real price. Anthropic reported that its multi-agent research system beat a single agent by 90.2% on an internal research evaluation, and that multi-agent systems used about 15 times the tokens of a chat interaction, with token usage alone explaining about 80% of performance variance on the benchmark it analyzed [1]. Deeper research is better research, bought with more reading. Spend it where a decision will change.

**2. Whether a pattern in the roll-up should change strategy.** Adopt it, test it, or leave the strategy alone. A pattern earns the right to change the ICP only after it survives a test on data it was not found in (see "How you will know").

**3. What a synthetic panel may be used for.** The evidence below supports a narrow answer:

| Use | Allowed? | Why |
|---|---|---|
| Generating hypotheses, objections, and questions to ask real customers | Yes | Cheap breadth; errors are caught by the human study that follows |
| Pretesting a survey or interview guide for confusing wording | Yes | The panel stands in for a careful reader, not for the market |
| Shortlisting which variants deserve a human test | Yes, with calibration | Aggregate direction is where panels are strongest; a variant the panel cuts is untested, not refuted |
| Evaluating your own personalization output for likely misreadings | Yes, builder-side | A quality check inside the accuracy stack (Chapter 20) |
| Sizing demand, setting price, or deciding a launch | No | Low variance and drift make wrong answers look certain |
| Writing simulated answers into a customer's record | Never | A simulation is not a fact about a person |

## What you need to know

### Account research has two jobs and one unit

Research has two jobs, and most systems do only the first. The first is to add what you do not have. The second is to challenge what you already believe (Chapter 7 calls this Enrichment as Cross-Examination). A CRM field that says fifty employees, next to current sources describing a national operation, is not a fact. It is a disagreement. A fact from a single source is a hypothesis; the same fact from an independent second source is usable.

In the governed personalization engine I built for go-to-market teams, research draws on several sources and tolerates partial failure: one failed source does not fail the account, and a source that keeps failing is dropped for the rest of the batch. An implausible value (a headcount of 100 beside clear multinational markers) is hidden from the trusted fields. The policy is hide when uncertain: a missing number costs less than a confident wrong one.

Research also has a unit: the entity, not the record. On a production pipeline I run, keying research to the company with a seven-day freshness window raised the share of records with rich research from 60.3% to 93.5% while spend fell (Chapter 5, Research Once per Entity).

Then the step most teams skip: the output of account research is not only a brief for a person. It is a write to memory in three forms (Chapter 10, Memory Layers): typed properties, evidence-level memories with sources, and a synthesis. The brief is read once, by one account manager. The properties are read by every roll-up that follows. **The brief serves a meeting; the properties serve the strategy.**

### You cannot group by a paragraph

The Boundary Rule from Chapter 10 decides whether a roll-up is possible: if you would put it in a WHERE clause or a chart, it is a property. Before any account is researched, define the strategy-grade properties it must fill. A minimal set for Larkspur:

```yaml
purchase_trigger:     # enum: new_site, compliance_deadline, fleet_growth, incumbent_failure, new_leader, cost_mandate, other, unknown
primary_job:          # enum: the job the customer hired the product for
competitor_evaluated: # list of named competitors, or none
stated_objection:     # enum, plus evidence link
outcome:              # won, lost, churned, expanded, open
loss_or_churn_reason: # enum, plus evidence link
evidence_type:        # stated | observed | inferred
```

Each property carries a description written as an instruction to the extractor: what to capture, where to find it, what it is not (Chapter 8, Extract with a Contract). Each value carries its source and whether the customer stated it, you observed it, or a model inferred it (Chapter 7, Inferred Lane). Keep enums stable; when one changes, re-extract history (Chapter 11).

That engine writes research as structured account intelligence, not one long paragraph: buying triggers, leadership changes, acquisitions, competitors, strategic priorities, each a field trust-filtered on its own. The step I am still building attaches source, confidence, and freshness to each field rather than to the record. One rule from the self-hosted memory system I built: when a research job's output is bound straight to a typed property, the binding stays optional, so an empty or invented value is never written into a field a roll-up will count.

### Rolling up: count, then characterize

The roll-up that produced Larkspur's depot finding follows six steps, and the order matters.

1. **Fix the denominator first.** Include lost deals, churned accounts, and accounts that never bought. An ICP built only from winners describes who bought, not what separates buyers from everyone else.
2. **Count deterministically.** Group typed properties by outcome in SQL. No model is involved in the arithmetic. My memory system exposes this as a property filter (purchase trigger is new site, outcome is won) that never calls a model, so the same query returns the same count every time.
3. **Weight by account, not by volume of evidence.** Large accounts generate more calls, tickets, and notes. Counting mentions instead of accounts drifts the strategy toward whoever talks most.
4. **Characterize second.** For each pattern that survives counting, a model reads a stratified sample of the underlying evidence, winners and losers alike, and names what it sees. Every sentence in the resulting memo links to the count and to example evidence. It is the review I run on a campaign's early cohorts: every flagged record plus a stratified sample of clean ones, with questions asked of the population, such as "Which facts appear unreliable?" and "What do the strongest outputs have in common?" For strategy, ask the second of a sample that includes the weakest accounts.
5. **Report stated and inferred separately.** "Customers told us" and "we inferred" are different strengths of evidence, and the strategy memo shows both.
6. **Test out of time** before adopting anything (see "How you will know").

The same machinery produces market insight: competitor mentions, loss reasons, and requested capabilities become quarterly time series instead of anecdotes. Counting is deterministic; naming is generative; never mix them in one step.

![Each account with a pending decision (a renewal, an expansion, a new site) gets research that reads the cache first and fills only the gaps, then writes to memory in three forms: typed properties, evidence, and a synthesis; the brief serves a meeting, the properties serve strategy. Across every account the properties are rolled up in six steps: fix the denominator to include won, lost, and churned; count by outcome in SQL; weight by account; let a model characterize the patterns; separate stated from inferred; and test out of time. The result is a situational ICP (multi-site operators within a year of adding a site) that becomes a versioned guideline every agent reads for scoring, research depth, routing, and framing, and that sets which accounts deserve deep research next.](/images/handbook/pb-rs-research-to-icp.svg)

*Figure RS.1. Research written as typed properties on every account can be counted; count first, name second, test out of time, and the ICP it produces sets the next round of research.*

The ICP that comes out tends to change shape, from a firmographic filter to a situation: multi-site operators within a year of adding a site. That is harder to buy as a list and far more useful for deciding which accounts to research deeply, closing the loop to decision 1.

### Synthetic panels: what the evidence supports

A synthetic panel is a set of simulated respondents produced by a language model, each conditioned on a description of someone, and asked questions as if it were a customer. There are three grades, and the difference between them matters more than the model used.

- **Persona panels** condition the model on demographic or firmographic descriptions. Argyle and colleagues showed in 2023 that GPT-3, conditioned on backstories drawn from real survey participants, reproduced response patterns of US subgroups, a property they named "algorithmic fidelity," producing what they called "silicon samples" [2].
- **Grounded panels** condition each simulated respondent on a real individual's own data. Park and colleagues built agents from two-hour interviews with over a thousand people; on held-out survey items the combined interview-and-survey agents reached 86% of participants' own two-week test-retest consistency, against 74% for agents given demographics only [3].
- **Calibrated hybrids** combine synthetic and human responses statistically. Wang, Zhang, and Zhang showed that in conjoint studies this approach cut the human data needed by 24.9% to 79.8% while reducing estimation error, and that naively substituting synthetic answers for human ones produced no savings, because of systematic bias [4].

The case for is real. Brand, Israeli, and Ngwe found GPT produced realistic willingness-to-pay estimates and textbook patterns such as declining marginal utility, and that fine-tuning on prior survey data from the same category improved alignment for existing and new features [5]. In *Nature* in 2026, Ashokkumar and colleagues had GPT-4 simulate responses to 70 preregistered US survey experiments (469 effects, 119,330 participants); its predicted effects correlated strongly with actual ones and matched pooled human forecasters, including for studies published after its training cutoff, though it tended to overestimate effect sizes [6].

The case against is as well documented. Bisbee and colleagues found synthetic responses had too little variance, that 48% of regression coefficients differed significantly from the human survey, with the sign flipping in 32% of those, and that identical prompts gave different distributions between April and July 2023 [7]. Verasight compared LLM responses with a June 2025 poll of 1,500 US adults: the best model was off by 4 points on average and the worst by 23; on a zoning question one model reported 59% support where the real figure was 28%, and the 29% of real respondents who said "don't know" became 0% [8]. Wang, Morgenstern, and Dickerson argue from how models are trained, and show in human studies, that LLMs standing in for participants misportray and flatten identity groups [9].

Read together, the pattern is consistent. Panels are good at the center of a distribution and poor at its edges and its uncertainty: "don't know" vanishes, minority positions converge on the mean, and low variance makes a wrong answer look significant. Strategy lives at the edges: a new segment, an unmet need, the undecided. **A synthetic panel is a well-read guess about the average customer. Strategy is usually about the customers who are not average.** Chapter 3 reaches the same verdict for UX research: synthetic users draft; people validate.

![Real respondents spread across the range of answers, while a synthetic panel answers in a narrow, tall peak at the center: edges are pulled toward the mean and low variance makes a wrong answer look certain. In the poll the chapter cites, 29% of real respondents said don't know and the model's share was 0%. Panels help with hypotheses and questions to ask, pretesting survey wording, shortlisting variants with calibration, and checking your own output; they mislead on sizing demand, setting price, deciding a launch, and never belong in a customer record. The workable flow is synthetic panel, then shortlist, then a human study, then the decision.](/images/handbook/pb-rs-synthetic-panel-edges.svg)

*Figure RS.2. Panels are good at the center and poor at the edges, so they shortlist and humans decide.*

I build synthetic-panel systems, and the version that behaves best is the grounded one: each simulated respondent is conditioned on one real customer's memory rather than on a persona description. That is the roll-up run in reverse, and it inherits every weakness of the memory beneath it.

## The action

None of this is real-time. All three jobs run asynchronously, precompute their results, and serve them to the systems that act (Chapter 17).

**Account research** is event-triggered: a renewal window, a new opportunity, a leadership change or new site. It reads the entity cache first, researches only the gaps, and writes typed properties, evidence, and an account plan with every inference labeled as a hypothesis. Freshness policies (Chapter 11) decide when a property must be re-verified. In my memory system, a long research run can execute as a background job on the model provider's infrastructure, holding a credential minted for that job alone: limited to saving, retrieving, and searching memory, revoked when the job ends, with every write attributed to the job. Routine refreshes run as scheduled prompts.

**The roll-up** runs monthly for market insight and quarterly for the ICP. Its output is a strategy memo in which every claim carries a count, a denominator, and example evidence. My engine's campaign reports already aggregate stored intelligence this way (population, score distribution, engagement, common signals); the ICP roll-up is that report run over outcomes. When the ICP changes, it does not stay a slide. It becomes a versioned guideline that every agent reads (Chapter 18): lead scoring, research depth, routing, and message framing all consume the same definition. Governance is a code path; so is strategy.

**Synthetic panels** run in two places. Before human research, they pretest instruments and shortlist variants, so the expensive human study asks better questions. After generation, they sit inside the accuracy stack: simulated recipients grounded in memory flag messages a customer would likely find wrong, irrelevant, or confusing, and those go to review. The objective of that check is accuracy and appropriateness, never pressure, and that objective belongs in governance, not in the prompt.

## What can go wrong

### Failure story: The Confident Panel

A year before the planning meeting, Larkspur tested a fuel-analytics add-on on a synthetic panel of 1,500 personas described as fleet managers, with sizes and industries sampled to match the customer base (a fictional composite of a pattern that the studies above predict). Seventy percent said they would pay. The confidence interval was narrow, because the simulated respondents barely disagreed with one another. The add-on launched. Large fleets bought it. Single-site operators, most of the customer base, mostly did not, and the few who did churned within two quarters.

The panel answered as a well-read model expects a fleet manager to answer, and much public writing about fleet management is about large fleets. The panel had no way to say "I don't know what I would use this for," which is what the real small operators would have said. **The Confident Panel** is the anti-pattern: synthetic agreement read as market evidence, with low variance mistaken for certainty. The fix was the rule in the decision table: panels shortlist, humans decide.

### Other risks to govern

- **Survivor ICP.** The profile is built from winners only and describes who bought, not what distinguishes them.
- **The Loudest Account.** Evidence is weighted by mentions, so the accounts with the most calls and tickets write the strategy.
- **Echoed Corroboration** (Chapter 13). Several research runs appear to confirm a fact because they all inherited one unverified write. Count independent sources, not agreeing ones.
- **Synthetic Leak.** A simulated answer is written into a real customer's record and later used to personalize. Tag synthetic output at the source and keep it out of customer memory entirely.
- **Reversing the roll-up onto a person.** A segment trait is not a personal fact. "Multi-site operators tend to struggle with dispatch" may guide research; it may not be asserted to one customer as something you know about them (Chapter 19, not legal advice). Check that using customer data for aggregate analysis is compatible with the purposes it was collected for.

### What is overstated

Pitches for "digital twins of your customers" often cite a high correlation between synthetic and human answers. Ask what was correlated: aggregate means can correlate well while segments are badly wrong. I part company with the popular reading of Argyle and of Brand, Israeli, and Ngwe, which treats their findings as permission to replace respondents. Both papers are more careful than that: Brand and colleagues position models as a supplement to human studies, with the most reliable gains when prior human data from the same category and population is already in hand [5]. I also part company with blanket dismissal. The grounded and calibrated results [3][4][6] are real, and they are improving.

Roll-ups overreach too. The depot finding is a correlation: a second site might be why customers buy, or only when they happen to evaluate. Chapter 21 guards against reading history as cause.

## How you will know

**Account research.** Grade a monthly sample of briefs against sources for unsupported and stale claims. Track the decision-change rate: how often research changed the plan it was commissioned for. Research that never changes a decision is a cost to cut. Track cost per researched decision (Chapter 12).

**The ICP and strategy.** Treat the ICP as a prediction and test it like one. First, out of time: find the pattern on older data, then check it on the most recent two quarters it never saw. Then prospectively: tag new opportunities with fit under both ICPs and compare win rate, time to value, and twelve-month retention, keeping part of prospecting on the old definition as a control (Chapter 21). A new ICP that does not beat the old one on data it was not built from is a story, not a strategy. An out-of-time test is only honest if you can see what was known then. My memory system keeps every superseded property value with the period it held, so the test can read each account as it looked while the deal was open, not as later research rewrote it.

**Synthetic panels.** Keep a calibration ledger per question type. Before trusting a panel on pricing, feature preference, or message comprehension, run it on questions where you already hold human answers, and record four numbers: average error, the ratio of synthetic to human variance, the worst subgroup error, and the "don't know" rate against the human rate. Record which model version produced them. Then track forward hit rate: how many shortlisted variants won the human test. Set thresholds by consequence; wording tolerates error that price cannot.

## Reader Q&A

**Can we stop interviewing customers now?**
No. Panels are cheapest when they make interviews better: better questions, better shortlists, fewer wasted sessions. The strongest results in the literature all depend on real human data, either as the grounding for each simulated respondent or as the calibration sample.

**A vendor offers "digital twins" of our buyers. What should we ask?**
What each twin is grounded in (a persona description or a real person's data); error by subgroup, not only on average; the "don't know" rate; how results changed across model versions; and whether any synthetic output can reach a customer record.

**How many accounts before a roll-up means anything?**
Enough that each pattern holds in both won and lost groups with room to spare, and survives an out-of-time check. A pattern found in forty accounts is a hypothesis for the next research cycle, not a strategy.

## For your AI

```yaml
playbook: RS
title: "Research and Strategy"
question: "How does understanding individuals roll up into strategy?"
concepts:
  - name: Research Depth Follows Decisions
    definition: "Account research depth (deep, standard, refresh, none) is set by the value of a pending decision on that account, not by a blanket policy."
  - name: Two Jobs of Research
    definition: "Research adds what is missing and challenges what is held; a single-source fact is a hypothesis until independently corroborated."
  - name: Strategy-Grade Properties
    definition: "Typed, enumerated properties (purchase trigger, job, objection, competitor, outcome, loss reason, evidence type) filled by every account's research so it can be counted across accounts."
  - name: Count, then Characterize
    definition: "Roll-ups count typed properties deterministically by outcome, weighted by account, then let a model name patterns from a stratified evidence sample; every claim links to a count and evidence."
  - name: Situational ICP
    definition: "An ideal customer profile defined by the situation that separates winners from losers (for example, recently added a site), rather than by firmographics alone."
  - name: Synthetic Panel
    definition: "LLM-simulated respondents conditioned on personas (persona panel), on real individuals' data (grounded panel), or combined statistically with human samples (calibrated hybrid)."
  - name: Calibration Ledger
    definition: "Per question type and model version: average error, synthetic-to-human variance ratio, worst subgroup error, don't-know rate, and forward hit rate against human tests."
decision_rules:
  - if: "an account has no pending decision"
    then: "do not commission new research; refresh only what a freshness policy requires"
  - if: "a roll-up pattern is found"
    then: "test it out of time and prospectively against the current ICP before adopting it"
  - if: "the roll-up counts mentions rather than accounts, or excludes lost and churned accounts"
    then: "rebuild the denominator before drawing conclusions"
  - if: "a synthetic panel result would set price, size demand, or decide a launch"
    then: "treat it as a hypothesis and run a human study"
  - if: "the question concerns a small segment, a new category, or events after the model's training cutoff"
    then: "do not use a synthetic panel as evidence"
  - if: "the panel's model version changed since last calibration"
    then: "re-run the calibration ledger before relying on it"
  - if: "a researched value is implausible against other evidence, or one research source keeps failing"
    then: "hide the value from trusted properties and build from the remaining sources"
  - if: "an out-of-time test is run on properties that research has since rewritten"
    then: "read the property values as they held at the time, from property history"
  - if: "output is synthetic"
    then: "tag it at the source and never write it into a customer record"
  - if: "a segment-level insight is about to be used in a message to one person"
    then: "treat it as inference: guide research with it, do not assert it"
assessment_questions:
  - "Is account research written into shared memory as typed properties, or only as briefs and documents?"
  - "Which properties would you need to count to explain why you win and lose, and are they captured today?"
  - "Does your ICP analysis include lost, churned, and non-buying accounts?"
  - "Where do you use or plan to use synthetic respondents, and what human data calibrates them?"
  - "How would you know if your current ICP is wrong?"
patterns: [Research Depth Follows Decisions, Research Once per Entity, Hide When Uncertain, Count then Characterize, Quote-Backed Insight, Calibration Ledger, Panels Shortlist Humans Decide]
anti_patterns: [The Confident Panel, Survivor ICP, The Loudest Account, Synthetic Leak, Echoed Corroboration]
links: [ch5, ch7, ch8, ch10, ch11, ch13, ch17, ch18, ch19, ch20, ch21]
maturity_dimension: understanding
as_of: 2026-09-26
```

## References

1. Anthropic (2025-06). "How we built our multi-agent research system." https://www.anthropic.com/engineering/multi-agent-research-system
2. Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., Wingate, D. (2023). "Out of One, Many: Using Language Models to Simulate Human Samples." *Political Analysis* 31(3). https://arxiv.org/abs/2209.06899
3. Park, J. S., Zou, C. Q., Kamphorst, J., et al. (2024, rev. 2026). "LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals" (earlier title: "Generative Agent Simulations of 1,000 People"). arXiv:2411.10109. https://arxiv.org/abs/2411.10109
4. Wang, M., Zhang, D. J., Zhang, H. (2024, rev. 2026). "Large Language Models for Market Research: A Data-augmentation Approach." arXiv:2412.19363. https://arxiv.org/abs/2412.19363
5. Brand, J., Israeli, A., Ngwe, D. (2023). "Using GPT for Market Research." Harvard Business School Working Paper 23-062; ACM EC 2024. https://www.msi.org/working-paper/using-gpt-for-market-research/ ; https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4395751
6. Ashokkumar, A., Hewitt, L., Ghezae, I., Willer, R. (2026-07-08). "Large language models can predict the results of social science experiments." *Nature* 656:115-122. https://www.nature.com/articles/s41586-026-10742-x
7. Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., Larson, J. M. (2024). "Synthetic Replacements for Human Survey Data? The Perils of Large Language Models." *Political Analysis* 32(4). https://www.cambridge.org/core/journals/political-analysis/article/synthetic-replacements-for-human-survey-data-the-perils-of-large-language-models/B92267DC26195C7F36E63EA04A47D2FE
8. Morris, G. E., and the Verasight data team (2025-08-18, updated 2026-09-01). "Your Polls on ChatGPT." Verasight. https://www.verasight.io/reports/synthetic-sampling
9. Wang, A., Morgenstern, J., Dickerson, J. P. (2025). "Large language models that replace human participants can harmfully misportray and flatten identity groups." *Nature Machine Intelligence* 7:400-411. https://www.nature.com/articles/s42256-025-00986-z
