One Customer, Many Systems

How do we unify a customer across CRM, warehouse, and tools without merging two people?

Working draft, revised September 2026 · 25 min read · Markdown for your AI

Questions this chapter answers

  • Our customer lives in the CRM, the warehouse, the product database, and the support desk. How do we make them one customer?
  • What is identity resolution, and how accurate does it have to be?
  • Do we need a CDP, or is the warehouse enough? Where does AI memory sit relative to both?
  • When should two records that look alike stay separate?
  • Which system wins when two of them disagree?

The short answer

Every personalization system eventually has to answer a question that sounds trivial: is this the same person? The CRM has a contact. The product has a user. The support desk has a requester. The event pipeline has an anonymous browser that later filled in a form. Each system is right about its own records and silent about everyone else's.

Unifying them is identity resolution, the least glamorous and most consequential step in the stack. Memory, decisions, generation, and measurement all inherit its mistakes. A missed match makes you forgetful: you ask a customer something you already know. A false match makes you reckless: you tell one person something that belongs to another. The second error costs far more, and most systems are tuned as if the two were equal.

Three decisions matter more than any tool choice. Match deterministically where strong shared keys exist and probabilistically only where you must, with explicit confidence. Treat a merge as a reversible, logged event, not an overwrite. And let identity confidence control how specific you are allowed to be: when the system is unsure who someone is, it should say less, not guess more.

On the stack: CRM, warehouse, CDP, and event pipelines each do one job well. None of them alone is the customer's memory. Decide which system is the source of record for each field, write it down, and make every integration obey it.

The sentence to take away: if identity confidence falls, personalization specificity should fall with it.


Two Sam Patels

Larkspur Systems is a fictional composite company used throughout this handbook.

Larkspur sells field-service and fleet software to about 40,000 customer accounts. One of them is a transport group with two subsidiaries that buy separately: a freight business and a rail-maintenance business. Each has its own contract, renewal date, and negotiated price. Both use the parent's email domain.

Both also employ someone named Sam Patel.

A cleanup job, written to reduce duplicates before an AI renewal pilot, matched contacts on first name, last name, and email domain. The two Sam Patels became one contact, and one email address survived. Three weeks later the renewal agent, working exactly as designed, pulled the freight subsidiary's contract, discount, and usage history and wrote a clear, well-grounded renewal note. It went to the Sam Patel at the rail subsidiary, who forwarded it to his procurement lead with one line: "Why are they charging us more?"

Nothing in the generation was wrong. Every fact in the email was true, sourced, and current. It was true about someone else.

I call this The Wrong Merge, and it deserves study because it defeats every safeguard downstream of it. Grounding checks pass, because the facts are in memory. Tone checks pass. A human reviewer approves, because the email is good. The only layer that could have stopped it is the one that decided two records were one person.

This problem is older than the CRM

In 1946 the US public-health statistician Halbert Dunn put it in terms that still hold: "Each person in the world creates a Book of Life. This Book starts with birth and ends with death. Its pages are made up of the records of the principal events in life. Record linkage is the name given to the process of assembling the pages of this Book into a volume." [1]

Your CRM, warehouse, and support desk each hold loose pages. Newcombe and colleagues ran the first computerized probabilistic linkage of vital records in 1959 [2], and in 1969 Ivan Fellegi and Alan Sunter gave the field the formal theory most modern matchers still descend from [3].

The failure mode is old too. On February 27, 2011, Sergio Ramirez went to a Nissan dealership in Dublin, California, to buy a car. His TransUnion credit report carried an alert that his name matched a name on the US Treasury's sanctions list. The dealer refused to sell to him; his wife bought the car in her own name. The Supreme Court's opinion records that TransUnion "did not compare any data other than first and last names" and that the product "unsurprisingly" generated many false positives [4]. A class of 8,185 people sued; in 2021 the Court held that the 1,853 whose reports had gone to third parties had standing to recover damages [4].

The stakes there were higher than a renewal email, but the mechanism is the same as the Sam Patel merge: attributes that feel like identity, treated as proof. Names are not identifiers. Domains are not identifiers. They are evidence, and evidence has weight.

Identity is a claim with a confidence, not a key

Entity resolution is the general problem: deciding which records refer to the same real-world thing, whether a person, company, product, or address. Identity resolution is the customer-data case: which records, identifiers, and events belong to the same person or organization, including stitching anonymous browsers and devices to known people. The mathematics is shared; the evidence differs.

Before matching anything, decide what kinds of things exist. It is tempting to create types called Customer, Lead, Partner, and Employee. That becomes brittle quickly. One person can be a product user, the billing contact on an account, a champion on an open deal, and next year a buyer somewhere else. Mixing identity with role creates duplicates and hides change.

Use a small set of durable entities instead:

EntityLarkspur exampleWhy it is separate
PersonEach Sam PatelA person persists while jobs and roles change
OrganizationThe group; each subsidiaryContracts belong to legal entities, not people
AccountThe freight subsidiary's Larkspur accountBilling, entitlements, and renewals attach here
RoleFleet operations lead at the freight subsidiary, since MarchA dated link between a person and an organization
HouseholdRare in B2B; central in consumer businessesShared addresses, devices, and payment

"Lead", "customer", and "champion" then become dated relationships among entities rather than containers. A person works at an organization; an organization owns an account; a person is the billing contact for an account from one date to another. This is what lets a system notice that someone left instead of writing to their old job, and it would have told Larkspur it had two people with two roles at two organizations, not one person with a messy history.

Model entities, not dossiers. A dossier is a pile of facts attached to a name. An entity has stable identity, typed relationships, and a history.

Deterministic first, probabilistic where the evidence is thin

Deterministic matching links records that share a strong identifier exactly: a product user ID, a CRM ID carried into the warehouse, a verified email address, a billing account number. It is fast, explainable, and auditable. It only works when a clean shared key exists, and often one does not: people use two addresses, company names appear with and without legal suffixes, trials start from personal email.

Probabilistic matching weighs many weaker signals together. In the Fellegi-Sunter model [3], each compared field (name, email, phone, company, address) has two probabilities: how often it agrees when two records truly match (m), and how often it agrees by coincidence when they do not (u). Agreement on a rare surname is strong evidence; agreement on "John Smith" is weak. The weights add up to a score compared against two thresholds: above the upper, link; below the lower, non-link; in between, a possible link that goes to review.

That three-way outcome is the most useful idea in the field, and most customer-data pipelines throw it away by forcing every pair into "merge" or "don't."

A record pair's match score is compared against two thresholds. Below the lower threshold is a non-link: the records stay separate entities. Between the thresholds is a possible link: a candidate link only, with no facts crossing it, high-consequence pairs routed to review, and any model reasoning logged as evidence. Above the upper threshold is a link. Field agreements carry different weights: agreement on a rare surname is strong evidence, agreement on "John Smith" is weak.

Figure 9.1. Fellegi-Sunter gives three outcomes; the possible-link band is where most pipelines lose information by forcing a yes or no.

Open-source tools make this practical at company scale. Splink, built by the UK Ministry of Justice's data linking team, implements Fellegi-Sunter with parameters estimated by expectation-maximization; its authors report linking a million records on a laptop in about a minute, with users including the UK Office for National Statistics and the Australian Bureau of Statistics [5]. Zingg takes an active-learning approach on Spark, Databricks, and Snowflake [6]; Senzing sells a commercial engine [7]. That is landscape as of September 2026, not a recommendation.

Two mechanics decide whether any of this runs. Blocking: Larkspur's 250,000 contacts make about 31 billion possible pairs (250,000 squared, halved). You compare only within blocks of plausible matches (same domain, same postcode, similar surname), and a pair that never shares a block can never match, so test your blocking rules for lost recall. Normalization: lowercase emails, strip legal suffixes, reduce domains to their registrable form, and enforce one rule that is violated constantly: a personal email provider is not an employer. A Gmail address should not inherit Google's firmographics. In the governed personalization engine I built, a freemail address is a flag with consequences: company research suppressed, firmographics hidden, lower-specificity copy, extra QA checks. The same over-trust of domains merged the two Sam Patels.

Language models are good in the gray zone, reading two records and explaining whether "Harrow Rail Services Ltd" and "Harrow Rail" are the same operating company. They are poor as the primary matcher: expensive, hard to reproduce, and confidently wrong. In the self-hosted memory system I built, that split is enforced. A deterministic finder proposes pairs from three probes: a strong key already pointing at a different record (two imports disagreeing about who owns it), the same normalized name and the same parent organization (a name alone never qualifies), and a shared secondary identifier such as a phone number. A group of more than 50 records sharing one name and parent, or one phone number, is treated as a placeholder value (a switchboard number, a default), not 50 duplicates. A model then judges each pair and must return merge, with a confidence and a reason, or keep distinct. Merges below a 0.9 confidence floor are refused, and an applied keep-distinct verdict stops the pair from being proposed again: knowing two Sam Patels are two people is evidence too.

The two errors are not symmetric

A false merge joins two different entities. A missed merge leaves one entity split across records. Tighten thresholds and false merges fall while missed merges rise.

Most teams tune for the error they can see. Duplicates show up on every dashboard; false merges are invisible until they hurt someone. For personalization that is backwards:

  • A missed merge makes the system forgetful: a repeated question, a welcome email to someone already onboarded, a long-standing account treated as cold. Annoying, recoverable, usually private.
  • A false merge makes the system wrong about who someone is: one person's data disclosed to another, one account's pricing applied to a different account. It can breach a contract or a privacy obligation, and it contaminates every memory, score, and summary built on the merged record.

So the rubric I use for merge confidence is expressed as what the system is allowed to do:

Merge confidenceTypical evidenceWhat the system may do
CertainShared system-issued ID, or a verified email the person controlsMerge automatically; person-level facts allowed
HighSeveral independent strong signals agreeLink automatically; person-level facts on owned channels
PossibleMixed evidence: the Fellegi-Sunter gray zoneCandidate link only; no facts cross it; review high-consequence pairs
WeakName only, or name plus a shared domainDo not link

The Sam Patel merge was weak evidence treated as certain.

Here is the hinge of this chapter: if identity confidence falls, personalization specificity should fall with it. A system sure of the person can use person-level facts. A system sure of the company but not the person should personalize to the account and role. A system sure of neither should be honestly general. Identity confidence is not a data-quality footnote. It is an input to every decision about what the system may say.

A three-step ladder in which identity confidence falls from top to bottom. When the system is sure of the person, it may use person-level facts. When it is sure of the company but not the person, it personalizes to the account and role. When it is sure of neither, it stays honestly general. The resolution service returns one of four permitted levels: person, account, segment, or general.

Figure 9.2. Confidence controls specificity: the less sure the system is of who someone is, the less specific it is allowed to be.

In the governed personalization engine I built, this is code, not a prompt instruction. Relationship statements ("your renewal is coming up", "your current contract") are blocked unless the relationship is known, and the copy is reframed as an evaluation. Weak evidence steps a record down a ladder: person and account, account, segment, approved static fallback, fail closed. It would not have caught the Sam Patel email, though: the bad merge made the relationship look known.

A merge is an event, not an overwrite

The common implementation picks a survivor record, copies some fields over, and archives the loser. It is simple, and it destroys the evidence you need when the merge is wrong. Once the two Sam Patels were one row, there was no clean way back: whose phone number, whose tickets, whose call notes?

The alternative is to link, not overwrite. Keep every source record intact with its source system and native ID. A separate identity layer says "these source records resolve to entity E-1234", with confidence, evidence, the deciding rule or model, and a timestamp. A merge is a log entry; an unmerge is another. Views are rebuilt from links, so reversing a bad decision restores the prior state instead of requiring archaeology.

My own memory system uses a close cousin: a redirect with a journal. Both IDs stay valid and reads of the absorbed one land on the survivor; properties fold only where the survivor has no value or the absorbed one is strictly newer; the journal keeps a pre-image, so unmerge restores the old state. The limits: memories written after the merge stay with the survivor, and the journal expires (90 days by default). Merges also start propose-only; an operator promotes the judge, not each verdict, once its proposals hold up.

This matters more once AI memory sits on top. Extracted facts, account summaries, and scores were all written against an entity ID. When the entity splits, each memory must follow the source record it came from, which is only possible if every memory points to its source. Provenance is what makes identity mistakes repairable. W3C's PROV model offers a vocabulary to borrow (an entity was generated by an activity, was attributed to an agent) [8].

Anonymous to known is a promotion, not a merge

Much behavior arrives before you know who someone is. Event pipelines give a browser or device an anonymous ID and link it to a known one when the person logs in or fills a form. Treat that transition as a promotion with its own rules:

  • A browser is not a person. Shared laptops and office machines mean one anonymous ID can be several people. Attach anonymous history only for the session that led to identification, or within a window you can defend.
  • Anonymous identifiers are fragile. As of September 2026, Chrome still allows third-party cookies by default while Safari and Firefox block them [9], and WebKit has capped script-set first-party cookies since 2019 (seven days under ITP 2.1) [10]. Do not build identity that needs a cookie to survive for months.
  • The identifying moment carries evidence. A verified login is certain. A typed-in email on a form is weaker: people mistype, and people enter a colleague's address.

Packaged tools encode similar protection. Segment, for example, lets teams set per-identifier limits and a priority order; by default a profile holds one user_id, and when an event would push a profile past a limit, the lower-priority identifier is demoted and resolution retried rather than profiles being merged past the limit [11]. Whatever your tool, find the merge protection and configure it deliberately.

Where each system is good, and where memory sits

The CRM (Salesforce, HubSpot, and peers) is the system of record for relationships and commercial workflow: accounts, contacts, deals, owners, activities. It carries the scars of every import and is rarely the best place to resolve identity at scale.

The warehouse or lakehouse (Snowflake, Databricks, BigQuery, and peers) holds every source side by side, which makes it the natural home for matching jobs.

The customer data platform. The CDP Institute defines a CDP as "packaged software that creates a persistent, unified customer database that is accessible to other systems" [12]. Packaged CDPs bundle event collection, identity, profiles, and activation. The composable CDP builds the same capabilities on the warehouse and uses reverse ETL to push modeled fields and audiences back into operational tools. The market has moved both ways: Twilio paid $3.0 billion for Segment in 2020 [13], Salesforce launched a "zero copy" partner network in 2024 to use warehouse data without moving it [14], and in 2025 mParticle merged with Rokt and Fivetran agreed to acquire Census [15]. Read that as consolidation, not a verdict.

Event pipelines (Segment and peers) collect behavior and handle anonymous-to-known stitching.

Where AI memory sits. None of these was designed for the reader that now matters most: an agent that needs, in one call and under a token budget, the typed facts, evidence, and synthesized understanding of one resolved entity (Chapter 10). The placement rule is simple: memory sits downstream of identity resolution and upstream of decisions. It consumes resolved entity IDs and never invents its own. If each agent keeps its own idea of who a customer is, you have rebuilt the duplicate problem inside your AI layer, only faster.

The system-of-record map

Unification fails less on algorithms than on one unanswered question: when two systems disagree, which wins? Answer it field by field.

FieldSource of recordOn conflict
Contract terms, renewal dateBilling / contract systemBilling wins; flag CRM mismatch
Account owner, deal stageCRMCRM wins
Product usage, active usersProduct data via warehouseWarehouse wins
Job title, roleMost recent verified sourceEnrichment may challenge, never overwrite
Consent, communication preferencesPreference centerAlways wins
AI-extracted factsMemory layerWritten to namespaced fields; never overwrites human entries

Two write-back rules keep integrations from corrupting what they integrate. Agents write to namespaced fields, never on top of fields people maintain. Every write is logged with writer, time, and diff, and bulk write-backs start in dry-run. I apply both in an open-source CRM operations library I maintain, and they are the difference between a revenue leader trusting agent writes and banning them.

Enrichment gets one more rule: it should challenge what you hold, not only fill blanks. When a vendor's headcount disagrees with three current sources, surface the conflict instead of letting the newest value win.

Identity as a first-class subsystem

Put the pieces together and identity stops being a cleanup job. It becomes a subsystem with a data model, an owner, and a service contract. I call the shape the Identity Spine:

  1. Stable internal IDs for every person, organization, and account, never changed or reused. Source IDs are attributes, not keys.
  2. An identifier table: every email, phone, CRM ID, product ID, anonymous ID, and domain with its source, first-seen and last-seen times, and verification status.
  3. Typed, dated links: person to organization, person to account, account to parent, each with type, dates, and source. A role is a link, not a column.
  4. A merge log: confidence, evidence, rule version, and timestamp for every link. Unmerge is a first-class operation.
  5. One resolution service: memory, decisions, generation, and analytics all ask it "who is this?" and receive an entity ID, a confidence, and a permitted specificity level.

Five source systems (CRM, warehouse, product database, support desk, event pipeline) feed the Identity Spine, which holds four layers: stable internal IDs that are never changed or reused; an identifier table with source, first and last seen, and verification status; typed, dated links where a role is a link rather than a column; and a merge log with confidence, evidence, and rule version. A single resolution service sits on top and answers "who is this?" for memory, decisions, generation, and analytics. Every answer carries an entity ID, a confidence, and a permitted specificity.

Figure 9.3. The Identity Spine: every consumer asks one resolution service, so "confidence controls specificity" is enforced in one place.

The fifth element is the one teams skip. Without it, marketing resolves identity one way, support another, and each new agent a third.

Resolve by every identifier the business uses (phone, device ID, a CRM or deal ID), with one normalization shared by writes and reads. In one release of my memory system, records saved under a CRM key could not be read back by it, so a CRM-keyed sync minted duplicates. Silent splits come from asymmetry as often as from bad matching.

The spine also makes relationships queryable. Store every relationship identifier at write time (a contact carries both the person's email and the company's domain) and "every open ticket at this contact's company" becomes a lookup rather than a guess.

Identity also has a scope: the same account can sit in two campaigns, the same record in two tenants. In that engine, organization and campaign are resolved before any research or generation, and page identity is scoped by campaign so one program cannot overwrite another's page for the same company. Correct content in the wrong tenant or campaign is still a critical failure.

What this does not do

There is no perfect match rate. Accuracy depends on the data, not just the algorithm. In healthcare, where investment is large and errors are dangerous, research cited by Pew (by the contractor Audacious Inquiry, for the federal health IT office) found match rates "as low as 50 percent" between organizations, even ones on the same electronic health record vendor [16]. Your CRM is not cleaner than a hospital's records. Plan for residual error and design downstream systems to survive it.

"Golden record" and "360-degree view" are aspirations, not deliverables. I agree with the CDP Institute's goal of a "persistent, unified customer database" [12]. Where I part company is with how the idea is usually sold: as one overwritten row per person that everyone can trust equally. A unified customer is a set of linked source records with confidence and provenance. A flattened profile that hides which fields are certain and which are guesses is less trustworthy than the fragments it replaced.

Zero-copy does not resolve identity. Querying the warehouse in place removes a copy; it does not decide whether two rows are one person.

Models do not replace matching discipline. A model asked "are these the same person?" answers fluently every time; measure it against labeled pairs like any other matcher.

Resolution is not permission. Linking product usage to a marketing profile is easy; whether you may use one for the other depends on consent and purpose (Chapter 19).

At scale

  • Resolution becomes incremental. New records match against existing entities as they arrive; periodic full runs catch drift. Keep both on the same rules, or they will disagree.
  • The gray zone outgrows human review. Route only high-consequence pairs (open deals, active contracts) to people; leave the rest unlinked.
  • Identity is a production API. It needs an owner, versioned rules, and change control. A threshold change is a release.
  • Corrections must propagate. An unmerge has to reach memory, audiences, CRM fields, and cached summaries, so track which consumers read which entity.
  • Measurement depends on it. Holdouts (Chapter 21) assume one person sits in one group. Duplicates leak people across groups; false merges mix them.

Failure story: The Wrong Merge

  • What happened: two people with the same name at sibling companies sharing a domain were merged; one subsidiary's pricing went to the other.
  • Root cause: name plus domain treated as certain; parent and subsidiaries modeled as one organization; survivor-record merge with no log.
  • Fix: one organization entity per contracting subsidiary; name plus domain demoted to weak; contract facts limited to contacts linked to that account at high confidence or better; link-and-log merges so the split was reversible.

The anti-pattern in one line: deduplicating for tidiness before deciding what a false merge costs.

Patterns

Identity Spine. Problem: each system and agent resolves identity differently. Solution: stable internal IDs, identifier table with provenance, typed dated links, merge log, one resolution service. Tradeoff: a real subsystem to own; consumers give up private matching logic.

Link, Don't Overwrite. Problem: merges destroy the evidence needed to undo them. Solution: keep source records intact, resolve through logged links, make unmerge an operation. Tradeoff: more complex queries; views must be materialized.

Confidence Controls Specificity. Problem: downstream systems personalize equally hard whatever identity confidence is. Solution: the resolution service returns a permitted level (person, account, segment, general) and decisions and generation obey it. Tradeoff: some messages become generic; reduce that share by improving evidence, not by lowering the bar.

Leader questions

  1. Which system is the source of record for each customer field, and is that written down?
  2. If two records are merged by mistake today, can we undo it, and how long does it take?
  3. Do we measure false merges, or only duplicates?
  4. Do all our AI agents get the same answer to "who is this?"
  5. What share of personalized messages go out at person, account, and segment level?

Build checklist

  • Durable entity types defined; roles modeled as dated links.
  • Internal IDs issued, never reused; source IDs stored as attributes.
  • Normalization rules, including personal-email-is-not-an-employer and legal-suffix handling.
  • Deterministic rules for strong keys; Fellegi-Sunter-style matching for the rest, with link / possible / non-link thresholds.
  • Blocking rules tested for lost recall on a labeled sample.
  • Merge log with confidence, evidence, and rule version; unmerge tested end to end.
  • Model judges only finder-proposed pairs, returns merge or keep distinct, and starts in a propose-only stage.
  • Resolution by every business identifier, with one normalization shared by write and read paths.
  • Anonymous-to-known promotion rules, including shared devices.
  • System-of-record map by field with conflict precedence.
  • Write-backs to namespaced fields, logged with diffs, dry-run first.
  • One resolution service returning entity ID, confidence, and permitted specificity.

Metrics to watch

  1. False merge rate (precision) from a regularly labeled sample of links, reported beside the missed merge rate (recall).
  2. Gray-zone backlog: possible links waiting, and their age.
  3. Specificity mix: share of decisions at person, account, segment, and general level.
  4. Time to unmerge across every downstream consumer.
  5. Source-of-record conflicts by field, and their trend.

Reader Q&A

Do we need a CDP? Not necessarily. If data already lands in a warehouse and you can model it, a composable approach does the same work. A packaged CDP makes sense when you need real-time collection and activation without building them. Either way, the rules, thresholds, and system-of-record map are yours.

How accurate must matching be? Accurate enough for the action you take. For a generic newsletter, missed merges cost little. For anything carrying contract, pricing, health, or financial facts, false merges must be near zero, which means refusing weak links and accepting more duplicates.

Should we merge a contact who changed companies? The person is the same; the role is not. Keep one person, end-date the old role, and never carry the old employer's account facts into messages to the new one.

For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 9
title: "One Customer, Many Systems"
concepts:
  - name: Identity Resolution
    definition: "Deciding which records, identifiers, and events across systems belong to the same person or organization."
  - name: Entity Resolution
    definition: "The general problem of deciding which records refer to the same real-world thing; identity resolution is its customer-data case."
  - name: Fellegi-Sunter Model
    definition: "Probabilistic record linkage that weighs field agreements by m and u probabilities and classifies pairs as link, possible link, or non-link."
  - name: Identity Spine
    definition: "Stable internal IDs, an identifier table with provenance, typed dated links, a merge log, and one resolution service returning entity ID, confidence, and permitted specificity."
  - name: Merge Confidence Rubric
    definition: "Certain, high, possible, weak; each level defines whether records may be linked and which facts may cross the link."
  - name: System-of-Record Map
    definition: "A per-field table naming the authoritative source and which source wins on conflict."
  - name: Finder-Judge Split
    definition: "A deterministic finder proposes candidate pairs; a model only judges each as merge or keep distinct, under a confidence floor and an operator-gated stage."
  - name: Anonymous-to-Known Promotion
    definition: "Rules for attaching anonymous device or browser history to a known person, handled separately from ordinary merges."
decision_rules:
  - if: "two records match only on name, or on name plus a shared email domain"
    then: "do not link; treat as separate entities"
  - if: "an email address is from a personal email provider"
    then: "do not infer employer or firmographics from the domain"
  - if: "identity confidence for the recipient is below high"
    then: "personalize at account, segment, or general level; never use person-level or contract-level facts"
  - if: "a pair falls between the link and non-link thresholds"
    then: "record a candidate link, share no facts across it, and route high-consequence pairs to review"
  - if: "two systems disagree on a field"
    then: "apply the system-of-record map; enrichment may challenge but not overwrite"
  - if: "an AI agent or integration writes back to the CRM"
    then: "write to namespaced fields, log writer, time, and diff, and start in dry-run"
  - if: "a merge is found to be wrong"
    then: "unmerge through the merge log and propagate the correction to every consumer, including memory"
  - if: "a language model is used to judge a candidate pair"
    then: "judge only finder-proposed pairs; require merge or keep distinct with a reason; refuse merges below a confidence floor; start propose-only"
  - if: "a customer-facing statement asserts a relationship (renewal, contract, product ownership)"
    then: "block it unless the relationship is known at high confidence; otherwise reframe as an evaluation"
assessment_questions:
  - "Which systems hold customer identity today, and which one wins on conflict for each field?"
  - "How are duplicates found and merged today, and can a merge be undone?"
  - "Do you measure false merges separately from duplicates?"
  - "Do all agents and channels use one resolution service, or does each resolve identity itself?"
  - "How are anonymous visitors linked to known contacts, and how are shared devices handled?"
  - "Are subsidiaries and parent companies modeled as separate organizations with their own contracts?"
patterns: [Identity Spine, Link Don't Overwrite, Confidence Controls Specificity]
anti_patterns: [The Wrong Merge, Survivor-Record Merge, Domain Equals Employer, Per-Agent Identity]
maturity_dimension: identity_and_memory
related_chapters: [7, 8, 10, 11, 19, 21]

References

  1. Dunn, H. L. "Record Linkage." American Journal of Public Health 36(12): 1412-1416, 1946. https://pmc.ncbi.nlm.nih.gov/articles/PMC1624512/
  2. Newcombe, H. B., Kennedy, J. M., Axford, S. J., James, A. P. "Automatic Linkage of Vital Records." Science 130(3381): 954-959, 1959. https://www.science.org/doi/10.1126/science.130.3381.954
  3. Fellegi, I. P., Sunter, A. B. "A Theory for Record Linkage." Journal of the American Statistical Association 64(328): 1183-1210, 1969. https://www.tandfonline.com/doi/abs/10.1080/01621459.1969.10501049
  4. TransUnion LLC v. Ramirez, 594 U.S. ___ (2021), decided June 25, 2021. https://www.supremecourt.gov/opinions/20pdf/20-297_4g25.pdf
  5. Linacre, R., et al. "Splink: Free software for probabilistic record linkage at scale." Real World Data Science (Royal Statistical Society), 2023-11-22. https://realworlddatascience.net/applied-insights/case-studies/posts/2023/11/22/splink.html
  6. Zingg (open-source entity resolution). https://github.com/zinggAI/zingg (accessed 2026-09-26)
  7. Senzing (entity resolution software). https://senzing.com/ (accessed 2026-09-26)
  8. W3C. "PROV-DM: The PROV Data Model." W3C Recommendation, 2013-04-30. https://www.w3.org/TR/prov-dm/
  9. MDN Web Docs. "Third-party cookies." https://developer.mozilla.org/en-US/docs/Web/Privacy/Guides/Third-party_cookies (accessed 2026-09-26)
  10. WebKit. "Intelligent Tracking Prevention 2.1." 2019-02. https://webkit.org/blog/8613/intelligent-tracking-prevention-2-1/
  11. Twilio Segment Docs. "Identity Resolution Settings." https://www.twilio.com/docs/segment/unify/identity-resolution/identity-resolution-settings (accessed 2026-09-26)
  12. CDP Institute. "What is a CDP?" https://www.cdpinstitute.org/what-is-a-cdp/ (accessed 2026-09-26)
  13. Twilio Inc. Form 10-K for fiscal year 2020 (Segment acquisition). https://www.sec.gov/Archives/edgar/data/1447669/000144766921000070/twlo-20201231.htm
  14. Salesforce. "Salesforce Launches Zero Copy Partner Network." 2024-04-25. https://www.salesforce.com/news/press-releases/2024/04/25/zero-copy-partner-network/
  15. Rokt and mParticle merger announcement, 2025-01-16, https://www.prnewswire.com/news-releases/rokt-and-mparticle-merge-to-redefine-real-time-relevance-302352650.html ; Fivetran, agreement to acquire Census, 2025-05-01, https://www.fivetran.com/press/fivetran-signs-agreement-to-acquire-census-delivering-the-first-end-to-end-data-movement-platform-for-the-ai-era
  16. The Pew Charitable Trusts. "Enhanced Patient Matching Is Critical to Achieving Full Promise of Digital Health Records." 2018-10-02. https://www.pew.org/en/research-and-analysis/reports/2018/10/02/enhanced-patient-matching-critical-to-achieving-full-promise-of-digital-health-records

This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.

Get chapters by email as they are revised

Prefer LinkedIn? Subscribe to the newsletter there instead.