# 8. Preparing It

*The question: How do we clean, enrich, and use unstructured data?*

> **Questions this chapter answers**
> - How clean does our customer data actually need to be before AI can use it?
> - Our best knowledge lives in calls, emails, and documents. How do we turn it into something a system can act on?
> - What does enrichment really add, and when does it make things worse?
> - Where does personal data get handled, and why does the order of the pipeline matter?

## The short answer

Most companies prepare data for the wrong reader. They clean it for reports and dashboards, where a field is either filled or empty. The new reader is an AI system that has to decide what to say to one person, and it needs something different: facts it can trust, with a source, a date, and a clear meaning.

Four moves do most of the work. **Clean for the decision, not for perfection**: normalize the fields your decisions actually read, and let the rest wait. **Preserve, then process**: keep every raw source before any model touches it, so every derived fact can be traced, re-extracted, and corrected. **Extract with a contract**: turn calls, emails, and documents into typed facts using property descriptions precise enough to act as instructions to the extractor. **Make enrichment argue with what you have**: use outside data to cross-examine existing values, not only to fill blanks.

Sensitive data that has no business in memory is removed on the way in, before extraction and embedding, because what reaches a vector cannot be cleanly taken back out. The payoff is not tidier data. It is the ability to be specific when the evidence supports it and honestly general when it does not.

---

Here is a record from Larkspur Systems, a fictional composite company we follow through this handbook: a mid-market maker of field-service and fleet software with about 40,000 customer accounts. One account, Oakhurst Mechanical (also fictional), looks like this in the CRM: industry "Construction," employees "51-200," stage "Customer," renewal in November, three contacts, last activity 19 days ago.

Now read the transcript of Larkspur's 42-minute check-in call with Oakhurst's operations director. In it she says they are adding about forty vans in the spring. The dispatcher who championed Larkspur left in June. Finance is consolidating vendors this year. The mobile app's offline mode fails at rural job sites. And the renewal decision has moved from her to the CFO.

Five facts that change what Larkspur should say, to whom, and when. The CRM record contains none of them. The transcript sits in a call-recording tool, searchable by date and little else. That gap is the subject of this chapter.

## The record was never the whole customer

CRM fields hold what someone decided in advance was worth a column. What customers actually tell you (what changed, what worries them, who now decides) arrives as language: calls, emails, tickets, notes, documents.

You will often read that 80 to 90 percent of enterprise data is unstructured. I do not use that number. Its usual origin is a 1998 Merrill Lynch report that said "some estimates run as high as 80%," with no source, and Seth Grimes traced and debunked the rule in 2008 [1]. IDC's forecast that 80 percent of the global datasphere would be unstructured by 2025 counts video and sensor data, which says nothing about your customer records [2]. You do not need the statistic. Put one transcript next to its CRM record and count.

This is observed, not measured at population level: in the deployments I have worked on, the language channels held the facts that changed decisions, and the fields held the facts that routed them.

## Clean for the decision, not for perfection

Data quality is genuinely poor almost everywhere. In the best-documented study, Nagle, Redman, and Sammon had 75 executives each inspect the last 100 records their teams created; on average 47 percent of new records had at least one critical error, and only 3 percent of the resulting scores met a basic acceptable standard [3]. Gartner's widely quoted figure that poor data quality costs organizations $12.9 million a year on average comes from self-reported estimates by data-quality vendors' reference customers, so treat it as a signal of scale, not a measurement [4].

The trap is concluding that everything must be clean before anything can be personalized. That project never finishes. Clean what your first decisions (Chapter 6) read:

- **Identity keys**: emails lowercased, domains reduced to the registrable domain, and personal email providers never treated as employers. A Gmail address does not make someone an employee of Google.
- **Enumerations**: one vocabulary for industry, stage, and region, mapped from every source's variants.
- **Dates and units**: one time zone convention, one currency, explicit ranges instead of provider buckets that masquerade as numbers.
- **Nulls with meaning**: "unknown," "not applicable," and "checked, found nothing" are three different states. Collapse them and you lose the difference between ignorance and evidence.

Cleaning is a cost; spend it where a wrong value would change what someone receives.

One rule belongs here, before any model runs: **structured facts stay structured.** In the personalization engine I built, a value already known exactly (a campaign ID, a score computed in code) is written as a value and skips extraction, because a model pass adds cost and turns a certain fact into a probable one. The reverse holds too: in the memory system I built, fields such as a person's name can be declared non-extractable, so they arrive only from a system of record.

## Preserve, then process

Before extraction, there is a rule that sounds like plumbing and decides whether the system can ever be trusted: **preserve, trace, then derive.**

When a transcript arrives, store the original first, with a stable source ID, system of origin, the entities it concerns, observation and retrieval times, and a processing state. Then run the models. A retry must not create a second, apparently independent memory, and lineage must point to the original record, not to the first agent's prose. One call may concern three people and two companies; that is one source with five associations, not five sources.

Only then can you re-extract when your schema improves (it will), explain any statement by walking back to the words someone actually said, and delete properly, because you know every derivative of a source. And an empty check is not evidence: "nothing new in this pull" is not "this account has no activity." Update when you last looked; create no memory.

Derivation must never gate preservation: a document is committed before property extraction runs, and a failed extraction leaves it stored and says so.

![A pipeline of six steps: a source arrives, the raw source is preserved first, content is redacted before the model, then extracted and embedded, the output is redacted again, and derived facts are written with lineage. The preserved original is stored with a stable source ID, system of origin, the entities it concerns, processing state, and observation and retrieval times, and a dashed lineage line runs from every derived fact back to that original rather than to an agent's prose. A retry creates no second memory, and one call about three people and two companies is one source with five associations. Preserving first is what makes it possible to re-extract when the schema improves, explain any statement by walking back to the words actually said, and delete properly because every derivative of a source is known.](/images/handbook/ch07-preserve-then-process.svg)

*Figure 8.1. Preserve, trace, then derive: store the original before any model runs, and every fact can be traced, re-extracted, and deleted.*

Estimate the feed before you turn it on. A useful planning shape:

`source checks x material-change rate x processing cost + entities x analytical tasks x review cycles x activation rate`

Providers may charge for every check, including empty ones; models only consume tokens for records that survive filtering. Model inference is often not the largest cost: data rights, subscriptions, identity resolution, and correction can cost more. Collecting everything is not intelligence. Start with sources that change a defined decision, measure their yield, and turn off the ones that change nothing.

## Redact on the way in

A transcript like Oakhurst's can contain a card number read aloud, a personal mobile number, and an integration key pasted into the meeting chat. The question is not whether to remove sensitive data but where removal actually protects anything.

The answer is before extraction and before embedding. Once personal data is embedded, its signature is in the vector and can influence similarity search even if you filter the text at retrieval time. Retrieval-time filtering hides the problem; it does not remove it.

The rule I recommend is tiered. Secrets, financial identifiers (card numbers validated with the Luhn checksum, IBANs), and government identity numbers are always removed, because no personalization decision needs them; memory keeps "identity was verified," never the passport number. Contact data is configurable: a sales team needs email addresses to match records, a healthcare deployment must strip them. Custom patterns cover identifiers no built-in tier knows, such as policy numbers or internal case IDs. Where linking is needed without the value, a consistent hash says "these five documents mention the same address" without keeping it.

Redaction has a price: what it removes, the extractor cannot extract. Reversible tokenization resolves that for outside models. In the memory system I built, emails, phone numbers, and a supplied list of names become format-preserving stand-ins before the call and are restored locally before anything is written. It is an egress control, not an at-rest one: the record holds real values, so the data is pseudonymized, not anonymized. And a token the model mangles stays a lost value, never guessed back to the wrong person.

The embedder is a separate channel, and tokens cannot protect it because they are not stable between a save and a later query: redact before embedding, or embed locally for regulated data. Map every point where content leaves (extraction, embedding, query embedding, answer synthesis) and log each external call with a hash of what was sent, never the text.

Redaction runs twice: before the model sees the content, and again on the extracted output, to catch what the pattern matcher missed or the model reconstructed. Three production lessons. Early extractors sometimes wrote the placeholder itself into a field, setting an email property to `[EMAIL]`; validate for placeholders. A privacy option that silently does nothing is worse than an error: where a save path could not apply redaction, we made it refuse the option rather than store the body unredacted. And make the guarantee a test: every build of the memory system I built proves that no reversible token can reach a stored record, memory, or graph edge.

## Extract with a contract

Here is the sentence this chapter turns on: **the property description is the extractor's instruction.** What you write to define a field is what the model reads when it decides what to pull from a transcript. A vague description produces vague extraction, and no model upgrade fixes that.

The first extraction prompt is written by hand and works for ten fields. By thirty it contradicts itself, and teams break each other's results by editing one shared prompt. The fix is to move the instructions into each property and compose the prompt at run time from the properties relevant to that content.

A property description that works as a contract has four parts:

1. **What to capture**: the meaning in this organization, and why it is used.
2. **Where to find it**: the signals and sources (transcripts, signatures, support tickets, order forms).
3. **What it is not**: the boundary with neighboring properties and common look-alikes.
4. **Examples**: one clear case, one that needs transformation, one where nothing should be extracted.

For Larkspur, one property looks like this:

```yaml
name: fleet_expansion_plan
type: text
update: replace          # newest plan supersedes older plans
description: >
  Capture: a stated, dated plan to add vehicles, technicians, or service
  regions, with size and timing when given. Sales uses it to time expansion
  conversations; Customer Success uses it to plan onboarding capacity.
  Where to find it: renewal and check-in calls, emails from operations or
  fleet managers, purchase requests, job postings for dispatchers.
  What it is NOT: general optimism ("we are growing"), a competitor's plans,
  a vendor's forecast, or plans the speaker attributes to someone else
  without confirming them.
examples:
  - input: "We're adding about forty vans in the spring."
    output: "Adding ~40 vans, spring (stated by operations director)"
  - input: "Business is good, we'll probably grow."
    output: null
```

![A four-row table showing the parts of a property description that works as a contract, filled in for Larkspur's (fictional) fleet_expansion_plan property. What to capture: a stated, dated plan to add vehicles, technicians, or service regions, used by Sales for timing and by Customer Success for capacity. Where to find it: renewal and check-in calls, operations emails, purchase requests, and job postings for dispatchers. What it is not: general optimism, a competitor's plans, a vendor's forecast, or unconfirmed plans attributed to someone else. Examples: "We're adding about forty vans in the spring" becomes "Adding ~40 vans, spring," while "Business is good, we'll probably grow" becomes null. When extraction is wrong, fix the description and examples before reaching for a bigger model.](/images/handbook/ch07-property-contract.svg)

*Figure 8.2. Extract with a contract: four parts turn a field definition into an instruction the extractor can follow, including when to extract nothing.*

Ownership becomes local: the team that owns a field owns its definition, and editing one cannot silently degrade another. Diagnosis becomes local too: when extraction is wrong, fix the description and examples before reaching for a bigger model, which buys reasoning depth but not clarity. The research agrees: Sainz and colleagues showed that models trained to follow detailed annotation guidelines generalize to unseen extraction tasks, and that detailed guidelines are key to good results [5].

A schema only catches what its designer anticipated, so the same pass should also extract open-set observations: short, sourced statements the schema had no field for. In a controlled test across 20 samples on a system I built, 38 percent of the valuable information was captured only by open-set extraction, 12 percent only by schema extraction, and 16 percent by neither [6]. Neither form alone was enough, which is why the architecture I published pairs open-set atomic facts with schema-enforced typed properties in one memory model [10]. Telling the extractor what kind of entity it is reading matters too: the same company web pages, read "as a company record, like an analyst," produced products, pricing, positioning, and named customers instead of a list of names and titles [7]. This is the quiet decision inside every pipeline: what you choose to extract is the kind of understanding of a person you are building.

## Enrichment should argue with what you have

Enrichment is usually sold as filling blanks. That is half the job. The other half is cross-examination: using new sources to challenge values you already hold.

If a provider reports 100 employees and several current sources describe a global enterprise, the number deserves suspicion; the conflict should surface before any message is written, not be settled by whichever value makes the copy more interesting. I prefer deterministic trust rules: reject counts outside a plausible range, reject values that look like a provider's bucket floor, reject a value that differs from explicit research by a factor of five or more. Once a value is rejected, the model must not be allowed to reason its way back into using it.

Cross-examination is easier on fields than on prose. The engine I built increasingly stores account research as structured fields (headcount, leadership changes, buying triggers, competitors) rather than one paragraph, so each field is trust-filtered on its own and a known fact is rendered, not regenerated. The next step, confidence, provenance, and freshness on every field, is still being standardized.

Uncertainty needs its own state: known, likely, uncertain, rejected. For anything customer-facing, the policy is **hide when uncertain**. That does not freeze personalization; it moves it down a ladder from person-and-account to account to segment, without treating the downgrade as a failure. A specific but false message is worse than a general but useful one.

![Conflicting evidence, a provider reporting 100 employees against current sources describing a global enterprise, flows into deterministic trust rules: reject counts outside a plausible range, reject values that look like a provider's bucket floor, and reject values five times or more off explicit research, with no reasoning back into a rejected value. The result is an explicit state: known, likely, uncertain, or rejected. Uncertain and rejected values are hidden from customer-facing output, which steps personalization down a ladder from person and account, to account, to segment, and stepping down is not a failure.](/images/handbook/ch07-enrichment-cross-examination.svg)

*Figure 8.3. Enrichment as cross-examination: trust rules settle conflicts before any message is written, and uncertain values move personalization down a level instead of into the copy.*

Decay is why enrichment is continuous. The usual claim that B2B data decays 22.5 percent a year has no primary source I can find. A defensible proxy is job tenure: the U.S. Bureau of Labor Statistics reported median tenure with the current employer of 4.1 years in January 2026, with 20.6 percent of wage and salary workers at one year or less [8]. Titles, teams, and authority move on roughly that clock, which is why Oakhurst's champion could leave in June while the CRM still listed her.

## What this does not do

Extraction does not make a fact true. It makes a claim structured. Structured outputs guarantee shape, not accuracy: a perfectly typed field can hold something the speaker never said. Provenance is what lets you check.

Pattern-based redaction and tokenization are high precision, not complete coverage: spelled-out numbers, unusual formats, and unlisted names escape them. Even dedicated tools say so: Microsoft's Presidio documentation states there is no guarantee it will find all sensitive information [9]. Redaction is one control among several (access scoping, retention limits, deletion), never the whole answer.

Single-pass extraction has a ceiling. In the test above, 16 percent of valuable information was missed by both methods, often implications that span several sources. And the standard memory benchmarks measure whether a system recalled a fact correctly, not whether it extracted something an agent can act on [7]. Evaluate on your own decisions.

I also disagree with the popular framing behind the "80 percent unstructured" line [1]. It makes unstructured data sound like a volume problem solved with storage. It is a meaning problem, solved with definitions.

## At scale

At 250,000 contacts, four things change. **The schema becomes shared infrastructure**: dozens of properties across teams stay workable only when each carries its own contract, each call includes only the properties relevant to that content, and names are canonicalized on save so "Deal Stage" and "deal stage" never become two silent keys. **Duplicates dominate**: when five overlapping sources were processed for one entity in a controlled test, 83 percent of candidate memories duplicated something already stored [6], so write-time deduplication is not optional. **Replay becomes routine**: every schema, redaction, or model change raises the question of what to re-extract. Without preserved originals, the answer is "all of it, and hope." **Loads go through a queue, not a loop**: a CRM export, a spreadsheet, or a year of history enters as one bulk job, never as thousands of looped single saves. In the memory system I built, that queue lives in the database, so jobs survive a restart; each row retries with backoff and failures are reported row by row; and ingest dedupes by a content and entity hash, so re-running an import after a timeout does not write twice. A provider's batch interface cuts the model price roughly in half for up to a day of latency, and a failed batch falls back to the normal worker, so no record is lost.

## Failure story: The Quality Trap

We once ran extraction over a set of company websites. The first pass returned 12 memories for one company, five of them names and titles. We added aggressive splitting to "capture more," and the count jumped to 101. The dashboard looked like a breakthrough.

Then we read the output: 149 near-duplicate pairs, 14 junk entries such as copyright notices and "has a careers page," and a product's feature list shattered into ten fragments. Two rules fixed it: keep lists together unless an item carries its own detail, and filter boilerplate. The count dropped to between 18 and 22, and every remaining memory was usable [7]. More facts in the store means more noise in every context window that reads them. Count is not quality.

## Patterns

**Extract with a Contract.** *Problem*: one hand-written prompt cannot scale past a few dozen fields owned by many teams. *Solution*: each property carries the four-part description; prompts are composed per call from relevant properties; open-set extraction runs alongside. *Tradeoffs*: authoring work per property, and an evaluation harness to prove it.

**Preserve then Process.** *Problem*: derived facts cannot be traced, replayed, or deleted. *Solution*: store the raw source with identity, origin, timestamps, and processing state before any model runs; derive idempotently; keep lineage. *Tradeoffs*: raw content carries its own retention and redaction obligations.

**Enrichment as Cross-Examination.** *Problem*: enrichment layers conflicting values onto records and the model picks one. *Solution*: compare sources with deterministic trust rules, give uncertainty an explicit state, and hide uncertain values from customer-facing output. *Tradeoffs*: less specific output for some records, by design.

## Leader questions

1. Which three decisions are we preparing data for, and which fields do they actually read?
2. When our system states a fact about a customer, can we show the original words or record it came from?
3. Where in our pipeline is personal data removed: before or after it reaches a model and an index?
4. When two sources disagree about a customer, who or what decides, and is that rule written down?
5. How many of our extracted facts were used in a decision last month?

## Build checklist

- [ ] Clean identity keys, enumerations, dates, and units for your first decisions; treat personal email domains as non-employers.
- [ ] Store every raw source before processing, with source ID, origin, entities, timestamps, and processing state.
- [ ] Make extraction idempotent; retries never create new memories; load backfills and imports as bulk jobs, not looped saves.
- [ ] Write known values directly, never through extraction; declare the fields that must never be inferred.
- [ ] Redact secrets, financial, and identity numbers before extraction and embedding; configure contact data per use case; tokenize identifiers bound for an outside model; treat the embedder as its own channel; scan outputs again, including for placeholder values.
- [ ] Write every property description in four parts, with a negative example.
- [ ] Run open-set extraction in the same pass as typed extraction.
- [ ] Encode trust rules for enriched values and a hide-when-uncertain policy.
- [ ] Deduplicate at write time; estimate the feed cost before switching a source on.

## Metrics to watch

- **Traceability rate**: share of customer-facing facts that link to a preserved source.
- **Extraction precision per property**: sampled against human labels, using each property's own criteria.
- **Noise and duplicate ratio** in newly written memories.
- **Enrichment conflict rate**: share of enriched values rejected or flagged, by provider.
- **Source yield**: decisions changed per thousand source checks.

## Reader Q&A

**Do we need a data-cleansing project before we start?** No. Clean the fields your first decisions read, and let the rest wait until a decision needs it.

**Should we store transcripts, or only what we extract from them?** Both, under retention rules. The extraction is what agents read; the original is what lets you verify, re-extract, and delete.

**Can a larger model replace good property descriptions?** Rarely. A stronger model reasons more; it does not know what your organization means by "decision maker." Fix the description first.

## For your AI

```yaml
chapter: 8
concepts:
  - name: Extract with a Contract
    definition: Each property carries what to capture, where to find it, what it is not, and examples; this description is the extractor's instruction.
  - name: Preserve then Process
    definition: Store the raw source with identity, origin, timestamps, and processing state before any model derives facts from it.
  - name: Enrichment as Cross-Examination
    definition: Outside data is used to challenge existing values as well as fill blanks, under deterministic trust rules.
  - name: Redact on the Way In
    definition: Remove sensitive data before extraction and embedding, and scan extracted output again.
decision_rules:
  - if: "a field is not read by any prioritized decision"
    then: "defer cleaning it"
  - if: "an enriched value conflicts with current evidence beyond a set threshold"
    then: "reject it and hide it from customer-facing output"
  - if: "extraction for a property is wrong"
    then: "revise the property description and examples before changing model tier"
  - if: "a source check returns nothing new"
    then: "update last-checked time; create no memory"
  - if: "a value is already known exactly from a system of record or code"
    then: "write it as a structured value; do not route it through extraction"
  - if: "content is sent to a model or embedder outside your infrastructure"
    then: "redact always-removed tiers, tokenize identifiers for the chat model, and redact or run locally for the embedder"
  - if: "loading a backfill, CRM export, or spreadsheet"
    then: "submit it as one durable, idempotent bulk job with per-row status, not a loop of single saves"
assessment_questions:
  - "Which systems hold calls, emails, and documents, and is any of it extracted into typed fields today?"
  - "Are raw sources retained with lineage to every derived fact?"
  - "At which pipeline step is personal data redacted, and is the embedding call covered?"
  - "How are conflicts between enrichment providers and CRM values resolved?"
patterns: [Extract with a Contract, Preserve then Process, Enrichment as Cross-Examination]
anti_patterns: [The Quality Trap, Retrieval-Time Redaction, The Placeholder Value, Clean Everything First, The Silent Privacy Option]
maturity_dimension: understanding
```

## References

1. Seth Grimes, "Unstructured Data and the 80 Percent Rule," Breakthrough Analysis / Clarabridge Bridgepoints, 2008-08-01. http://altaplana.com/articles.html (catalog: http://metadatace.cci.drexel.edu/omeka/items/show/14556). Traces the claim to Shilakes and Tylman, Merrill Lynch, 1998.
2. VentureBeat, "Report: 80% of global datasphere will be unstructured by 2025" (IDC forecast). https://venturebeat.com/data-infrastructure/report-80-of-global-datasphere-will-be-unstructured-by-2025
3. Tadhg Nagle, Thomas C. Redman, David Sammon, "Only 3% of Companies' Data Meets Basic Quality Standards," Harvard Business Review, 2017-09-11. https://hbr.org/2017/09/only-3-of-companies-data-meets-basic-quality-standards
4. Gartner, "Data Quality: Why It Matters and How to Achieve It." https://www.gartner.com/en/data-analytics/topics/data-quality
5. Oscar Sainz et al., "GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction," ICLR 2024. https://arxiv.org/abs/2310.03668
6. Hamed Taheri, "Dual Memory: Free Text and Typed Properties," hamedtaheri.com. /articles/dual-memory-free-text-and-typed-properties
7. Hamed Taheri, "Beyond Fact Count: Measuring What Actually Matters in Agent Memory Extraction," hamedtaheri.com, 2026-04-08. /articles/beyond-fact-count-agent-memory
8. U.S. Bureau of Labor Statistics, "Employee Tenure in 2026," released 2026-09-24. https://www.bls.gov/news.release/tenure.nr0.htm
9. Microsoft Presidio, FAQ. https://microsoft.github.io/presidio/faq/
10. Hamed Taheri, "Governed Memory: A Production Architecture for Multi-Agent Workflows," arXiv:2603.17787, 2026-03-18. https://arxiv.org/abs/2603.17787

Further reading on the site: "Schema-Guided Extraction," "Guided Memory Extraction," "4-Tier PII Redaction," "Two-Phase Redaction," and "Building an Enterprise Personalization Engine" (sections on enrichment and data trust).
