Questions this chapter answers
- The data is ours and the inference is correct. Why does the customer still feel watched?
- What should our systems never infer or store, even where the law would allow it?
- What does the law actually require of AI personalization, in principles that will still hold next year?
- How do we build privacy into the memory layer instead of wrapping a policy around it?
- When a customer says "forget me" or "stop marketing to me", what has to happen, and where?
This chapter is not legal advice. It sets out principles that have been stable for years. Jurisdiction-specific rules and dates change often; they live in the dated companion note, Regulation as of September 2026. Check it, and check with counsel, before you ship.
The short answer
Personalization runs on two tests, and most teams only run one. The first is accuracy: is the fact true? The second is appropriateness: would this person expect us to know it, did it come to us for this purpose, and does using it here serve them? A message can pass the first and fail the second, and when it does, the customer does not experience a clever system. They experience surveillance.
The law encodes a floor under the second test. Across the EU, the UK, California, a growing list of US states and Canada, the same ideas recur: collect for a stated purpose and use for that purpose; keep only what you need; treat inferences about health, sexuality, religion, politics and ethnicity as the sensitive data they reveal; let people stop direct marketing, absolutely; and put a human and an explanation behind automated decisions that seriously affect someone. Regulators also treat claims like "anonymized" and "we don't sell data" as promises they will check.
For a leader, three moves matter more than any single statute. Write a never-store and never-infer list and make the system enforce it. Build privacy into the write path of customer memory, because nothing downstream can fully undo what was stored or embedded. And make deletion and objection reach every representation of a person, including the tools you synced them to. If you cannot prove a deletion, you have not done one.
Compliance is the floor. The customer's expectation is the standard they will actually hold you to.
One set of facts, two messages
Larkspur Systems (a fictional company, a composite of many real ones) sells field-service and fleet software. Two messages go to the same fleet operations manager in the same week. Both rest on the same two facts: idle time at her Denver depot is up this month, and she has looked at the pricing page for Larkspur's idle-reduction add-on.
The first appears inside the product, on the dashboard she opens every morning, signed by the account manager she talks to monthly: "Idle hours at Denver are up this month. You were looking at the idle-reduction add-on; here is what it would have flagged over the last thirty days, using your own data."
The second arrives at 11:40 p.m., at a personal email address the enrichment vendor supplied, from a sales rep she has never heard of: "Saw you checking our pricing late last night. Looks like Denver is having a rough month."
Same facts, same company. One reads as service; the other reads as being watched. Nothing in the second message is false. What changed is expectation (she knows her telematics feed the dashboard; she did not expect a stranger to narrate her browsing), purpose (she shared depot data to run her fleet, not to be prospected), channel (her personal inbox, near midnight), and who is speaking (a relationship, versus a system that has clearly been watching).
Figure 19.1. Nothing in the second message is false; expectation, purpose, channel and speaker are what turn service into surveillance.
Aguirre and colleagues tested this with personalized ads and found that when data collection was overt, personalization raised click-through; when it was covert, the same personalization raised feelings of vulnerability and lowered click-through [1]. People do not object to being known. They object to discovering they were watched.
Accurate is a property of the fact; appropriate is a property of the use
Here is the sentence this chapter turns on: accuracy belongs to the fact, appropriateness belongs to the use, and your system has to check both. A true fact can be the wrong thing to say, to this person, in this channel, today.
Stacks that check anything check accuracy: freshness, provenance, confidence (Chapters 11 and 20). Appropriateness is left to the copywriter's taste, which does not scale to a quarter of a million messages. Make it a gate, with four questions a decision layer can evaluate:
- Expectation. Would this person expect us to know this? Declared data and data from the product they use with us pass easily. Browsing, enrichment, and anything inferred need a reason.
- Purpose. Did the fact come to us for a purpose this use serves? Depot telemetry shared to run a fleet serves an operational insight. It does not obviously serve a cold prospecting email.
- Benefit. Does using it here help the person, in this moment, more than it helps us? If the only beneficiary is the conversion rate, treat that as a warning.
- Say-how-we-know. If they asked "how do you know that?", would the honest answer make the message better or worse? If worse, do not use the fact. Use it in a way where the source is obvious, or fall back to something honestly general.
The fourth question is the one teams skip, and the most useful: it turns a vague feeling into a test anyone can run. The dashboard message passes it; the midnight email fails it.
Figure 19.2. Accuracy is checked on the fact; appropriateness is a four-question gate on the use, and a fact must pass both.
A caution on one historical fix. In Charles Duhigg's 2012 account of Target's pregnancy-prediction score, the company reportedly learned that customers reacted badly to obviously targeted baby coupons, and responded by mixing them among unrelated offers so the targeting looked random [2]. (The widely retold story of a father discovering his daughter's pregnancy is single-sourced; I do not rely on it.) Hiding the inference is not the fix. It fails the say-how-we-know test by design. If you have to disguise how you know something, you should not be using it.
The law draws a few lines, and they are stable
Statutes change; the principles under them have been steady. I state them in GDPR terms, the most developed text, and note where other regimes converge.
Purpose, minimization, storage limits. GDPR Article 5 requires that personal data be collected for specified purposes and not used in ways incompatible with them, that it be limited to what is necessary, and that it be kept no longer than needed [3]. Article 25 requires data protection by design and by default [3]. For a memory system: every fact should know its source and permitted uses, and some should expire.
A lawful basis, and an absolute right to stop marketing. GDPR Article 6 lists the lawful bases; Recital 47 says direct marketing "may be regarded" as a legitimate interest, which is permission to weigh, not a blank check [3]. Article 21 is sharper: a person can object to processing for direct marketing, including profiling for it, and then the data "shall no longer be processed for such purposes" [3]. There is no balancing test. Build objection as a hard suppression, not a preference score.
Inferences about sensitive traits are sensitive data. Article 9 restricts processing of data revealing health, sex life or sexual orientation, religious or political beliefs, racial or ethnic origin, trade union membership, and genetic and biometric data [3]. In the OT case (C-184/20, 2022), the EU Court of Justice held that data "liable indirectly to reveal" such traits falls under that protection [4]. That is the legal anchor for the rule practitioners need most: if your model infers it, you have processed it. US regulators arrive at a similar place through enforcement: the FTC's 2023 GoodRx order banned the company from sharing health data for advertising [5].
Automated decisions that matter need a human and an explanation. GDPR Article 22 restricts decisions based solely on automated processing that produce legal or similarly significant effects [3]. The CJEU has read it broadly: in SCHUFA (C-634/21, 2023) a credit score computed upstream counted as the decision when the lender relied on it decisively [6], and in Dun & Bradstreet Austria (C-203/22, 2025) "meaningful information about the logic involved" meant an intelligible account of how the person's data led to the result, not source code [7]. Most marketing personalization is not an Article 22 decision. Pricing, credit terms, eligibility, and account restrictions can be. California's automated decision-making rules, the UK's rewritten Articles 22A to 22D, and Colorado's replacement statute all regulate the same zone, with different scopes and dates (see the note).
Personalization is lawful; manipulation is not. The EU AI Act's Article 5 prohibits AI that uses subliminal, purposefully manipulative or deceptive techniques, or exploits vulnerabilities due to age, disability or social or economic situation, to materially distort behavior and cause significant harm [8]. The Commission's February 2025 guidelines say that advertising which personalizes content based on user preferences is "not inherently manipulative" unless it uses those techniques [8]. That is the legal line between the subject of this handbook and the thing it refuses to teach.
Claims about privacy are themselves regulated. The FTC's position is that hashing an identifier does not make data anonymous, because a hash is still a persistent identifier [9]. In the Rite Aid order (2023), the FTC held a company accountable for the error rate of its AI system and required deletion of data and models derived from it [10]. Say only what your architecture guarantees.
Inference is now where most sensitive data comes from
The privacy conversation still pictures sensitive data as something a customer types into a form. In AI personalization it mostly arrives by inference, and the capability is not new; its price is. In 2013, Kosinski, Stillwell and Graepel used the Facebook Likes of more than 58,000 volunteers to predict private attributes; they distinguished gay from straight men 88% of the time and Democrats from Republicans 85% of the time [11]. Fewer than 5% of the gay users had liked anything explicitly about it. In 2023 and 2024, Staab and colleagues showed that general-purpose language models reading ordinary forum posts could infer personal attributes such as location, income and sex with up to 85% top-1 and 95% top-3 accuracy, at roughly 100 times lower cost and 240 times the speed of human labelers, and that common text anonymization did not reliably stop it [12]. Those are best results on labeled benchmarks, not a claim about any individual. Across the decade between them, the headline accuracy figures look similar; what fell was the cost and time of getting them.
For a builder, the consequence is concrete. Your extraction pipeline will infer sensitive traits whether or not anyone asked. Point a language model at a support transcript, ask for "anything relevant to the account," and it will record that the contact is on medical leave, going through a divorce, or observing a religious holiday next week. All of it is now in memory, retrievable by every agent that reads the record.
So enforce the never-infer list at extraction, not at generation. Say what not to capture in the property schema itself (Chapter 8's extract-with-a-contract), then check the output, because instructions reduce the problem and do not remove it. Better still, drop forbidden properties from the extraction spec before the model sees it. Store the operational fact, not the sensitive one: "unavailable until May 12; route to deputy" serves the account; "recovering from surgery" does not.
Inference is also often wrong, and fluency hides it. A wrong sensitive inference is a double failure: accuracy and privacy at once.
Privacy belongs in the write path
The most important architectural decision in this chapter is where privacy runs. My answer, after getting it wrong once in the memory system I build, is: on the way in, before anything is embedded or extracted [13].
The reasoning is mechanical. In a memory layer, content goes to a language model for extraction and to an embedding model for retrieval. Once an identifier is in a vector, filtering text at read time hides it from the screen; it does not remove its influence on similarity search. Retrieval-time filtering is cosmetic. Only write-time controls are guarantees.
Figure 19.3. Privacy runs on the way in: redact, extract under a never-infer contract, check again, then store; only write-time controls are guarantees.
The mechanics are in Chapter 8 ("Redact on the way in"): secrets, financial and identity numbers removed by default; contact data configurable per deployment (a consistent hash is pseudonymous, not anonymous); two phases, before the model sees the content and again on what it extracted [13][14]. In the system I built, an operator can switch redaction off globally, per tier or per save, so keeping it on is a deployment rule, not a product guarantee. Three properties matter to a reviewer. Detection is a swappable layer you can audit, extend or replace; enforcement (which calls are intercepted, the restore step, the audit record) stays in the core. On a detection error, the system falls back to minimal redaction, or can be set to refuse the call; a pass-raw mode exists and should stay off [19]. And each external model call can log the entity types and counts removed plus a hash of the text sent, never the text, so "what did we send outside last month?" is a query.
Pattern-based redaction is precise on identifiers and blind to meaning; it will not catch a diagnosis described in plain words. Redaction handles identifiers; the extraction schema handles meaning.
Scoping is the other half, and it runs on two axes. Every memory belongs to one entity, and retrieval filters by entity key before similarity search runs; Chapter 13's leakage test shows why [15][18]. The other axis is the reader. In the self-hosted memory system I built, each person and agent can hold its own key, bound to a role and a namespace, enforced by row-level security inside Postgres and failing closed; a read-only agent is not even offered write tools. Test the policy with every key type: one release let a namespace-scoped admin key read every namespace, because the role check ran before the namespace gate. Isolation must be enforced by storage, never by semantics.
The same holds on the way out: the governed personalization engine I built passes customer-facing payloads through a serializer that allow-lists fields, so raw research or seller-only notes cannot leak just because someone added an internal property.
Every place content leaves is a privacy decision
When I mapped egress for that self-hosted system, it was a closed list: up to three model calls on a save (extraction, embedding, a document rewrite), up to three on a retrieve (query embedding, intent classification, answer synthesis), and two for hosted agent jobs. Nothing is stored off the box, but protection is not uniform, and a builder has to know where it thins [19]:
- The embedder is its own channel. Stand-ins would not match between a save and a later query, so its inputs can only be redacted: names and emails still reach it. For regulated data, embed locally.
- Query intent classification is redact-only, since tokenizing would corrupt the search terms. Route it to a local model for full query privacy.
- Hosted agent jobs are covered only where the system is in the path. The job's text fields and tool results are tokenized, and echoed stand-ins are restored before any write, so a save never keys on a fake identifier; the rest of the provider-native request passes through as written.
- Two paths are not covered: a locally run curation pass sends its tool calls to the configured provider unprotected, and an interactive assistant connected to memory receives real values.
Tokenization, reversibly swapping identifiers for stand-ins, is an egress control, not an at-rest one: the provider never sees the real values, but they are restored and stored, so the data is still personal data (a partner's diligence review caught that our setting's name implied otherwise).
The choices form a ladder, most private first: fully local models, air-gapped if needed (no content leaves, at a cost in model size and speed); a cloud model with tokenization; a cloud model with redaction only (names and emails leave); raw content, acceptable only when the provider sits inside your own compliance boundary. On the cloud rungs, provider order matters; for regulated deployments I rank: a model in your own cloud account, reached through the cloud's workload identity so there is no provider key to leak; then a provider direct under a signed data processing agreement with no training and short or zero retention; never an aggregator router. The processing agreement, and a health-data business associate agreement where eligible, is generally part of the provider's terms at no extra cost. Verify current attestations and get retention terms in writing. None of this is legal advice, and none of it makes a deployment compliant by itself.
Deletion is a test of your architecture, not your intentions
"Forget me" sounds like one operation. In an AI memory system it is many: typed properties, extracted memories, vectors, a long-form synthesis, cached context, drafts in a queue, decision traces, logs, backups, and copies synced to the CRM, the email platform and the ad audience. Delete the row and leave the synthesis, and the next agent still knows everything.
Three rules make deletion real:
- Keep a deletion manifest per entity. Every representation written for a person registers itself, so erasure has a list to walk instead of a guess. Propagate to downstream systems and record their confirmation.
- Regenerate what was derived. Summaries and syntheses that mentioned the person, or were built from deleted facts, get rebuilt, not just edited. If your design fine-tunes on customer data, know that the EDPB's view is that a model trained on personal data cannot always be considered anonymous [16], and the FTC has ordered deletion of models along with data [10]. The simplest defense is not training on customer records at all.
- Suppression survives deletion. Objection to marketing must outlive the record, or the next enrichment import will bring the person back and the next campaign will email them. Keep a minimal suppression entry (for example a keyed hash of the address, flagged for this sole purpose) that every send path checks.
In the memory system I built, erasure is never metered or license-gated, also scrubs background consolidation summaries that cited the person, and is never offered to autonomous agents.
Deletion is where you find out whether "privacy by design" was design or décor.
Channels have their own rules
In outline: US commercial email is opt-out (CAN-SPAM, which covers B2B mail); Canada is opt-in (CASL); the EU requires consent for most electronic marketing to individuals, with a narrow existing-customer exception. In the US, AI-generated voices count as "artificial" for telephone consent rules. Mailbox providers add their own bulk-sender rules: authentication, one-click unsubscribe, low complaint rates. The practical point: consent state and channel rules are inputs to the decision layer (Chapter 14), checked per person and per send. Thresholds, penalties and dates are in the companion note.
What this does not do
It does not make you compliant; sector rules (health, finance, children) and contracts add obligations not covered here. It does not fully solve detection: regex finds identifiers, not meaning; name protection built on a supplied list misses unlisted names in free text, so pattern tools alone fall short of US Safe Harbor de-identification, which removes all 18 listed identifier types, free-text names and dates included; model-based detection adds its own errors. And it does not remove a real tradeoff. Less data sometimes means a less specific message. When the only specific thing you could say fails the appropriateness gate, the right output is honestly general. That costs some lift. It buys the thing lift depends on.
And two popular claims deserve pushback. The first is that consumers "expect personalization," usually supported by McKinsey's 2021 finding that 71% of consumers expect it [17]. That is a survey answer about wanting relevance, not consent to any use of any data; the same research base shows people punishing covert collection [1]. The second is the "privacy paradox" argument that because people share data despite saying they care, they do not really care. Aguirre's result points the other way: people share readily when collection is visible, and pull back when it is hidden [1]. The paradox is mostly a transparency problem.
At scale
At a quarter of a million people, rare harms stop being rare: a sensitive inference that slips through once in ten thousand messages is dozens of people a month. The appropriateness gate has to run in code, on every decision, with its reasons in the decision trace (Chapter 18).
Scale multiplies jurisdictions, so policy is resolved per person (residence, consent state, channel) at decision time, not assumed per campaign. And scale multiplies copies: every new agent, cache and integration is one more place a deleted person can survive. Treat "can we find every copy of this person within a day?" as an operational metric, not an audit finding.
Failure story: The Get-Well Email
A Larkspur support call runs long. The customer's fleet lead mentions, in passing, that she will be out for six weeks after surgery and that her deputy will handle the route sync issue. The extractor, asked for "anything relevant to the relationship," stores a memory: contact recovering from surgery, returns mid-May. Every fact is accurate.
Two weeks later the renewal agent, grounding its draft in memory exactly as designed, writes to her: "Wishing you a smooth recovery. When you're back in May, let's talk about expanding to your Phoenix depot." It passes every accuracy check. She forwards it to her CEO with one line: "Why does our software vendor know about my surgery?"
Nothing was hallucinated or hacked. The system stored health information nobody needed, then used it in a sales message. The fix had three layers: an extraction contract that captures "unavailable until May 15; route to deputy" and never the reason; a never-infer check on extracted output; and an appropriateness gate that blocks health, family and personal-life facts from generated sales copy.
The anti-pattern is The Accurate Intrusion: a true fact, used where the person never expected it to travel.
Patterns
Write-Path Privacy. Problem: sensitive data reaches models, vectors and memory before any filter runs. Forces: extraction needs context; some identifiers are needed for linkage. Solution: tiered, two-phase redaction before extraction and embedding; always-on tiers for secrets, financial and identity numbers; configurable contact tier; audit record on provenance. Tradeoffs: false positives on numeric strings; misses meaning-level sensitivity, so pair with extraction contracts.
Purpose-Tagged Memory. Problem: facts collected for one purpose leak into another. Forces: agents read everything they can retrieve. Solution: every property and memory carries source, lawful basis or consent state, allowed purposes and expiry; the decision layer filters by purpose before a fact reaches a prompt; the appropriateness gate runs on the draft. Tradeoffs: more metadata per fact, and a schema discipline most teams have to learn.
Deletion Manifest. Problem: erasure misses derived and downstream copies. Forces: many representations, many systems. Solution: register every representation per entity at write time; erasure walks the manifest, regenerates derived syntheses, propagates downstream, records confirmations; keep a minimal suppression entry. Tradeoffs: write-path overhead, and discipline for every new integration.
Leader questions
- What is on our never-store and never-infer lists, and where in the system are they enforced?
- If a customer asked "how do you know that?" about our last campaign, would every answer make us look better?
- How long does a full erasure take, and how do we prove it reached every copy, including vendors?
- Which of our automated decisions could significantly affect someone (price, terms, eligibility), and who reviews them?
- What do our privacy claims to customers promise, and does the architecture guarantee each one?
Build checklist
- Redaction runs before extraction and embedding; secrets, financial and identity numbers are always removed.
- Post-extraction scan for reconstructed identifiers.
- Extraction schemas state what never to capture (special categories, personal life) and store operational facts instead.
- Every fact carries source, consent or lawful basis, allowed purposes and expiry.
- Every model and embedding call that can send content out is listed with its protection; the embedder counts as its own channel.
- Retrieval is scoped by entity key before similarity search; per-person and per-agent keys are enforced in the database; customer-facing payloads pass an allow-list serializer.
- An appropriateness gate (expectation, purpose, benefit, say-how-we-know) runs on generated drafts, with reasons in the decision trace.
- Objection to marketing is a hard suppression checked by every send path and survives deletion.
- A per-entity deletion manifest covers properties, memories, vectors, syntheses, caches, queues, logs, document history and downstream systems.
- No training or fine-tuning on customer records unless deletion from the model is solved.
- Consent state and channel rules are inputs to each send decision, resolved per person.
Metrics to watch
- Sensitive-inference rate: share of extracted memories flagged by the never-infer check (should trend to near zero, then stay there).
- Appropriateness-gate blocks: count and reasons per thousand drafts; a sudden rise usually means a new data source.
- Time to complete erasure, end to end, including downstream confirmations.
- Suppression leaks: messages sent to people who objected (target: zero; every one is an incident).
- "How did you know" complaints: replies and tickets questioning how you knew something, tagged and reviewed monthly.
Reader Q&A
Is B2B data covered by privacy law? Yes, where it identifies a person. A named contact's work email, title and behavior are personal data under GDPR and in scope for CASL and CAN-SPAM. Company-level facts about an organization generally are not.
If the data is public, can we use it? Public is not the same as expected. Public data still needs a lawful basis in the EU, and inferences drawn from it can be sensitive (the OT principle). Run the appropriateness gate regardless.
Is AI personalization "manipulation" under the EU AI Act? Not by itself [8]. Any system precise enough to serve a person is precise enough to pressure them; the law, and this handbook, draw the line at the second.
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 19
concepts:
- name: Appropriateness Check
definition: "Four-question gate on any personal fact before use: expectation, purpose, benefit, say-how-we-know. Separate from accuracy."
- name: Never-Store List
definition: "Data the memory layer removes on ingestion by default, kept always on as a deployment rule because the system allows it to be switched off: secrets, financial account numbers, identity document numbers."
- name: Never-Infer List
definition: "Traits the system must not extract or derive: health, sex life or orientation, religion, politics, ethnicity, union membership, genetic or biometric data, and personal-life circumstances; store operational facts instead."
- name: Write-Path Privacy
definition: "Tiered, two-phase redaction before extraction and embedding, because stored vectors and memories cannot be reliably filtered afterward."
- name: Deletion Manifest
definition: "Per-entity registry of every representation and downstream copy, walked on erasure; derived syntheses are regenerated; suppression survives deletion."
decision_rules:
- if: "a fact is accurate but the person would not expect us to know it, or its source would make the message worse if disclosed"
then: "do not use it; make the source obvious or fall back to an honestly general statement"
- if: "an extracted fact reveals or implies a special-category trait"
then: "discard it and store only the operational consequence (for example availability or routing)"
- if: "a person objects to direct marketing"
then: "hard-suppress across every channel and agent, with no balancing, and keep a minimal suppression entry after deletion"
- if: "an automated output sets price, terms, eligibility or access for a person"
then: "treat it as a possible regulated automated decision: human review path, explanation, contest route; check the companion note for jurisdiction"
- if: "a privacy claim (anonymized, not sold, deleted) is made to customers"
then: "verify the architecture guarantees it; hashed identifiers are not anonymous, and tokenization protects what a model provider sees, not what is stored"
assessment_questions:
- "Where does redaction run relative to extraction and embedding?"
- "Which traits does your extraction schema explicitly forbid capturing?"
- "Does every stored fact carry source, consent or lawful basis, allowed purposes and expiry?"
- "How many systems hold a copy of a customer, and can erasure reach all of them within a set time?"
- "Which jurisdictions do your customers live in, and is policy resolved per person at decision time?"
- "Do any automated outputs affect price, terms or eligibility?"
patterns: [Write-Path Privacy, Purpose-Tagged Memory, Deletion Manifest]
anti_patterns: [The Accurate Intrusion, Retrieval-Time Filtering as Privacy, Hashing as Anonymization, Tokenization Sold as At-Rest Protection, Hide the Inference, Deletion That Leaves the Summary]
maturity_dimension: governance_and_reliability
companion_note: regulation-as-of-2026-09
not_legal_advice: trueReferences
- Aguirre, E., Mahr, D., Grewal, D., de Ruyter, K., Wetzels, M. (2015). "Unraveling the personalization paradox: The effect of information collection and trust-building strategies on online advertisement effectiveness." Journal of Retailing 91(1). https://www.sciencedirect.com/science/article/abs/pii/S0022435914000669
- Duhigg, C. (2012-02-16). "How Companies Learn Your Secrets." The New York Times Magazine. https://www.nytimes.com/2012/02/19/magazine/shopping-habits.html
- Regulation (EU) 2016/679 (GDPR), Arts. 5, 6, 9, 17, 21, 22, 25; Recital 47. https://eur-lex.europa.eu/eli/reg/2016/679/oj
- CJEU, Case C-184/20, OT v Vyriausioji tarnybinės etikos komisija, judgment of 2022-08-01. https://curia.europa.eu/juris/liste.jsf?num=C-184/20
- FTC (2023-02-01). "FTC Enforcement Action to Bar GoodRx from Sharing Consumers' Sensitive Health Info for Advertising." https://www.ftc.gov/news-events/news/press-releases/2023/02/ftc-enforcement-action-bar-goodrx-sharing-consumers-sensitive-health-info-advertising
- CJEU, Case C-634/21, SCHUFA Holding (Scoring), judgment of 2023-12-07. https://curia.europa.eu/juris/liste.jsf?num=C-634/21
- CJEU, Case C-203/22, Dun & Bradstreet Austria, judgment of 2025-02-27 (press release 22/25). https://curia.europa.eu/site/upload/docs/application/pdf/2025-02/cp250022en.pdf
- Regulation (EU) 2024/1689 (AI Act), Art. 5; European Commission Guidelines on prohibited AI practices, 2025-02-04. https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5 ; summary: https://www.insideprivacy.com/artificial-intelligence/european-commission-guidelines-on-prohibited-ai-practices-under-the-eu-artificial-intelligence-act/
- FTC Office of Technology (2024-07-24). "No, hashing still doesn't make your data anonymous." https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/07/no-hashing-still-doesnt-make-your-data-anonymous
- FTC (2023-12-19). "Rite Aid Banned from Using AI Facial Recognition After FTC Says Retailer Deployed Technology without Reasonable Safeguards." https://www.ftc.gov/news-events/news/press-releases/2023/12/rite-aid-banned-using-ai-facial-recognition-after-ftc-says-retailer-deployed-technology-without
- Kosinski, M., Stillwell, D., Graepel, T. (2013). "Private traits and attributes are predictable from digital records of human behavior." PNAS 110(15). https://www.pnas.org/doi/10.1073/pnas.1218772110
- Staab, R., Vero, M., Balunović, M., Vechev, M. (2024). "Beyond Memorization: Violating Privacy Via Inference with Large Language Models." ICLR 2024. https://arxiv.org/abs/2310.07298
- Taheri, H. (2026-03-28). "4-Tier PII Redaction: How We Built Privacy Into the Memory Layer, Not Around It." https://hamedtaheri.com/articles/four-tier-pii-redaction
- Taheri, H. (2026-03-14). "Two-Phase Redaction: Scrubbing PII Before and After LLM Extraction." https://hamedtaheri.com/articles/two-phase-redaction-pii
- Taheri, H. (2026-03-14). "Zero Cross-Entity Leakage Across 3,800 Results." https://hamedtaheri.com/articles/zero-cross-entity-leakage
- EDPB (2024-12-17). Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models. https://www.edpb.europa.eu/documents/opinion-of-the-board-art-64/opinion-282024-on-certain-data-protection-aspects-related-to_en
- McKinsey & Company (2021-11-12). "The value of getting personalization right, or wrong, is multiplying." https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying
- Taheri, H. (2026-03-18). "Governed Memory: A Production Architecture for Multi-Agent Workflows." arXiv:2603.17787. https://arxiv.org/abs/2603.17787
- First-party: operator documentation for a self-hosted governed memory system built by the author (data flow and PII handling, access control, hosted jobs, partner brief), release 0.8.2, verified 2026-09-16.
This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised