Questions this chapter answers
- Who decides what our AI may say and do, and where does that decision actually live?
- How do company-wide rules and one customer's wishes both reach the agent at the moment it acts?
- If a customer complains about something our AI did, can we reconstruct what it knew, what it decided, and who allowed it?
- Which actions need a human, and how do we keep human review from becoming a rubber stamp?
- Do we need NIST, ISO 42001, or the EU AI Act to get started?
The short answer
Governance for AI personalization is not a policy document, a review board, or a dashboard. None of those can stop anything, because none of them sits in the path the agent takes when it acts. Governance is a code path: rules are executed inside the system, before the model reasons, around the tools it can call, and after it writes, so a rule holds whether or not the model remembers it.
It works at two levels. At the organization level, it carries the company's rules (brand, claims, pricing authority, compliance) with explicit precedence, so a campaign instruction can never override a legal constraint, and with versions, so every rule change is a release with an owner and a rollback. At the customer level, it carries each person's consent, preferences, contact rules, the promises made to them, and sensitive context, and treats those as hard constraints on every agent that touches that person.
Three disciplines make it real. Traceability: every consequential action leaves a decision trace of what was known, which rule versions applied, what was decided, and who allowed it. Risk tiers: put real gates only in front of the short list of actions that cannot be taken back (sending to a customer, moving money, making a commitment, changing a system of record, letting data leave). Usable human review: people approve the few decisions that deserve it, with the evidence in front of them, in a flow where saying "no" is fast and saying "yes" without looking is hard.
If you fund one thing, fund the irreversible-action list and the gates around it. It fits on an index card, and it is where a mistake stops costing time and starts costing the thing itself.
In November 2022, Jake Moffatt's grandmother died. He went to Air Canada's website to book a flight and asked the airline's chatbot about bereavement fares. The chatbot told him he could buy a regular ticket and apply for the bereavement discount afterward. That was not the policy. A page elsewhere on the same website said bereavement fares could not be claimed retroactively, and when Moffatt applied, Air Canada refused.
He took the airline to British Columbia's Civil Resolution Tribunal, and in February 2024 he won [1]. The damages were small: roughly the fare difference, plus interest and fees. The airline's argument is what makes the case worth remembering. In the tribunal's words, "Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission." It "makes no difference whether the information comes from a static page or a chatbot," and the airline had not explained why customers "should have to double-check information found in one part of its website on another part" [1][2].
Notice what failed. Air Canada had the correct policy. It was written down, published, and owned by someone. It simply never reached the component talking to the customer at the moment it mattered, and nothing checked the answer before it left. The company had a policy. It did not have governance.
Governance is a code path, not a document
Here is the sentence this chapter turns on: if you cannot name the component that enforces a rule, and say what happens when that component says no, you do not have a rule. You have a preference.
Most AI governance programs produce a policy, a committee, and a dashboard, and all three stand outside the execution path. By the time any of them has an opinion, the email is in the inbox.
The instinctive fix is to paste the rules into the system prompt. That holds for about a week. Prompts are unversioned, invisible to the teams who own the rules, and multiplied: ten agents across five teams become ten private interpretations of the company. When legal updates a policy, the change reaches the agents legal knows about and misses the rest. And a rule inside a prompt is a request, not a boundary. The model will usually honor it. "Usually" is a fine standard for tone and a poor one for a refund limit.
So split every rule into two kinds. Rules code can check (maximum discount, required unsubscribe link, no internal notes in customer output, no contact during a hold, no numeric claim without evidence) become deterministic checks. Rules only judgment can apply (tone, what counts as an overclaim, whether a message fits the customer's situation) go into the model's context as governed guidance and are then checked by semantic review and sampling. If a rule can be deterministic, it should not depend on the model remembering to obey it. On adaptive pages and screens the most important such rule is structural: exits, prices, consent, and cancel paths sit in a fixed frame no personalized slot can change (Chapter 3), and every generated value passes a render gate before anyone sees it (Chapter 16).
That split gives the first pattern, the Policy Sandwich: deterministic before, probabilistic in the middle, deterministic after. Before the model runs, code decides eligibility (may we contact this person, through this channel, about this?) and assembles the governed context. The model reasons and writes. Afterward, code validates the output and gates the action. The probabilistic layer is where the value is. The deterministic layers are where the guarantees are.
Add tools and you get three-layer guardrails: before (eligibility, consent, the right rules routed into context), during (the capability boundary: an agent can call only the tools, with only the parameters, its role allows), and after (output validation and an action gate in proportion to consequence). The middle layer is the one teams skip, and the one that holds when everything else fails. An agent that drafts retention offers should hold no credential that can issue a refund. Its worst case is the sum of its permissions, not the quality of its instructions.
In a self-hosted memory system I built, each agent's key carries a role, and the agent sees only the tools that role may call: a reader agent's write tools are not refused, they are absent. Autonomous loops are never offered hard delete. A job handed to an outside provider gets a credential minted for that job alone, limited to three memory tools and revoked when the job ends.
Figure 18.1. Deterministic before and after, probabilistic in the middle, with a capability boundary around the tools.
Organization level: precedence, owners, versions
A legal constraint, a brand preference, a campaign strategy, a channel limit, and an instruction for one paragraph are different kinds of statement. Mixed into one prompt, they compete on equal terms, and the model resolves the conflict however it resolves it that day. I treat them as a governance hierarchy, where a lower level can never override a higher one:
- Hard constraints (legal, safety, privacy; enforced in code where possible)
- Enterprise policy (claims, pricing authority, brand)
- Program or campaign contract
- Channel policy (length, format, required elements)
- Task instruction
- Customer context
- Model preference (what the model would do if left alone)
Figure 18.2. Rules of different authority do not compete on equal terms; precedence decides, and hard constraints always win.
Suppose the enterprise rule is that internal CRM notes never reach a customer, and a campaign manager writes "use every available insight to make this specific." The campaign must lose, and it should lose in code: the one serializer every customer-facing payload passes through strips internal fields whatever any instruction said. That is the general move at scale: find the last shared function before an important boundary and make it protective (the serializer, the sender that enforces suppression, the parser that rejects unsafe links), so every new program inherits the protection instead of relying on each author to remember it.
The most useful enterprise rules are allow-lists, not prohibitions. In a governed personalization engine I built, each campaign carries a capability map of what the AI may recommend, so the model is asked not "what would sound good here?" but "which approved capability fits this buyer?" Selected engines limit numeric proof to customer-approved statistics. A customer-status rule forbids "your renewal is coming up" unless that relationship is known; otherwise the copy is reframed as an evaluation. Air Canada's chatbot had no such list: it stated a policy nobody had approved.
Three practices keep the hierarchy honest.
Every rule has an owner. Marketing owns voice, sales operations owns discount authority, legal owns claims and consent language, support owns service promises. The governance layer does not make policy. It makes sure every agent, on every platform, receives the current approved version of each owner's policy by reference, and nothing else. In the layer I built, hard constraints are marked always-on and critical, so they are never trimmed when the context budget is tight, and authoring quality turned out to decide discoverability: in our internal tests, well-written rules were 20 to 50 percentage points more likely to reach the right task than poorly written versions of the same rules [3], and the architecture paper reports 92 percent governance routing precision in controlled experiments [16]. A rule titled "Q3 update" is governance in name only. Not every rule is for every agent, either: in the self-hosted memory system I built, a guideline stays org-wide by default but can be scoped to its author, a group, one person, or whoever can already see a given record.
Every rule has a status and a version. Draft, approved, deprecated; agents see only approved. A rule change is a code change in everything but syntax, so its record carries the previous and new version, the reason, the evidence, the approver, the validation results, and the rollback target. Then you can say "unsupported claims fell after v8" instead of "we tuned the prompt last week." Reference each program's rules by a stable identifier, never a loose name: "Q3 enterprise" and "Q3 enterprise EMEA" are different contracts. And make the rollback real. My memory system once let a document update silently replace the prior body; now the old version is snapshotted in the same transaction as the change, and a restore can itself be undone.
Model changes are releases too. In January 2024 the parcel company DPD disabled the AI element of its customer chatbot after it swore at a customer and wrote a poem criticizing the company. DPD's explanation: "An error occurred after a system update" [4]. A new model, prompt, or guideline that reaches customers deserves the same regression checks as a code deploy.
Customer level: every person carries constraints
Organization-level rules say what the company will do. Customer-level governance says what the company will do to this person. It is the half most programs forget, because it looks like data rather than policy. Each record should carry, as typed fields with provenance rather than free-text notes:
- Consent and lawful basis per purpose and channel, with source and date.
- Contact rules: opt-outs, frequency caps, quiet periods, preferred channel, relationship owner.
- Stated preferences: "email only," "not before our fiscal close," "talk to my deputy."
- Promises made: a price hold, a service credit, "no outreach until the review is done." A promise is a constraint on every future action, not a line in a ticket.
- Sensitive context: an open escalation, a dispute, a bereavement, a legal hold. These often change whether to act at all, not only the tone.
- Suppression: an explicit do-not-contact or do-not-personalize state that overrides every program.
The precedence rule is simple: a customer-level constraint outranks a campaign, and it is checked in the "before" layer by code, so the agent is never invited to write the message. This is where governance meets memory (Chapters 10 and 11). A promise that lives as prose in a support ticket will not be found by the agent that needs it, or will be weighed as one fact among many. The Boundary Rule applies with force: if you would put it in a WHERE clause, it is a property. "Do not contact before March 1" is a WHERE clause.
The same layer should hold what the system declines to use: facts that are accurate and lawful but wrong to act on (Chapter 19). Appropriateness rules need a governed home, or they exist only in someone's judgment.
Traceability: can you answer in thirty seconds?
When a customer, a regulator, or your CEO asks "why did the AI say that?", there are two possible answers: a week of log archaeology, or a lookup that takes thirty seconds because every consequential action wrote a decision trace when it happened. A useful trace records:
- What was known: the customer facts used, each with source, time, and confidence, and the customer-level constraints checked.
- What governed it: rule versions, program contract, channel policy, model and prompt versions.
- What was decided and why: the candidate actions considered (including doing nothing), the one chosen, the stated rationale, every validator result.
- Who allowed it and what happened: automatic or approved, by whom and when; what was actually sent or changed; what the customer did next.
Three cautions. A model's explanation of its own choice can be plausible and wrong, so the trace must hold the inputs and checks that make the rationale verifiable. The trace is itself customer data, with the same access, retention, and deletion obligations as the memory it points to (Chapter 19). And the trace is evidence only if the model cannot write it: in the memory system I built, the system stamps who made each automated change, every agent tool call that changes something writes an audit row with the caller and the shape of its arguments (never the content), and audit rows can be cryptographically chained so an edited history is detectable.
The trace turns incidents into lookups and lets you correlate quality with versions (Chapter 20). Outside parties increasingly expect it. As of September 2026, the EU AI Act's Article 50 transparency duties apply from August 2, 2026, including telling people they are interacting with an AI system unless that is obvious [5]. A disclosure is a string your agent emits; a record of what it did and who authorized it is a log line with an actor, an action, and a timestamp. Both are cheap to design in and expensive to retrofit, exactly like authentication. (Timing for other obligations moved with the 2026 Digital Omnibus [6]; keep dates in the companion note. Not legal advice.)
Risk: sort actions by whether they can be undone
Heavy review everywhere is unaffordable, so uniform oversight quietly becomes light review everywhere, including on the actions that cannot be undone. Governance with stakes makes the opposite trade.
| Tier | What it covers | Default control |
|---|---|---|
| 0. Internal, reversible | drafts, notes, summaries, proposed record changes | automatic; logged |
| 1. Customer-visible, low stakes | website modules, in-product recommendations, anything replaced on next render | automatic with validation and sampled review |
| 2. Customer communication | messages that make claims or cite facts about the person | validation plus semantic review; humans review pilots and flagged cases |
| 3. Irreversible or binding | the list below | hard gate: limits in code, human approval above thresholds |
The irreversible-action list for an AI acting on customers has five groups:
- Send: a delivered message cannot be unread.
- Money: refunds, credits, discounts, price changes, write-offs.
- Commitments: statements of policy, terms, eligibility, or promises. Air Canada's chatbot made one its owner then had to honor [1]. In April 2025 the support bot of the coding tool Cursor told users a "one device per subscription" policy existed; it did not, users cancelled, and a co-founder apologized for an "incorrect response from a front-line AI support bot" [7][8]. In December 2023 a car dealership's chatbot agreed to sell a new Tahoe for one dollar and called it "a legally binding offer" [9]. Different causes, one lesson: an agent that can state terms needs a boundary on which terms it may state.
- Records: deleting, merging, or overwriting a system-of-record field; changing consent or opt-out state.
- Data leaving: exports, partner sharing, one customer's data surfacing in another's context.
That list fits on an index card. Everything else is recoverable by definition.
Figure 18.3. Sort actions by whether they can be undone, and spend real gates only on the five that cannot.
The control hierarchy: approval is the fourth control, not the first
For each item on the list, apply these in order:
- Remove the authority. The cheapest control is the capability that does not exist.
- Make it reversible. Irreversibility is often a default setting, not a law of nature. Delay the send by fifteen minutes; stage the record change as a proposal; soft-delete with a restore window; cap and expire the offer.
- Bound it in code. Thresholds, allow-lists of claims and offers, per-customer rate limits, the customer-level constraints above.
- Require approval where consequence justifies it: above a discount threshold, for a new claim class, for a first send in a new program.
- Monitor and test. Log every gate decision, override, and bypass; replay past cases against new rules before release.
- Support the reviewers. Rehearse the approval path and reward the person who says "not yet."
The order prevents a common mistake: asking a person to compensate for authority the organization could have removed upstream. A human approving two hundred refunds an hour is not a control. That human is a latency.
The fix for volume is to approve the judge, not each verdict. The memory system merges duplicate customer records, a records action. Code finds candidate pairs, the model judges each, and every merge is journaled and undoable for ninety days by default. There is deliberately no per-merge approval: the routine starts by only proposing, and an operator promotes it one stage at a time on reviewed runs. Demotion is always allowed and takes effect on the next write. The stage is a property of the call path, not an instruction in the prompt, so a model cannot talk its way past it (Chapter 11). Close the side doors too. In the same system, every safety control on a background curation routine (stage, write ceiling, change ledger) hangs off its recipe, and for a while a generic scheduling path could start one with no recipe, and so with none of those controls. That path now refuses it. A control that one path can skip is a control on the main path only.
The same logic should govern changes to the system itself. In the personalization engine, that loop is operator-run today: a person uses AI to reason across the first real cohort, classifies each finding by root cause, and decides the fix. The design I am productizing is two approvals. A routine proposes a fix for a recurring problem (illustratively: "11 of 83 pages use unsupported relationship language; add one guideline change, one deterministic guard, six regression examples"). Approval one: the diagnosis is reasonable, implement and prove it outside production against replayed, held-out, and regression sets. Approval two: it passed, promote it. What should hold is the credential underneath: the analysis worker gets read-only production access and a branch, and no deployment or data-mutation rights, so nothing it holds can ship anything. That is human-supervised autonomy: autonomy to observe, diagnose, propose, implement outside production, and validate; production authority held by people.
Engaging humans: the hard part is usability, not design
A human can sit in the loop (approves before the action; right for tier 3, and tier 2 during pilots), on the loop (watches behavior in near real time and can pause, override, or roll back; right for mature tier 1 and 2), or over the loop (sets rules, reviews samples and traces, changes the system; right for everything).
The EU AI Act's human-oversight article is a good design template even where it does not legally bind. (As of September 2026, Article 14 applies to systems the Act classes as high-risk; most marketing personalization falls outside that class, though uses such as credit decisions or life and health insurance pricing may fall inside [10].) It asks that overseers can understand the system's limits, "remain aware of the possible tendency of automatically relying or over-relying on the output" (automation bias), interpret its output, "override or reverse the output," and stop it "in a safe state" [10]. Read those as five requirements for a review screen.
Oversight rarely fails because the principle is wrong. It fails because it is too slow, adds friction to routine work, and gets waived for important accounts at quarter end. A review step the organization sets aside whenever it is in a hurry is not a control. It is a policy with a known bypass built in. So design review like any product for busy users:
- Show the evidence, not just the output: the facts the message relies on, with source and date, and which validators flagged.
- Make "no" cheap: one click with a reason code that feeds the learning loop (Chapter 20).
- Put friction only where the consequence is. If people route around review because it is slow for routine work, the tiers are wrong.
- Watch for rubber-stamping. Approval rates at or near 100 percent, or review times too short to read the evidence, mean the humans have become part of the automation. Seed known-bad cases and check they are caught.
- Never let urgency waive the gate. The path for "this must go now" is a faster approver, not no approver.
A control that survives contact with a real deadline is the only kind that counts.
The frameworks are maps, not mechanisms
Your customers' procurement teams will ask about these, so know them.
- NIST AI RMF 1.0 (January 2023) is voluntary and organizes the work into four functions, Govern, Map, Measure, Manage, with Govern cutting across the rest [11]. Its Generative AI Profile, NIST AI 600-1 (July 2024), lists twelve generative-AI risks, including confabulation (NIST's term for fabricated output), data privacy, information integrity, and human-AI configuration [12].
- ISO/IEC 42001:2023 is the first certifiable management-system standard for AI: policies, roles, risk assessment, continual improvement [13].
- The EU AI Act draws legal lines. Article 5 prohibits manipulative or deceptive techniques that cause significant harm; the Commission's guidelines say preference-based personalization is not inherently manipulative unless it uses such techniques [14]. Article 50 sets disclosure duties [5].
My qualification of all three: they tell you what to manage and are silent on where the control sits. An organization can map every NIST subcategory to a policy document and still run an agent that can issue any refund it likes. Aligning or certifying does not, on its own, settle any privacy or AI-law obligation (Chapter 19). Use the frameworks to structure the program; use this chapter's patterns to decide what runs in the path.
What this does not do
Governance lowers the rate and cost of mistakes. It does not eliminate them.
Deterministic checks cover only what they detect. A grounding check catches unsupported revenue figures because someone wrote a detector for them. It will not catch the claim class nobody anticipated.
Tests are snapshots. When I tested our governance layer against fifty designed scenarios meant to push agents past ten kinds of policy (discount limits, confidential figures, unbacked SLA claims), compliance held in all fifty [15][16]. That is a signal about the architecture, not a guarantee about the world. A policy that does not cover a situation offers no protection, and a model update can shift behavior at the margins. Treat any 100 percent as the start of a larger test set.
"A human will approve it" is not automatically a safeguard. Without evidence on the screen and time to read it, it produces automation bias with a signature attached.
Governance can strangle value. A program that routes every tier 1 recommendation through approval will be abandoned within a quarter, and the abandonment will take the tier 3 gates with it. Proportion is the whole discipline.
At scale
Larkspur Systems (a fictional composite company used throughout this handbook) runs agents across its six teams (marketing, sales, customer success, support, product, and events), touching about 40,000 accounts and 250,000 contacts. At that scale, two things change.
Customer constraints must be global. A quiet period set by support must stop marketing's agent, the sales sequence, and the renewal bot, which works only if the constraint lives in shared customer memory, not one team's tool. The most common failure at scale is not a rogue model. It is several well-behaved agents that never heard each other's promises.
Oversight shifts from items to populations. Nobody reviews 250,000 messages. You review the pilot cohort, then distributions: flagged-claim rates, suppression hits, overrides, complaints per thousand sends, by program and rule version. The trace makes this possible; the tiers make it affordable.
Failure story: The Broken Promise
After a painful outage, Larkspur's account manager called a fleet customer's operations director and promised two things: a service credit, and no sales or marketing outreach until the post-incident review closed. She logged both in the support ticket, as a note.
Two weeks later the expansion agent ran its weekly scan. Usage on the account had spiked (the customer was re-running jobs the outage interrupted), and a model scored it as a strong upsell candidate. The agent checked consent: valid. Frequency caps: clear. It wrote an accurate, grounded email about capacity add-ons, passed every validator, and sent it. The operations director forwarded it to the account manager with one line: "So much for your word."
Every fact was true and every component behaved. The agent chose a locally reasonable action that broke a customer promise, because the promise lived in prose in another team's system and never became a constraint. The fix was not a better model. It was a typed customer-level field (outreach hold, reason, owner, expiry), written by the support workflow, read by every agent's "before" layer, outranking every campaign.
Patterns
Policy Sandwich. Problem: rules in a prompt are followed usually, not always. Forces: the model's judgment is the value; some rules cannot tolerate "usually." Solution: deterministic eligibility and context before the model; deterministic validation and gating after. Tradeoffs: rules moved into code must be written and maintained; judgment rules stay probabilistic and need sampling.
Irreversible-Action Gate. Problem: uniform oversight becomes light oversight everywhere. Forces: review is expensive; permanent mistakes cost more. Solution: enumerate send, money, commitments, records, and data leaving; apply the control hierarchy to each. Tradeoffs: latency on gated actions; tiers must stay honest or gates creep into everything.
Governance Hierarchy. Problem: rules of different authority collide in one context. Forces: many owners, many programs, frequent change. Solution: explicit precedence, customer constraints above campaigns, owned and versioned rules served by reference, hard constraints enforced at shared choke points. Tradeoffs: up-front ownership work; conflicts surface and must be resolved by people instead of silently by the model.
Leader questions
- Name the component that enforces our discount limit. What happens when it says no?
- If a customer asked why our AI said something last Tuesday, how long until we can show what it knew, which rule version applied, and who allowed it?
- Where does a promise one team makes to a customer live, and which agents can read it?
- Which irreversible actions can any agent take today without a gate?
- What is the approval rate in our review queues, and how do we know reviewers are reading?
Build checklist
- Write the irreversible-action list for every agent.
- Remove tools and credentials agents do not need; separate read, propose, and execute rights; mint per-job credentials that expire.
- Add reversibility: send delays, staged record changes, soft deletes, expiring offers.
- Move every code-checkable rule out of the prompt into validators or gates.
- Define the governance hierarchy; enforce hard constraints at shared choke points.
- Give every rule an owner, status, and version, with change records and rollback targets.
- Model consent, holds, promises, preferences, and sensitive context as typed customer fields checked before any agent acts.
- Emit a decision trace for every tier 2 and tier 3 action.
- Build review screens with evidence, one-click rejection with reasons, and seeded test cases.
- Treat model, prompt, and guideline changes as releases with replay, held-out, and regression checks.
Metrics to watch
- Gate outcomes by tier: blocked, approved, overridden, bypassed per thousand actions; every bypass explained.
- Trace completeness: share of tier 2 and 3 actions with a full decision trace.
- Customer-constraint violations: actions that contradicted a hold, opt-out, or promise.
- Review health: approval rate, median review time, catch rate on seeded bad cases.
- Rule freshness: share of agent calls served the current approved version of each rule.
Reader Q&A
Does every customer-facing message need human approval? No. It needs a tier. Pilots, new programs, new claim classes, and tier 3 actions get humans; mature, validated work runs with sampling and population monitoring.
Is our AI vendor responsible if its agent says something wrong? Do not count on it. Air Canada argued its chatbot was responsible for its own actions, and the tribunal called that "a remarkable submission" [1]. Customers and tribunals tend to hold the deploying company to what its agent said. Not legal advice.
Who should own AI governance? Domain owners own their rules; a platform owner owns the layer that serves and enforces them; a named executive owns the irreversible-action list.
For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 18
concepts:
- name: Governance is a code path
definition: "Rules are executed inside the system (before the model, around its tools, after its output), so a rule holds whether or not the model remembers it."
- name: Policy Sandwich
definition: "Deterministic eligibility and context assembly before the model, probabilistic reasoning and writing in the middle, deterministic validation and action gating after."
- name: Three-layer guardrails
definition: "Before (eligibility, consent, routed rules), during (capability boundary on tools and credentials), after (output validation and action gate)."
- name: Governance hierarchy
definition: "Explicit precedence: hard constraints > enterprise policy > program contract > channel policy > task instruction > customer context > model preference; customer-level constraints outrank campaigns."
- name: Customer-level governance
definition: "Per-person consent, contact rules, preferences, promises, sensitive context, and suppression stored as typed fields and enforced as constraints on every agent."
- name: Decision trace
definition: "Record per consequential action of what was known (with provenance), which rule and model versions governed it, what was decided and why, who allowed it, and what happened."
- name: Irreversible-action list
definition: "The short list of actions that cannot be undone: send, money, commitments, records, data leaving."
- name: Human-supervised autonomy
definition: "Design pattern: agents may analyze, diagnose, propose, and validate outside production; production changes require two human approvals, and the agent holds no production write credentials."
- name: Allow-list governance
definition: "Enterprise rules written as what the agent may say (capability map, approved statistics, known customer relationships) rather than only what it may not."
- name: Approve the judge, not each verdict
definition: "A human promotes an automated routine's stage (propose-only, limited, full) on evidence instead of approving each item; demotion is always available."
decision_rules:
- if: "a rule can be checked by code"
then: "enforce it in a validator or gate, not only in the prompt"
- if: "an action is on the irreversible-action list"
then: "apply the control hierarchy in order: remove authority, make reversible, bound in code, require approval above threshold, monitor"
- if: "a customer-level constraint (hold, opt-out, promise) conflicts with a campaign"
then: "the customer-level constraint wins, evaluated before the agent is invoked"
- if: "a rule, prompt, or model version changes"
then: "treat it as a release: record owner, reason, approver, validation results, and rollback target"
- if: "human approval rate is at or near 100 percent or review time is too short to read evidence"
then: "suspect rubber-stamping; seed known-bad cases and redesign the review screen or the tiering"
- if: "an agent may name capabilities, statistics, or a customer's relationship with you"
then: "restrict it to approved allow-lists; if the relationship is not known in the data, reframe as an evaluation"
- if: "a routine takes many similar consequential actions (merges, record changes)"
then: "make each action journaled and undoable, start the routine propose-only, and promote it one stage at a time on reviewed evidence"
- if: "a lower-level instruction would override a hard constraint"
then: "the hard constraint wins, enforced at a shared choke point"
assessment_questions:
- "Which component enforces each of your hard rules (discount limit, internal-data exposure, suppression), and what happens when it refuses?"
- "Where do promises and preferences stated by customers live, and can every agent read them before acting?"
- "Which irreversible actions can any agent take today without a gate?"
- "Can you reconstruct, for a given customer message, the facts used, rule versions, model version, and approver?"
- "Who owns each governance domain, and how does a rule change reach every agent?"
- "What are the approval rate and median review time in your human review queues?"
patterns: [Policy Sandwich, Irreversible-Action Gate, Governance Hierarchy, Three-layer guardrails, Two Approvals, Shared Choke Points, Allow-list governance, Approve the judge not each verdict]
anti_patterns: [The Broken Promise, Governance by Document, Prompt as Policy, Uniform Oversight, Rubber-Stamp Review]
maturity_dimension: governance_and_reliabilityReferences
- Moffatt v. Air Canada, 2024 BCCRT 149, British Columbia Civil Resolution Tribunal, February 2024. https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
- McCarthy Tétrault, "Moffatt v. Air Canada: A Misrepresentation by an AI Chatbot" (quotes from the decision). https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot ; see also American Bar Association, Business Law Today, February 2024. https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/
- H. Taheri, "14 Agent Configs, 3 Teams, Zero Source of Truth," hamedtaheri.com, 2026-03-14 (internal routing experiments: 25 governance variables, 5 categories, 20 task types). https://hamedtaheri.com/articles/fourteen-agent-configs-zero-source-of-truth
- ITV News, "DPD disable AI chatbot after it swears at customer and calls company 'worst delivery service'," 2024-01-19. https://www.itv.com/news/2024-01-19/dpd-disables-ai-chatbot-after-customer-service-bot-appears-to-go-rogue ; TIME, https://time.com/6564726/ai-chatbot-dpd-curses-criticizes-company/
- EU AI Act (Regulation (EU) 2024/1689), Article 50, transparency obligations; Commission FAQ. https://artificialintelligenceact.eu/article/50/ ; https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act (status as of September 2026)
- Regulation (EU) 2026/1744, "Digital Omnibus on AI." https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng (status as of September 2026; details in companion note)
- AI Incident Database, Incident 1039 (Cursor support bot, April 2025). https://incidentdatabase.ai/cite/1039/
- Fortune, "Cursor's AI glitch triggers viral fallout," April 2025. https://fortune.com/article/customer-support-ai-cursor-went-rogue/
- AI Incident Database, Incident 622 (Chevrolet of Watsonville chatbot, December 2023). https://incidentdatabase.ai/cite/622/ ; GM Authority, https://gmauthority.com/blog/2023/12/gm-dealer-chat-bot-agrees-to-sell-2024-chevy-tahoe-for-1/
- EU AI Act, Article 14, human oversight. https://artificialintelligenceact.eu/article/14/
- NIST, AI Risk Management Framework 1.0 (NIST AI 100-1), 2023-01-26. https://www.nist.gov/itl/ai-risk-management-framework
- NIST, Generative Artificial Intelligence Profile (NIST AI 600-1), 2024-07-26. https://doi.org/10.6028/NIST.AI.600-1
- ISO/IEC 42001:2023, Information technology: Artificial intelligence: Management system. https://www.iso.org/standard/42001
- EU AI Act, Article 5, and Commission Guidelines on prohibited AI practices, 2025-02-04. https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5
- H. Taheri, "Adversarial Governance Compliance: Our Methodology and What Near-Perfect Accuracy Tells Us," hamedtaheri.com, 2026-03-14. https://hamedtaheri.com/articles/adversarial-governance-compliance
- H. Taheri, "Governed Memory: A Production Architecture for Multi-Agent Workflows," arXiv:2603.17787, March 2026. https://arxiv.org/abs/2603.17787
This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.
Get chapters by email as they are revised