# 20. Accuracy and Quality

*The question: How do we achieve and keep accuracy at scale?*

> **Questions this chapter answers**
> - How accurate is "accurate enough" for personalization a customer will read?
> - Which checks should be code, and which need a model or a person?
> - Can a model check its own work, and can AI reliably supervise AI?
> - How do we find the errors no single review will ever see?
> - When something goes wrong, how do we fix the right thing instead of adding another paragraph to the prompt?

## The short answer

Accuracy in personalization is not "the model rarely hallucinates." A message is accurate when it is about the right person, uses facts trustworthy enough for the surface, makes only claims that follow from those facts, follows the campaign's rules, renders correctly, and arrives. A perfect sentence delivered to the wrong account is inaccurate. So is a correct page that never loads.

That makes accuracy a property of the system, built in layers, cheapest and most certain first:

1. **Deterministic controls** for everything code can decide: schemas, validators, allowlists, data-trust rules, and an abstention policy that makes the system say less when it knows less.
2. **Self-review** inside each generation step: cheap, useful, never the final word.
3. **An AI judge** for meaning code cannot check, calibrated against human labels, because judges have measurable biases.
4. **A stratified human sample**, so reviewers see the clean-looking records where blind spots hide.

The standard throughout: *specific and true, or honestly general.*

Keeping accuracy is a different job from reaching it. The first real cohort is part of system design, because it exposes defect patterns no test set anticipated. Every defect is diagnosed to the layer that caused it before anything changes, and every fix is validated against three sets: the failures that exposed it, a held-out sample, and previously good cases.

For a leader, the test is simple. Ask for the deterministic defect rate across the whole last cohort, and the semantic defect rate on a random sample of records that passed every automated check. If your team can give you both numbers, you have a quality system. If they show you five excellent examples, you have a demo.

## Two thousand messages, and the pattern no one could see

*Larkspur Systems is a fictional company used throughout this handbook. Its numbers are illustrative.*

Larkspur's first production cohort was 2,000 renewal-season messages to fleet operations leaders: an email, a personalized landing page, and a one-page brief for the account manager. Before launch, the team had read 25 of them closely. All 25 were excellent.

Then the engine ran a deterministic scan across all 2,000 rendered artifacts. It found 42 pages showing a headcount of exactly 10,001, a data provider's bucket floor that a trust rule should have rejected. It found 17 landing pages whose hero displayed a literal fragment of JSON, all from one template version that mapped the wrong zone key. It found 9 emails thanking a former customer, whose contract had lapsed, "as a valued customer." None of it was visible in the 25-record review, because none of it lived in those 25 records.

The more instructive finding came from the semantic pass. An AI reviewer read every flagged record plus a random sample of 200 records that had passed every check, and it noticed something no single record revealed. For accounts with thin data, the engine framed the message around "scaling your fleet" in almost four cases out of ten, including accounts whose telemetry showed a shrinking vehicle count. Each message was fluent on its own. Together they were a pattern, and it traced back to one worked example in the campaign guideline that happened to describe a growing fleet.

The fix was not a better model and not a longer prompt. It was a rejected-value rule, a repair in the shared rendering function, a customer-status check former customers could not pass, and a new guideline example. Four defects, four layers, four owners.

Larkspur's direct-mail team had caught the same 10,001 floor a year earlier, but that rule lived in the print workflow. A rule that lives in one channel protects one channel.

Accuracy is found in populations, fixed at the layer that caused it, and proved on records the fix was not designed around.

## Accurate means right person, true claims, right contract, delivered

Model benchmarks teach a narrow definition of accuracy: did the output contain a false statement? In an enterprise engine the test is the one in the short answer, plus two clauses: hard controls override softer preferences, and failures are visible, bounded, and recoverable.

Each clause belongs to a different layer: identity (Chapter 9), memory and freshness (Chapters 10 and 11), the generation contract (Chapter 15), governance (Chapter 18), and rendering and delivery. Accuracy breaks at whichever layer nobody tested. A second axis sits beside truth: appropriateness. A correct fact can still be wrong to use on this channel, for this person (Chapter 19). Truth checks and appropriateness checks are different checks, and both need an owner.

The standard that holds this together is **specific and true, or honestly general.** Verified, current, account-level evidence earns specificity. Thin evidence falls back to what is true for the company, then the industry, then an approved static version. Moving down that ladder is a designed outcome, not a failure.

> Specificity should degrade before reliability does.

## The accuracy stack runs cheap and certain first

Here is the sentence this chapter turns on: **never ask a probabilistic component to guarantee something a deterministic one can check.** Models reason. Guidelines express judgment. Code enforces invariants. Put as questions: AI asks whether this is semantically good and relevant; code asks whether it is structurally allowed and safe; operations asks whether the intended artifact actually arrived. Quality works when each control does the job it is good at, and fails when one control is asked to do all of them.

| Layer | What it answers | Cost per record | Coverage |
|---|---|---|---|
| Deterministic controls | Is this permitted and structurally safe? | Near zero | 100% of records |
| Self-review | Does this draft hold up against its context? | One extra pass, context already loaded | 100% of generations |
| AI judge | Is this semantically right for this person and contract? | A model call per record reviewed | Flags plus a stratified sample; full census for high-risk launches |
| Human sample | Is the judge right, and what is everyone missing? | Expert time | Small, stratified, recurring |

Anything the first layer can catch should never consume the attention of the fourth.

![The accuracy stack has four layers that run in order. Deterministic controls ask whether a record is permitted and structurally safe, at near-zero cost, on 100% of records. Self-review asks whether a draft holds up against its context, for one extra pass on 100% of generations. An AI judge checks semantic fit for this person and contract at a model call per record reviewed, on flags plus a stratified sample; a small, stratified, recurring human sample spends expert time checking whether the judge is right and what everyone is missing.](/images/handbook/ch18-accuracy-stack.svg)

*Figure 20.1. Cheap and certain controls run first, so scarce expert attention goes only to what code cannot decide.*

## If a rule can be deterministic, it should not depend on the model remembering it

Deterministic controls sit on both sides of generation.

**Before generation, constrain the context.** Data-trust rules decide which values the model may see. An employee count that matches a provider's bucket floor, falls outside a plausible range, conflicts five-fold with explicit research, or sits at 100 beside every marker of a multinational is rejected and removed from context entirely; the model must not be able to reason its way back into using it. An abstention policy turns uncertainty into an explicit state (known, likely, uncertain, rejected), and the rule for customer-facing copy is short: hide when uncertain.

**During generation, constrain the shape.** OpenAI reported its August 2024 model scored 100% on a complex schema-following evaluation in strict mode, against under 40% for an earlier model ([OpenAI, 2024](https://openai.com/index/introducing-structured-outputs-in-the-api/)); Anthropic documents the same schema guarantee ([Anthropic docs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs)). But a schema guarantees shape, not truth. A field typed as a string will always contain a string, not always the right one.

**After generation, check the claims.** A validator extracts concrete claims (named products, numbers, dates, events, relationship statements) and tests each against the verified context and campaign rules: no customer-status claim for a prospect, no leaked internal source name, no unsafe URL, every high-risk claim grounded in allowed evidence. "Your industry faces margin pressure" and "your company acquired a competitor for $480 million" are not the same kind of sentence, and the second deserves the strictest check.

Two lessons from operating this layer. When a section fails, replace the whole section with a true, general fallback; surgical edits to generated prose wreck its grammar. And put checks at shared choke points, the last function every output passes through before a boundary, so a new campaign inherits the protection without anyone remembering to add it.

The same rule applies to the plumbing under the checks: success must mean landed. In my memory system, a saturated database once accepted saves and returned success while writing nothing, and a failed embedding once came back as a success with an empty vector. Both now fail loudly, at the cost of a slower save. A quality stack reading from memory that silently did not update is checking the wrong facts.

A grounding check is not a universal factuality oracle; it covers only the claim classes it knows. In the engine I built it covers acquisitions, funding, and certain numeric assertions; an acquisition the evidence does not contain is flagged, rewritten, or held. But even a narrow check makes the most expensive mistakes impossible rather than unlikely.

## Self-review is a draft, not a verdict

Each generation step reviews its own output before the deterministic layer runs: re-read every specific claim, confirm it appears in the supplied context, remove it if it does not, and fill any gap with something true and general. It is cheap because the context is already loaded.

The research explains why it helps and why it cannot be the last word. Self-Refine had a model critique and revise its own output and reported roughly 20% average improvement across seven tasks ([Madaan et al., 2023](https://arxiv.org/abs/2303.17651)). Huang et al. found that on reasoning tasks "LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction" ([Huang et al., ICLR 2024](https://arxiv.org/abs/2310.01798)). A critical survey resolves much of the tension: self-correction "works well in tasks that can use reliable external feedback," and the authors found no prior work showing success from prompted-LLM feedback alone except in exceptionally suited tasks ([Kamoi et al., TACL 2024](https://arxiv.org/abs/2406.01297)).

I read that as a design instruction. Self-review works in personalization when it is not purely intrinsic: the model checks its claims against an external source, the verified customer context, rather than its own sense of what seems right. Two techniques make that explicit. **Chain-of-Verification** drafts, plans verification questions, answers them independently "so the answers are not biased by other responses," then revises ([Dhuliawala et al., 2023](https://arxiv.org/abs/2309.11495)). **SelfCheckGPT** samples several responses and treats disagreement between them as a fabrication signal, with no external database needed ([Manakul et al., EMNLP 2023](https://arxiv.org/abs/2303.08896)).

Self-review remains a soft control. The same probabilistic component should never be the sole judge of whether it followed the rules.

> A model grading its own homework has produced a better draft, not a review.

## AI can supervise AI, if you supervise the supervisor

Code cannot judge whether a sentence implies a relationship that does not exist, or whether a brief is too generic for an account the company knows well. For those questions, an independent AI reviewer is the only thing that scales.

The evidence that it can work is real, and so are the caveats. Zheng et al. found strong model judges "achieving over 80% agreement" with human preferences, "the same level of agreement between humans," and named position, verbosity, and self-enhancement biases ([Zheng et al., NeurIPS 2023](https://arxiv.org/abs/2306.05685)). Wang et al. showed that by changing only answer order, a much smaller model "could beat ChatGPT on 66 over 80 tested queries with ChatGPT as an evaluator" ([Wang et al., 2023](https://arxiv.org/abs/2305.17926)). Panickssery, Bowman, and Feng found evaluators score their own outputs higher than others' while humans rate them equal ([Panickssery et al., 2024](https://arxiv.org/abs/2404.13076)).

A judge is an instrument, and instruments need calibration:

- **Judge against the contract, not against "good."** The reviewer first resolves what the campaign promised (company name only, fully personalized, or a brief and score with no customer copy), then gets the trusted context and the final customer-facing output. A company-level experience should not fail for omitting a name; a brief promised as deeply personalized should not pass with generic industry copy. Written principles as the rubric is the move Constitutional AI made at training time ([Bai et al., 2022](https://arxiv.org/abs/2212.08073)).
- **Score atomic claims.** FActScore splits a generation into atomic facts and computes the share supported by a knowledge source; ChatGPT scored 58% on biographies ([Min et al., EMNLP 2023](https://arxiv.org/abs/2305.14251)). Here, the knowledge source is the customer's verified memory.
- **Evidence before verdict, positions swapped.** Ask the judge to cite supporting or contradicting context before rating, and swap order when comparing variants.
- **Different model family** from the generator where practical, to blunt self-preference.
- **Measure the judge** against a human-labeled set, as a first-class metric.
- **Challenge the blockers.** When the judge finds an issue serious enough to stop a release, a second reviewer tries to refute it first. False blockers cost campaigns.
- **Approve the judge, not each verdict.** In the memory system I built, an AI judge decides whether two records are the same person. Verdicts below a confidence floor are refused outright; a new judge only proposes until an operator reviews its proposals and promotes it, one stage at a time.

The output is a release verdict (GO, GO WITH FIXES, NO GO) with versioned thresholds a person can read. That is how QA ends in the personalization engine I built. For high-value launches, a domain expert also reads the first outputs before scale: a risk-based calibration gate, not a human review of every record.

## QA the thing the recipient sees, across the whole population

**Check the rendered artifact.** A perfect zone in a database can still become a broken page: wrong template version, wrong zone key, a hidden field, a stale cache, a page identity shared with another campaign, a missing URL, a static label that contradicts the generated copy. So QA runs on what the customer receives: the rendered page, the assembled email, the brief the rep opens, the print-ready PDF, the exact payload a downstream system accepts.

**Check everything cheap, sample everything expensive.** Deterministic checks run on 100% of records, which gives a real baseline. My population scan looks for classes code can name: a missing artifact, an empty stub brief, a suspicious firmographic, a personal email address that inherited a corporation's profile, a risky claim, a score of 88 labeled LOW. Semantic review then runs on every flag plus a stratified random sample of records that looked clean, stratified by evidence tier, persona, industry, template, and channel. The flags confirm known risks. The clean sample estimates what automation is missing, the number most teams never know.

> Sample the clean records, because that is where your blind spots live.

## The first real cohort is part of system design

A campaign is not finished when it passes its test records. The first few hundred real records reveal behavior the design team did not anticipate, because production data is stranger than any test set.

Site reliability engineering solved the release half of this long ago. Google's SRE workbook defines canarying as "a partial and time-limited deployment of a change in a service and its evaluation," judged on a few attributable metrics, "perhaps no more than a dozen" ([Google SRE Workbook](https://sre.google/workbook/canarying-releases/)). A pilot cohort is a canary with one addition: the reviewer looks for patterns, not only thresholds. I use AI as an analyst across the population, asking what no single record can answer: What errors repeat? Which personas expose a guideline gap? Where is the engine too generic for what we know? Which fallbacks fire too often? Which facts are consistently unreliable? Which strong outputs share a pattern worth keeping?

The protocol: define the contract; run a small real cohort; deterministic QA on all of it; semantic review on flags plus a stratified clean sample; population pattern analysis; diagnose each pattern to a layer; fix, validate, expand. My steps: synthetic records, then the first 10 to 25 real ones, then 50 to 100, then a few hundred, each read as a population before the next opens.

![The pilot cohort protocol runs in seven steps: define the contract, run a small real cohort, run deterministic QA on all of it, and run semantic review on every flag plus a stratified clean sample. Then comes population pattern analysis, asking questions such as which errors repeat and which fallbacks fire too often, followed by diagnosing each pattern to a layer, and finally fixing, validating, and expanding. It is a canary with one addition: the reviewer looks for patterns, not only thresholds.](/images/handbook/ch18-pilot-cohort-protocol.svg)

*Figure 20.2. The first real cohort is a canary: find population patterns and diagnose their layer before expanding.*

## Diagnose the layer before changing the system

The most common failure in AI operations is solving every problem by adding a paragraph to the prompt. The prompt grows, rules conflict, and nobody can say which instruction caused which behavior. Classify the cause first:

| Root cause | Typical response |
|---|---|
| Factual or safety defect | Deterministic prevention, a guideline rule, a regression test |
| Identity or data-quality edge case | Trust rule or classification, plus a safer fallback |
| Guideline defect | Update the guideline or its examples |
| Generator defect | Change the output contract or prompt architecture |
| Template or render defect | Fix the shared rendering function |
| Delivery defect | Retry, recovery, routing, or observability |
| Platform pattern (recurs across campaigns) | A shared platform mechanic, not a local patch |
| Campaign preference | Local configuration or guideline, not a platform change |
| Quality opportunity (correct, but could be better) | Messaging or guideline tuning, not a defect fix |
| Performance hypothesis | A controlled experiment, not a safety rule |

In my own memory system, validation before a release caught a weaker model putting an email address into the record-id field, creating orphan duplicates. A prompt paragraph would have lowered the rate; a code guard that coerces identifier-shaped values into the identity field removed the class.

## Two loops, three sets

Quality learning and performance learning share telemetry but must not share authority. The **quality loop** observes a defect, classifies its cause, fixes the guard, guideline, or fallback, and confirms recurrence drops. The **performance loop** forms a hypothesis about a content pattern, runs a controlled variation, and measures the outcome (Chapter 21). A system must never relax a factual guard because a bolder message earned more clicks.

> Correctness defines the feasible region. Optimization happens inside it.

Every change is validated against three sets, because a fix tested only on the records that exposed the problem overfits:

- **Replay set:** the records that exposed the defect.
- **Held-out set:** similar records not used to design the fix.
- **Regression set:** previously good records and known edge cases.

A change is ready when it reports: known failures fixed, held-out improved, good cases unchanged, no new blockers. Illustratively: 11 of 83 records show unsupported relationship language; the fix is one guideline update, one guard, six regression examples; replay 11 of 11, held-out 29 of 30, regression 40 of 40. Guidelines are versioned like code, with reason, evidence, validation results, approver, and rollback target.

Every QA run also writes a trend-ledger row (error rates, sample size, verdict, issue classes, versions). That turns "is today's sample okay?" into "are we improving or regressing across versions?"

![Two learning loops draw on shared telemetry but hold separate authority. The quality loop observes a defect, classifies its cause, fixes the guard, guideline, or fallback, and confirms recurrence drops; the performance loop forms a hypothesis, runs a controlled variation, and measures the outcome. Every change is validated on three sets: a replay set of the records that exposed the defect, a held-out set of similar records not used to design the fix, and a regression set of previously good records and known edge cases.](/images/handbook/ch18-two-loops-three-sets.svg)

*Figure 20.3. Quality and performance learning share data, not authority, and no change ships until it passes replay, held-out, and regression sets.*

Autonomy fits here only under Chapter 18's two approvals: a scheduled routine may run QA and propose a change with evidence, but a person approves the diagnosis and, separately, the proven release. In my engine, QA, the verdict, and the ledger run today and a person diagnoses; the loop that proposes and validates on the three sets is the next step, not a result.

## Quality is also maintenance

Accuracy decays even when nothing in the engine changes, because the world changes under the memory. Continuous background curation of customer memory, which this handbook calls dreaming, is the maintenance side of quality (Chapter 11). The quality system's job is to measure whether it works: how often generation abstains for stale evidence, and how often a curated fact is later contradicted. In my memory system, every curation pass reports writes applied, proposed, and refused; a record that fails three passes in a row is quarantined for a week; and two routines that keep rewriting the same property are reported as an oscillation rather than left to fight silently.

## What this does not do

A chapter about accuracy that overstated its own case would refute itself, so here are the limits.

**Hallucination is not solved.** Vectara's grounded-summarization leaderboard, a narrow and favorable task, showed rates from about 1.8% to 24.2% across listed models as of September 2026 ([Vectara](https://github.com/vectara/hallucination-leaderboard)). Open generation about a specific person is harder. The stack contains errors; someone still owns the residual.

**A judge is not an oracle.** Eighty percent agreement is the level at which humans agree with each other, so a judge will be confidently wrong on a meaningful share of hard cases. A judge never measured against human labels is an opinion with a number attached.

**Deterministic checks are silent about unnamed errors.** They make known classes impossible and say nothing about the rest, which is why the clean sample exists. On my own system, a configuration name mismatch once disabled semantic ranking, so retrieval quietly returned the newest memories instead of the most relevant, and every check passed. My paper on governed memory names silent quality degradation as a structural problem for this reason ([Taheri, 2026](https://arxiv.org/abs/2603.17787)).

**Small evaluations are coverage tests.** I have published a result I am proud of and careful about: 100% policy compliance across 50 adversarial governance scenarios in our own tests (the paper reports 100% adversarial governance compliance for the same architecture). It shows the architecture held under the pressure we designed, not under pressure we did not imagine. Treat any vendor's accuracy number the same way, including mine.

**Accuracy is the floor, not the goal.** A message can be specific, true, and useless. Whether it was worth sending is a decision problem (Chapter 14) and a measurement problem (Chapter 21).

I also disagree with a popular shortcut: "add a reflection step and the model will fix itself." Reflection earns its place only when it has evidence to reflect against (Huang et al.; Kamoi et al.).

## At scale

At scale, rates become counts. A 2% defect rate is 40 messages in a 2,000-record pilot and 20,000 across a million. The reader who gets one does not experience a rate; they experience a company that got them wrong.

Three things change with volume. The human sample stays small while the population grows, so stratification matters more than size. Cost pushes semantic review toward sampling, which makes judge calibration more important. And every model update is a release: in January 2024 a parcel carrier disabled its support chatbot after a system update left it swearing at a customer ([ITV News, 2024](https://www.itv.com/news/2024-01-19/dpd-disables-ai-chatbot-after-customer-service-bot-appears-to-go-rogue)). Model and guideline changes go through the same three sets and canary as code. Across many teams, protections in shared functions are inherited by every new campaign; protections in individual prompts are inherited by none. Once several campaigns run at once, the next step I would take (a design, not yet a result) is scheduled portfolio QA: the cheap deterministic scan across every active campaign, with semantic review aimed at flags, random clean samples, and any campaign whose trend-ledger row regressed.

## Failure story: The Self-Graded Exam

*A composite of a pattern that recurs.*

A team used the same model that generated its messages, with a "review this for accuracy" prompt, as the quality gate. The gate passed 97% of outputs, and leadership approved a full launch on that number. The first complaint came within days: a message congratulating an account on an expansion that had been cancelled. A human review of 100 passed records found unsupported claims in a double-digit share, clustered in accounts with thin data. The judge shared the generator's blind spot: it had no evidence to check against, and a measurable tendency to favor text like its own.

The fix was the stack: trust rules and abstention before generation, claim-level grounding after it, a judge from another model family scoring atomic claims against memory, and a human-labeled set to measure that judge. **The Self-Graded Exam** is a quality number produced by the component being graded.

## Patterns

**1. The Accuracy Stack.** *Problem:* one control, usually the prompt, is asked to guarantee truth, compliance, and quality at once. *Forces:* code is cheap and certain but narrow; models are broad but probabilistic; humans are accurate but scarce. *Solution:* deterministic controls, self-review, a calibrated independent judge, and a stratified human sample, in that order. *Tradeoffs:* more components; trust rules need upkeep; calibration needs a labeled set.

**2. The Pilot Cohort Protocol.** *Problem:* test records lack production's defect classes, and demo reviews hide the long tail. *Forces:* speed to launch versus exposure. *Solution:* treat the first real cohort as a canary, with population-level pattern analysis and layer diagnosis before expanding. *Tradeoffs:* a slower first launch; tooling to render and inspect final artifacts.

**3. Three Validation Sets.** *Problem:* fixes tuned on the failures that exposed them quietly break other cases. *Forces:* pressure to ship; few failing examples. *Solution:* report replay, held-out, and regression results for every change, and promote only with human approval and a rollback target. *Tradeoffs:* the regression set is ongoing work; small held-out sets are noisy.

## Leader questions

1. What was our deterministic defect rate across 100% of the last cohort, and the semantic defect rate on a random sample of records that passed every check?
2. Which quality checks are code, and which depend on a model remembering an instruction?
3. How well does our AI judge agree with human reviewers, and when did we last measure it?
4. When we fixed the last quality problem, which layer did we change, and how do we know nothing else broke?
5. Can any automated process change production behavior without a person approving it?

## Build checklist

- [ ] Data-trust rules reject implausible or conflicting values, and rejected values leave the context.
- [ ] An explicit specificity ladder (person, account, segment, approved static, fail closed).
- [ ] Structured outputs everywhere, treated as shape guarantees only.
- [ ] Claim-level validators for high-risk claims at shared choke points; failed sections replaced whole.
- [ ] Self-review that checks claims against supplied context.
- [ ] An AI judge scoring atomic claims against contract and memory, ideally from another model family.
- [ ] A human-labeled set measuring judge-human agreement, refreshed on a schedule.
- [ ] QA on rendered, delivered artifacts.
- [ ] Root-cause taxonomy for every defect; replay, held-out, and regression sets for every change.
- [ ] A trend ledger row per QA run (error rates, sample size, verdict, issue classes, versions).
- [ ] Versioned guidelines with two-approval promotion for automated proposals.

## Metrics to watch

- **Deterministic defect rate** across 100% of each cohort, by defect class.
- **Semantic material-issue rate** on the stratified clean sample.
- **Judge-human agreement**, tracked across model updates.
- **Abstention and fallback rate** by evidence tier (too low may mean invented specificity; too high, weak data).
- **Recurrence rate** of fixed defect classes.

## Reader Q&A

**How accurate is accurate enough?** It depends on the consequence of the surface. A wrong claim in an internal seller brief is caught by a person; the same claim in a customer email is not. Set thresholds per surface and claim class, and version them.

**Should the judge review every record?** For high-risk launches, a full census can be justified. In steady state, review every flag plus a stratified clean sample large enough per stratum to notice a pattern.

**Who owns accuracy?** Each layer has an owner, and one person owns the release verdict. If nobody owns the verdict, everyone owns the incident.

## For your AI

```yaml
chapter: 20
concepts:
  - name: Enterprise accuracy
    definition: "End-to-end property: right person, trustworthy facts, claims that follow from them, contract followed, correct render, confirmed delivery, bounded and recoverable failure."
  - name: Honestly General
    definition: "The accuracy standard (specific and true, or honestly general): use specificity only where evidence supports it; otherwise degrade to a less specific, true message."
  - name: Accuracy Stack
    definition: "Deterministic controls, then self-review, then a calibrated AI judge, then a stratified human sample; cheap and certain first."
  - name: Abstention policy
    definition: "Uncertainty is an explicit state (known, likely, uncertain, rejected); customer-facing copy hides what is uncertain."
  - name: Pilot Cohort Protocol
    definition: "Treat the first real cohort as a canary: deterministic QA on all records, semantic review on flags plus a stratified clean sample, population pattern analysis, layer diagnosis, then expand."
  - name: Three validation sets
    definition: "Every change is tested on replay (exposing failures), held-out (similar unseen records), and regression (previously good and edge cases) sets."
  - name: Two learning loops
    definition: "A quality loop that fixes defects and a performance loop that runs experiments; they share telemetry but not authority."
  - name: Trend ledger
    definition: "One row per QA run (deterministic and semantic error rates, sample size, verdict, issue classes, guideline and template versions), so quality is compared across versions, not judged per sample."
decision_rules:
  - if: "a rule can be checked by code"
    then: "enforce it deterministically; do not rely on a prompt instruction"
  - if: "a value fails a data-trust rule"
    then: "remove it from the model's context entirely"
  - if: "evidence for a specific claim is missing, stale, or conflicting"
    then: "degrade to account, segment, or approved static content; never invent to stay specific"
  - if: "a generated section fails validation"
    then: "replace the whole section with a true, general fallback"
  - if: "an AI judge is used for release decisions"
    then: "measure its agreement against human labels and use a different model family from the generator where practical"
  - if: "the judge raises a release blocker"
    then: "run a second reviewer that tries to refute it before stopping production"
  - if: "an AI judge makes decisions that write to production state"
    then: "refuse verdicts below a confidence floor, start the judge in propose-only mode, and promote it stage by stage on reviewed evidence"
  - if: "a defect is found"
    then: "classify the responsible layer before proposing a fix"
  - if: "a change is proposed"
    then: "validate on replay, held-out, and regression sets; promote only with human approval and a rollback target"
  - if: "a model or provider version changes"
    then: "treat it as a release and run the same validation and canary"
assessment_questions:
  - "What is your deterministic defect rate across 100% of the last cohort?"
  - "Do you review a random sample of records that passed every automated check?"
  - "Is QA run on the rendered, delivered artifact or on stored text?"
  - "Which model judges your outputs, and how often does it agree with human reviewers?"
  - "Can any automated process change production prompts, guidelines, or rules without human approval?"
patterns: [Accuracy Stack, Pilot Cohort Protocol, Three Validation Sets, Shared Choke Point, Specificity Ladder, Trend Ledger]
anti_patterns: [The Self-Graded Exam, The Prompt Paragraph Pile, The Demo-Account Sign-off, Fixes Tuned on the Failures Alone]
maturity_dimension: governance_and_reliability
```

## References

1. Madaan, A., et al. (2023). "Self-Refine: Iterative Refinement with Self-Feedback." NeurIPS 2023. https://arxiv.org/abs/2303.17651
2. Huang, J., et al. (2024). "Large Language Models Cannot Self-Correct Reasoning Yet." ICLR 2024. https://arxiv.org/abs/2310.01798
3. Kamoi, R., Zhang, Y., Zhang, N., Han, J., Zhang, R. (2024). "When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs." TACL. https://arxiv.org/abs/2406.01297
4. Dhuliawala, S., et al. (2023). "Chain-of-Verification Reduces Hallucination in Large Language Models." https://arxiv.org/abs/2309.11495
5. Manakul, P., Liusie, A., Gales, M. (2023). "SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection." EMNLP 2023. https://arxiv.org/abs/2303.08896
6. Zheng, L., et al. (2023). "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." NeurIPS 2023 Datasets and Benchmarks. https://arxiv.org/abs/2306.05685
7. Wang, P., et al. (2023). "Large Language Models are not Fair Evaluators." https://arxiv.org/abs/2305.17926
8. Panickssery, A., Bowman, S. R., Feng, S. (2024). "LLM Evaluators Recognize and Favor Their Own Generations." https://arxiv.org/abs/2404.13076
9. Bai, Y., et al. (2022). "Constitutional AI: Harmlessness from AI Feedback." https://arxiv.org/abs/2212.08073
10. Min, S., et al. (2023). "FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation." EMNLP 2023. https://arxiv.org/abs/2305.14251
11. Vectara. Hallucination Leaderboard (snapshot as of September 2026). https://github.com/vectara/hallucination-leaderboard
12. OpenAI (2024-08-06). "Introducing Structured Outputs in the API." https://openai.com/index/introducing-structured-outputs-in-the-api/
13. Anthropic. "Structured outputs" documentation (accessed September 2026). https://platform.claude.com/docs/en/build-with-claude/structured-outputs
14. Google. *The Site Reliability Workbook*, Chapter 18, "Canarying Releases." https://sre.google/workbook/canarying-releases/
15. ITV News (2024-01-19). "DPD disables AI chatbot after customer service bot appears to go rogue." https://www.itv.com/news/2024-01-19/dpd-disables-ai-chatbot-after-customer-service-bot-appears-to-go-rogue
16. Taheri, H. "How I Build an Enterprise Personalization Engine That Can Be Trusted at Scale." hamedtaheri.com, 2026-08-23. https://hamedtaheri.com/articles/building-enterprise-personalization-engine
17. Taheri, H. "Enterprise-grade accurate personalization at scale." hamedtaheri.com, 2026-07-13. https://hamedtaheri.com/articles/accurate-personalization-at-scale
18. Taheri, H. "Adversarial Governance Compliance: Our Methodology and What Near-Perfect Accuracy Tells Us." hamedtaheri.com, 2026-03-14. https://hamedtaheri.com/articles/adversarial-governance-compliance
19. Taheri, H. (2026). "Governed Memory: A Production Architecture for Multi-Agent Workflows." arXiv:2603.17787. https://arxiv.org/abs/2603.17787
