# 15. Writing for One

*The question: How do we generate for each person without inventing?*

> **Questions this chapter answers**
> - Can AI really write for each customer as well as a good person would?
> - How do we stop it from making things up about the people we write to?
> - How do we keep brand, tone, and claims consistent across a million generated messages?
> - What should the system say when it does not know enough?
> - Does one piece of generation logic have to be rebuilt for every channel?

## The short answer

Writing is no longer the hard part. A current model can produce a fluent, well-toned message for every person in your database, cheaply enough that the cost of the words barely registers. What it cannot do on its own is tell the difference between what it knows about a person and what merely sounds like something it could know. Language models are trained and evaluated in ways that reward a confident guess over an honest "I don't know" [3]. In a personalized message, a confident guess about the reader's own business is the most expensive kind of error you can make, because the reader is the one person guaranteed to catch it.

So the design goal changes. You do not ask the model for "a personalized email." You give it a contract: which claims it may make, from which evidence, in which tone, at what length, with what fallback. You break the message into small typed zones, each with its own job and evidence rule. You make every personal claim carry a receipt back to memory. When the evidence is thin, the system says less, on purpose, and stays useful by being honestly general. Then code, not the model, has the final word on the rules code can enforce, before one governed context is compiled into email, web, seller notes, or print.

For a leader, the test is simple. Ask your team: "Show me, for any sentence we sent, the evidence that allowed it." If they can, the program can scale. If they cannot, every additional message adds risk faster than it adds value.

The standard for this chapter, and for the handbook: **specific and true, or honestly general.**

## A beautiful email about an acquisition that never happened

*Larkspur Systems is a fictional composite company used throughout this handbook.*

Larkspur's AI SDR pilot had been running for six weeks when the reply came in. The message the pilot had sent to the operations director of a regional HVAC service company was, by any stylistic measure, excellent. It was short. It sounded like a thoughtful account executive. Its second sentence congratulated her on the company's recent acquisition of a competitor and suggested that merging two dispatch operations was exactly the moment to revisit fleet scheduling.

There had been no acquisition. A company with a similar name, two states away, had announced one. The pilot knew only rung-two facts about her (name, title, industry, company size), plus whatever an ungoverned web search on her name and company returned. That search returned the article, the prompt asked the model to "reference a recent company event," and the model did what it was asked: it found the most relevant-looking event and wrote about it with the confidence of someone who had checked.

Her reply was one line: "You clearly have no idea who we are."

Nobody at Larkspur had written a false sentence. The prompt was reasonable. The model was good. The article was real. The failure lived in the space between them, where nothing required the claim to be about the right company, to be sourced, or to be checked after it was written. That space is what this chapter is about.

## The writing got cheap; the right to be specific did not

Software can now produce natural language in the idiom, length, and tone a moment calls for, cheaply enough to do it separately for an enormous number of people. The old assumption that enough volume would reveal a template no longer holds. The evidence, read carefully, supports a narrower claim than the vendor pitch. In a study of 195 patient questions posted to a public forum, evaluators preferred chatbot answers to physicians' answers in 78.6% of evaluations and rated them more empathetic [7]. In a controlled three-party test, a persona-prompted GPT-4.5 was judged to be the human 73% of the time [8]. In marketing, AI-generated images matched or beat human-made ones, with up to 50% higher click-through than stock photography in a field test of more than 173,000 impressions [9]. And a Marketing Science study of machine-drafted SEO landing pages found the machine could produce unique, human-like content while concluding that the human editor remained essential [10].

Read together: in narrow, well-specified tasks, machine writing is at or above the quality most organizations could afford per record. None of those studies measured whether the writing was true about a specific reader, because none of them had to.

That is the gap. Fluency is not a guarantee of truth. A model can infer something useful about a person and confidently invent something false in the same paragraph, in the same assured voice; a 2020 study of summarization systems, asked only to condense a given article, found substantial unsupported content from all of them [1]. Retrieval [2] puts evidence in front of the model; it does not oblige the model to stay inside it. Kalai and colleagues argue the invention is not a mysterious defect: models are "optimized to be good test-takers, and guessing when uncertain improves test performance" [3]. A human writer who invents a customer's acquisition has a reputation to lose. The model has none. The accountability has to live in the system around it.

Measured rates make the stakes concrete. On Vectara's leaderboard, where models summarize a provided document and are scored on whether they added anything unsupported, rates ran from 1.8% for the best model to 24.2% for the worst, as of September 22, 2026 [4]. That is the easy case: one document, an instruction to use only it. A personalized message asks for more (combine facts, infer relevance, sound natural), which gives a model more room to fill gaps. My inference, not a measured number, is that open personalized writing without controls sits well above the summarization floor.

Here is the arithmetic, labeled as illustration. Larkspur has about 250,000 contacts. Suppose each message makes two personal claims and 2% are unsupported before checking: 10,000 unsupported claims per send. If checks catch 90%, 1,000 still reach readers. Change the assumptions and the number moves; it does not go to zero. The question is which errors escape, and what they cost.

## The model may write; it may not decide what is true

Here is the hinge of this chapter: **a generated sentence may be only as specific as the evidence behind it, and something other than the model decides how much evidence there is.**

Everything else follows. The model does semantic work: selecting, connecting, phrasing. Deciding which facts are trustworthy, which surfaces may carry them, and how specific this message is permitted to be are policy outputs computed before the model is called. I build this as a **generation contract**: a machine-readable specification per zone that turns "personalize this" into something testable.

```yaml
zone: account_relevance
purpose: "One sentence on why this offer matters to this account now."
max_words: 35
tone: "experienced operations advisor; no flattery; no urgency language"
allowed_evidence:
  - memory.account.industry          # confidence >= high
  - memory.account.recent_signals    # observed_at within 90 days, source attached
  - memory.contact.role
prohibited_context:                  # never enters the prompt at all
  - seller_internal_notes
  - unverified_relationships
forbidden_claims:
  - customer_status_unless_confirmed
  - unverified_numbers
  - events_without_source_and_date
  - internal_notes
fallback:
  tier: industry
  source: approved_static_copy
validation: [grounded_personal_claims, no_forbidden_claims, max_words]
```

`allowed_evidence` is the only material the zone may be specific about. `prohibited_context` is stronger than a prompt instruction: the zone never sees it. `forbidden_claims` names the claim classes that are expensive when wrong; with the relationship unknown, "your current contract" becomes an evaluation, not an ownership claim. `fallback` defines the safe version in advance. `validation` names the post-generation checks, so contract and QA agree on what "correct" means.

Had Larkspur's pilot run under that contract, the acquisition would have been an `event` without a verified entity match, filtered out before the model saw it. The model cannot misuse a fact it was never given.

## A message is a set of typed zones, not one long generation

An email looks like one piece of writing. Operationally it is several decisions: subject, opening relevance, the value claim, proof, call to action. A seller brief has an account snapshot, recent signals, conversation openers, and source references.

I decompose each into **typed zones**, and each zone carries its own job, evidence rule, length, risk level, and fallback. A generator producing a twelve-word headline cares about compression. A generator summarizing recent account signals cares about sources and dates. Asking one prompt to do both makes each worse and makes failures impossible to locate. With zones, a defect has an address: "the signal card cites events older than 90 days" is fixable; "the email feels off" is not.

In the governed personalization engine I built, a landing page is a governed shell of roughly 8 to 12 text zones (headline, value proposition, industry framing, proof, call to action); legal copy and navigation never vary. The value-claim zone is not asked what would sound good; it picks from the campaign's approved capability map, and in selected engines numbers come only from customer-approved statistics.

![An email is broken into five typed zones: subject, opening relevance, value claim, proof, and call to action. The opening relevance zone is expanded into its generation contract: its purpose is why this offer matters to this account now; its allowed evidence is industry, role, and sourced signals within 90 days; it forbids unverified numbers, unsourced events, and internal notes; it runs 35 words at most; its fallback is the industry tier with approved static copy; and it is validated for grounded claims, no forbidden claims, and length. Because each zone has its own contract, a failure has an address.](/images/handbook/ch14-typed-zones.svg)

*Figure 15.1. Each zone of a message carries its own contract, so a defect points to one zone instead of "the email feels off."*

Zones also make optimization honest. A cherry-picked win shows you the ceiling; at scale you live on the floor. Tune each zone against a spread of real records (the thin-data account, the company whose name collides with a product) until the output holds across all of them.

Decomposition invites a trap: a separate autonomous agent for every zone. A specialized worker earns its existence only with a distinct combination of context, instruction, output schema, tools, evaluation, and fallback. Often several zones come from one call and are split deterministically, with an expensive reasoning model reserved for the one high-risk synthesis step. The goal is clean responsibility boundaries, not the largest number of agents.

## Every personal claim carries a receipt

Governed memory (Chapters 10 and 11) stores each fact with its value, source, confidence, recency, and the surfaces it may appear on. Generation should preserve that chain rather than dissolve it into prose. The practical mechanism is a **claim grounding check**: extract the concrete claims from the output (every named company, event, number, date, product, and relationship) and test each against the evidence the zone was allowed to use.

This is the production form of an idea from research. FActScore breaks generated text into atomic facts and measures what fraction a reliable source supports; in its original evaluation ChatGPT's biographies scored 58% [11]. Chain-of-Verification has a model draft, plan verification questions, answer them independently, and then revise, which reduced hallucination across several tasks [12]. In a personalization engine, the knowledge source is not Wikipedia. It is the person's own record, and every claim should cite it.

Not every claim needs the same scrutiny. "Field-service companies are under pressure on technician utilization" is a segment-level truth. "You acquired a competitor last month" is a high-risk personal claim. Named events, amounts, customer-status claims, exact headcounts, and executive changes get the strictest rule: grounded in allowed evidence, or removed.

In the engine I built, deterministic grounding runs today for selected high-risk classes (acquisitions, funding, certain numeric assertions) and flags, rewrites, or holds the output. It is not a universal fact checker. Full provenance on every customer-facing statement (claim, source, date, confidence, surfaces, content version) is the next step, not a result I can claim.

When a zone fails, replace the whole zone with its approved general version. Do not edit phrases in place. Surgical edits to generated prose wreck its grammar and leave sentences that half-assert the removed claim. Replace the section; do not repair the sentence.

## Specific and true, or honestly general

When the evidence is thin, the system should say less, on purpose. I use three tiers, which I call the **Specificity Ladder**: the mechanism by which a message steps down, rung by rung, toward honestly general.

| Tier | What the evidence supports | What the copy does |
|---|---|---|
| **Full** | Trusted identity, corroborated account facts, a recent verified signal | Speaks to the specifics: the real event, the role, the named detail |
| **Account** | Solid company facts; the person or the recency is uncertain | Talks about the company and its situation; nothing framed as "recent" without a date |
| **Industry** | Little that can be verified for this account | Says what is genuinely true for companies like this one; no personal claims |

![The Specificity Ladder shown as three descending steps. At the full tier, evidence includes trusted identity and a recent verified signal, and the copy speaks to the real event, role, and detail. As evidence thins, the message steps down to the account tier, where company facts are solid but the person or recency is uncertain, and the copy talks about the company and its situation. It steps down again to the industry tier, where little can be verified, and the copy says what is true for companies like this one with no personal claims. The tier is set in code by evidence, never by ambition, and the copy never fabricates to climb a rung.](/images/handbook/ch14-specificity-ladder.svg)

*Figure 15.2. The Specificity Ladder: the message steps down as evidence thins, and honestly general is a valid place to land.*

The tier is chosen by evidence, never by ambition, and the copy never fabricates to climb one. It is also chosen in code. Research suggests models are partly calibrated about what they know when asked [5], which makes self-assessment a useful signal. It is not a policy. Chapter 9's rule applies here: if identity confidence falls, personalization specificity should fall with it, and that rule should not depend on the model noticing.

This reverses a common instinct. Teams treat the industry tier as failure and push to raise the "personalization rate." But a well-written industry message is a success; a specific-sounding message built on an invented detail is a failure even when it reads better.

## The model reviews; code decides

Before a zone's output is accepted, the model re-reads it against the context: does each specific claim appear in the evidence, does it answer the zone's job, did it invent a relationship or capability, did a source or the recipient's name leak where it should not? This is cheap and catches a lot, because the context is still in attention.

It is also a soft control. Huang and colleagues found that without external feedback, models do not reliably correct their own reasoning [13]. I never let the same probabilistic component be the sole judge of whether it followed the rules. After self-review, deterministic validation runs: required fields, length, prohibited phrases, formatting leakage, internal provider names, unsafe relationship claims, and the claim grounding check. If a rule can be deterministic, it should not depend on the model remembering to obey it. Deciding where each rule lives is a design act:

| Requirement | Where it lives |
|---|---|
| Sound like an experienced advisor | Guideline |
| Headline at most 14 words; HTTPS links only | Code |
| No ownership claims unless confirmed | Guideline plus a code relationship guard |
| Lower specificity when identity is weak | Code policy |
| Pick the most relevant proof point | Model, within approved evidence |

Use the model where judgment adds value. Use code where certainty is available.

One check is easy to forget: a call that returns "completed" with an empty load-bearing section has not succeeded. **Technical completion is not semantic completion;** retry boundedly, then fall back or hold.

![A five-step pipeline alternates between code and the model. Code supplies evidence with source, date, and confidence, and code sets the tier (full, account, or industry). The model writes the zone inside its contract and then self-reviews, which is only a soft control. Code then validates the output with deterministic rules and the claim grounding check. For each extracted personal claim there are three outcomes: grounded claims that match allowed evidence let the zone ship; an ungrounded claim replaces the whole zone with its approved general fallback; and if the fallback also fails, the system fails closed and does not send.](/images/handbook/ch14-model-writes-code-decides.svg)

*Figure 15.3. The model writes and reviews, but code sets the tier and has the final word: replace the section; do not repair the sentence.*

Structured outputs belong in this layer, understood correctly. Constrained decoding guarantees that a response parses against your schema; Anthropic's documentation describes "schema-compliant responses through constrained decoding," with some JSON Schema features unsupported [6], and OpenAI reported its 2024 model scoring 100% on a complex-schema evaluation [14]. That solved a real engineering problem. It did not solve truth. A perfectly valid `{"recent_event": "Acquired Northwind Services"}` is still false. **Structured outputs guarantee the shape of a claim, not its truth.** On the way back into memory, the self-hosted memory system I built treats a generated value bound to a record property as optional by default, so an empty or invented answer is not written; requiring it surfaces the miss instead.

One more boundary: treat model output as untrusted input wherever it is rendered. Escape it, parse and allowlist its links, normalize structured values before templates see them. A generated sentence that reaches HTML is a security surface as well as a content one.

## One context, compiled into many channels

A mature engine does not build a separate customer model per channel. Channels compile from one governed context, and a **channel compiler** applies each surface's policy. A verified facility expansion might become one sentence in an email, a card on a landing page, a sourced talking point in a seller brief, and nothing on the outside of a mailer, because that surface has a different privacy contract. The channel changes the expression, not the underlying truth, which is why the same facts never contradict each other across surfaces (Playbook P9).

The seller brief is the dangerous neighbor: its hypotheses must never be copied into a buyer email. So whatever leaves passes an explicit serializer that allowlists fields; a new internal property is never exposed merely because someone added it.

The same logic extends beyond text. Personalized pages, offer selection, and adaptive product interfaces are zones too: the model chooses among approved components, offers, and layouts and fills their slots; it does not emit free-form HTML or invent a discount. Template-and-zone pages ship today. Fully generated interfaces composed per person exist as demos and early products; treat broad production use as forecast, and apply the same contract when it arrives (see the Closing). Chapter 16 carries the contract beyond the message: zones across whole pages and product screens, a render gate in front of every generated value, and notifications as a decision; Chapter 3 gives the usability evidence for keeping the frame fixed.

## What this does not do

**It does not make generation true in general.** A grounding check covers only the claim classes it can detect. It catches the invented acquisition; it may miss a wrong implication ("as you scale your fleet" to a company that is shrinking). Semantic review and human sampling still matter (Chapter 20).

**It does not rescue bad memory.** If the record holds the wrong employer, a perfect grounding check faithfully grounds the wrong claim. Identity and freshness (Chapters 9 and 11) come first.

**It will not win every test.** Honestly general copy can open worse than fake-specific copy in the short run. I accept that trade, because the fake-specific version spends trust the business needs later; but measure it rather than assume it (Chapter 21).

**It does not settle "as good as a human."** The studies above show preference and indistinguishability in narrow tasks [7][8][9]. The same Marketing Science study that found machine drafts human-like kept a human editor in the loop [10]. The honest comparison is not your best account executive on her best day. It is the attention you could actually afford for each of 250,000 people.

A popular claim deserves naming: that constrained decoding or "JSON mode" prevents hallucination. It prevents malformed output. Anyone who tells you schema compliance is a factuality control is selling the wrong guarantee.

One more use for the contract: its `forbidden_claims` list is also where you write down what the system must never use, even when it would predict a response.

## At scale

At a hundred messages, a person can read every one. At 250,000, the system must answer three questions without a person in the loop: what did we say, what evidence allowed it, and which tier did each message run at. That requires the claim ledger to be stored, not just computed, and the tier distribution to be a dashboard metric. A sudden rise in industry-tier messages usually means an upstream research source failed quietly. A sudden fall usually means a check stopped running.

Cost scales with fan-out: zones times channels times retries times review passes. Run cheap deterministic checks first, spend model intelligence on judgment, and sample semantic review across apparently clean records rather than reviewing everything. Two patterns from the engine I built keep retries cheap. Conditional best-of-two: score the first set deterministically (blank, wrong length, repetitive opening, missing approved proof); if clean, ship it; if not, generate once more and keep the lower-penalty set, paying double only where the first draft failed. Partial retry: when two emails of a sequence fail, regenerate those two, not the sequence. Beyond those two, treat guidelines, contracts, and validators as production artifacts: version them, test them against replay, held-out, and regression sets, and promote them with approval. In my engine, a person diagnoses and decides each fix today; three-set validation with approval-gated promotion is the design I am building toward (Chapters 18 and 20).

## Failure story: The Hallucinated Detail

Back to Larkspur's acquisition email, diagnosed by layer rather than blamed on "the AI."

- **Identity:** the research result was matched to the account by name similarity alone. No entity confidence, no domain agreement.
- **Context:** the article entered the prompt as a bare string: no source, no date, no confidence, no allowed surfaces.
- **Instruction:** "reference a recent company event" demanded specificity without permitting abstention. It asked the model to climb a tier the evidence did not support.
- **Validation:** nothing extracted the claim or checked it. The self-review, had there been one, would have found the article in context and approved it.

The fix was not a better prompt. It was an entity-confidence gate on research, provenance on every fact, a contract that forbids unsourced events, a grounding check, and an account-tier fallback. The anti-pattern: **specificity demanded, abstention forbidden.**

The failure class is not invented for the story. In the engine I built, acquisitions and funding are exactly the classes that earned deterministic grounding checks, and unsupported relationship language ("your renewal is coming up") earned its own guard. The strongest controls came from observed failures, each turned into a guideline, a code backstop, and a regression test.

## Patterns

**Generation Contract.** *Problem:* "personalize this" is not testable. *Forces:* campaigns differ; hard rules must not depend on prompt wording. *Solution:* per-zone, machine-readable spec of purpose, allowed evidence, forbidden claims, tone, length, fallback, and validation. *Tradeoffs:* upfront design work; contracts must be versioned like code.

**Claim Ledger (claim grounding check).** *Problem:* fluent text hides unsupported claims. *Forces:* not every claim can be checked; high-risk ones must be. *Solution:* extract atomic personal claims, match each to allowed evidence with provenance, store the result, replace the failing zone with its general fallback. *Tradeoffs:* detector coverage is partial; extraction adds latency and cost.

**Channel Compilers.** *Problem:* each channel rebuilds its own view of the customer, and surfaces contradict each other. *Forces:* channels have different lengths, audiences, and privacy contracts. *Solution:* one governed context; per-channel compilers that enforce surface policy and serialize only what may leave. *Tradeoffs:* compilers need their own tests; the rendered artifact, not the stored text, must be QA'd.

## Leader questions

1. For any sentence we sent last week, can we show the evidence that allowed it?
2. What share of messages ran at full, account, and industry tier, and who decided those thresholds?
3. Which claim classes are forbidden, and are they enforced in code or only in a prompt?
4. When a message fails validation, what does the customer receive instead?
5. Do our channels read from one customer context or several?

## Build checklist

- [ ] Every zone has a generation contract with allowed evidence, forbidden claims, and a fallback.
- [ ] Personalization tier is computed from identity and evidence confidence before generation.
- [ ] Facts enter the prompt with source, date, and confidence, never as bare strings.
- [ ] Rejected facts are removed from context, not flagged for the model to weigh.
- [ ] Self-review runs inside each zone; deterministic validation runs after it.
- [ ] High-risk personal claims are extracted and grounded; failures replace the whole zone.
- [ ] Structured outputs enforce shape; nothing in the design treats them as a truth check.
- [ ] Model output is escaped and link-allowlisted at every rendering boundary.
- [ ] Channel compilers enforce per-surface policy from one governed context; anything that leaves passes an explicit field allowlist.
- [ ] Contracts and guidelines are versioned and regression-tested before promotion.

## Metrics to watch

- **Ungrounded personal claim rate**, measured on sampled output after all checks (the escape rate, not the catch rate).
- **Tier distribution** over time (full / account / industry / fallback / fail closed).
- **Fallback trigger rate by zone**, which points to the zone or data source that needs work.
- **Cross-surface contradiction rate** from semantic QA.
- **Reader-flagged errors** ("you have the wrong company") per 10,000 messages.

## Reader Q&A

**Can we just use a bigger model and skip the checks?** Better models invent less, and the leaderboard spread shows the difference is large [4]. None is at zero, and at scale small rates become thousands of messages. Keep the checks; enjoy the lower fallback rate.

**Won't honest generality make us sound like everyone else?** Only if it is written badly. A sharp industry-level message that names a real pressure in the reader's world beats a fake-specific one that names the wrong event. Write the fallbacks as carefully as the full tier; many recipients will see them.

**Should a human approve every message?** Not at volume, and approval fatigue makes it unreliable anyway. Put people where judgment is scarce: approving contracts, reviewing samples, and signing off on high-consequence or physical sends (Chapter 18).

**Does this apply to chat, not just outbound?** Yes. In a conversation, each turn is a zone with the same contract. A tribunal held Air Canada responsible for what its chatbot told a customer [15]; the company owns what its generator says, whatever the channel.

## For your AI

```yaml
chapter: 15
concepts:
  - name: Generation Contract
    definition: "Machine-readable per-zone specification of purpose, allowed evidence, forbidden claims, tone, length, fallback, and validation."
  - name: Typed Zones
    definition: "Decomposition of a message or experience into small units, each with its own job, evidence rule, risk level, and fallback."
  - name: Claim Grounding Check
    definition: "Extract atomic personal claims from generated output and verify each against allowed, provenance-carrying evidence; replace the zone on failure."
  - name: Honestly General
    definition: "A lower personalization tier, chosen by evidence, that makes no personal claim it cannot support and remains useful."
  - name: Specificity Ladder
    definition: "The full, account, and industry tiers; the mechanism by which a message steps down toward honestly general as evidence thins. Distinct from Graceful Degradation (Chapter 17), which handles dependency and system failures."
  - name: Channel Compilers
    definition: "Per-surface transformers that express one governed context under each channel's length, audience, and privacy policy."
decision_rules:
  - if: "identity or evidence confidence for a zone is below its contract threshold"
    then: "drop to the account or industry tier before generation; never ask the model to decide"
  - if: "a generated high-risk personal claim has no matching allowed evidence"
    then: "replace the whole zone with its approved general fallback; do not edit the sentence"
  - if: "a rule can be expressed as code (length, links, forbidden phrases, internal fields)"
    then: "enforce it deterministically after generation, not only in the prompt"
  - if: "the fallback itself fails validation"
    then: "fail closed and do not send"
  - if: "a fact is allowed on one surface but not another"
    then: "enforce it in that surface's channel compiler"
  - if: "a call reports completion but a load-bearing section is empty, or the first set fails deterministic scoring"
    then: "retry boundedly (one alternative, lower penalty wins; only failed slots), then fall back or hold"
assessment_questions:
  - "Can you show, for any sent message, which evidence allowed each personal claim?"
  - "Is the personalization tier computed in code before generation, or left to the prompt?"
  - "Which claim classes are forbidden, and where is each enforced?"
  - "Do all channels read one governed customer context?"
  - "What does a recipient receive when generation or validation fails?"
patterns: [Generation Contract, Typed Zones, Claim Ledger, Channel Compilers, Specificity Ladder]
anti_patterns: [The Hallucinated Detail, Specificity Demanded Abstention Forbidden, Schema As Truth, Agent Zoo, Sentence Surgery]
maturity_dimension: experience_generation
```

## References

1. Maynez, J., Narayan, S., Bohnet, B., McDonald, R. (2020). "On Faithfulness and Factuality in Abstractive Summarization." ACL 2020. https://aclanthology.org/2020.acl-main.173/
2. Lewis, P. et al. (2020). "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." NeurIPS 2020. https://arxiv.org/abs/2005.11401
3. Kalai, A. T., Nachum, O., Vempala, S. S., Zhang, E. (2025). "Why Language Models Hallucinate." arXiv:2509.04664. https://arxiv.org/abs/2509.04664
4. Vectara Hallucination Leaderboard (HHEM), accessed 2026-09-26; README updated 2026-09-22. https://github.com/vectara/hallucination-leaderboard
5. Kadavath, S. et al. (2022). "Language Models (Mostly) Know What They Know." arXiv:2207.05221. https://arxiv.org/abs/2207.05221
6. Anthropic, "Structured outputs," Claude Platform documentation, accessed 2026-09-26. https://platform.claude.com/docs/en/build-with-claude/structured-outputs
7. Ayers, J. W. et al. (2023). "Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum." JAMA Internal Medicine. https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2804309
8. Jones, C. R., Bergen, B. K. (2025). "Large Language Models Pass the Turing Test." arXiv:2503.23674 (preprint). https://arxiv.org/abs/2503.23674
9. Hartmann, J., Exner, Y., Domdey, S. (2025). "The power of generative marketing: Can generative AI create superhuman visual marketing content?" International Journal of Research in Marketing 42(1). https://www.sciencedirect.com/science/article/pii/S0167811624000843
10. Reisenbichler, M., Reutterer, T., Schweidel, D. A., Dan, D. (2022). "Frontiers: Supporting Content Marketing with Natural Language Generation." Marketing Science 41(3):441-452. https://pubsonline.informs.org/doi/abs/10.1287/mksc.2022.1354
11. Min, S. et al. (2023). "FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation." EMNLP 2023. https://arxiv.org/abs/2305.14251
12. Dhuliawala, S. et al. (2023). "Chain-of-Verification Reduces Hallucination in Large Language Models." arXiv:2309.11495. https://arxiv.org/abs/2309.11495
13. Huang, J. et al. (2024). "Large Language Models Cannot Self-Correct Reasoning Yet." ICLR 2024. https://arxiv.org/abs/2310.01798
14. OpenAI (2024-08-06). "Introducing Structured Outputs in the API." https://openai.com/index/introducing-structured-outputs-in-the-api/
15. Moffatt v. Air Canada, 2024 BCCRT 149 (2024-02-14). https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
