The New Era

What exactly changed with AI, and where is it going?

Working draft, revised September 2026 · 27 min read · Markdown for your AI

Questions this chapter answers

  • What can AI actually do for personalization today that it could not do three years ago?
  • Can it really understand a customer, and can it write as well as a person?
  • How cheap is it now, and why is my AI bill still going up?
  • How long can an agent work without a person, and where will that be in two or three years?
  • What is real, what is a demo, and what is a forecast?

The short answer

What changed is the price of understanding one person. For the whole history of business communication, reading a customer's situation closely enough to say something specific and true required a skilled human's hour. That hour was the binding constraint, so depth went to the few accounts that could pay for it and everyone else got the template.

Language models removed most of that constraint for six kinds of work, which I call the capability stack: read, reason, decide, write, act, learn. Each layer is real today, each is uneven, and none is free.

Three facts should shape your planning. First, the price of a fixed level of model capability has fallen by roughly 10x to several hundred times per year depending on the task, but agents consume far more tokens per task, so your spend per decision is an architecture choice, not a market price. Second, the length of task an agent can complete without a person has been doubling roughly every six to seven months on software benchmarks, faster since 2024; that is a measured trend on narrow tasks, not a guarantee about your customers. Third, fluency is not truth. The system that writes a brilliant account brief can invent a detail with the same confidence.

The practical conclusion: deep personalization for every customer is now an engineering and governance problem rather than a staffing problem. Everyone rents the same models. The winners will be the companies that build the memory, the decisions, and the checks that turn cheap understanding into accurate action.


Three hours, and then four minutes.

Larkspur Systems (a fictional composite company used throughout this handbook) sells field-service and fleet software to about 40,000 accounts. Priya, an account manager, has a renewal call Thursday with Brightwater Fleet Services (also fictional), a regional operator with 300 vehicles. On Tuesday she reads eighteen months of CRM notes. She skims fourteen support tickets and notices that nine of them came from the same dispatcher. She listens to half of a recorded quarterly review in which the operations director says, almost in passing, that the board wants fuel costs down before the next budget cycle. She finds a job posting for a "telematics analyst," a role Brightwater has never had. Her brief: the dispatcher is struggling with routing, the director is under cost pressure, and the new hire suggests they are about to take their data seriously. Lead with fuel-cost reporting, offer routing training, and do not open with the price increase.

That brief took about three hours. At the loaded cost of skilled selling time from Chapter 4, roughly $50 to $72 an hour, it cost Larkspur at least $150 to $220, and more in practice, because senior account managers cost more than that median. It was worth it, and impossible to repeat for more than a few hundred of 40,000 accounts. Doing all of them would take 120,000 hours (a full three-hour brief, not the 20-to-60-minute first touch Chapter 4 priced), roughly sixty full-time people for a year. So Larkspur does what every company does: deep work for the top accounts, a segment email for everyone else.

Now run the same job through an agent. It reads the same notes, tickets, transcript, and posting, plus public news, and writes a comparable brief in a few minutes. A run of this shape consumes somewhere around a hundred thousand to a million tokens, on the order of $0.15 to a few dollars per account at current prices for capable models. The low end matches Chapter 4's $0.15 to $1 research run; the high end is higher because this is agentic, multi-step reading of a transcript and a long history, not a single research pass. Across all 40,000 accounts, that is roughly $6,000 to $120,000, depending almost entirely on how the system is built.

The ratio is the story: the price of a first-pass understanding of one account fell by two to three orders of magnitude. It is also the change most easily misread.

A log-scale bar chart comparing the cost of one account brief. A person working about three hours costs $150 to $220, more in practice; an agent working a few minutes costs $0.15 to a few dollars, two to three orders of magnitude less. Below, across 40,000 accounts at Larkspur (fictional), people would need 120,000 hours, roughly sixty full-time people for a year, while agents would cost roughly $6,000 to $120,000 depending almost entirely on how the system is built.

Figure 5.1. On a log scale, a first-pass account brief fell from hundreds of dollars of skilled time to cents or dollars of model time.

Understanding got cheap; being right did not

Here is the sentence this chapter turns on: the cost of understanding a customer used to decide who got understood; now the decision moves from the budget to the architecture.

When understanding was expensive, the budget did the rationing; nobody had to decide which of 40,000 accounts deserved depth, because the arithmetic of human hours decided. When understanding is cheap, that implicit policy disappears, and three explicit questions replace it. Which decisions deserve depth? How do we know the understanding is true? What happens when it is wrong at scale? Most of this handbook answers those questions.

The capability stack: read, reason, decide, write, act, learn

Six layers arrived, at different speeds, and they fail in different ways.

LayerWhat it does for personalizationMaturity (as of 2026-09)
ReadExtracts facts and signals from calls, emails, tickets, documents, imagesShips now, widely
ReasonInterprets ambiguous situations; forms a view of the accountShips now, uneven
DecideChooses the next action for one person, including doing nothingShips now inside designed systems; weak when left to the model alone
WriteGenerates language for one person at human qualityShips now; accuracy is the constraint
ActRuns a bounded loop of tool calls without a person between stepsShips now for bounded tasks; long unattended runs are early
LearnCarries what happened into the next decisionShips now as memory and feedback you build; not as the model changing itself

It reads more than the customer said

Start with reading, because leaders underestimate it. These systems do not only recover stated facts. They infer relationships among the facts and propose what those relationships imply. Give a model a customer's support history, a call transcript, and a job posting, and it may infer, as Priya did, that a team is struggling with one module or that a new hire signals a change in how the customer uses data.

Those are inferences, not knowledge. Some will be wrong, and fluency can make a weak inference sound much stronger than it is. But the most valuable thing in a model of a customer is often not a fact they stated. It is a pressure, preference, relationship, or unfinished problem inferred across facts. Until recently that required a thoughtful person reading closely. A machine can now attempt it across a whole customer base.

The first-era systems in Chapter 2 could not. A propensity model reads columns, and the director's aside about fuel costs never became a column. Most of what a company knows about its customers sits in calls, emails, notes, and documents, and for the first time it is readable at the scale of the customer base.

In a governed personalization engine I built, account research lands as structured fields, not one long paragraph: industry, buying triggers, leadership changes, technology initiatives, stated priorities, challenges. Each field can be trust-checked on its own, and each surface renders only the fields it needs instead of asking a model to restate them. The next step I am building attaches confidence, source, and freshness to every field.

It works through an account the way an expert would

The deeper change is that the same medium can be used to work through a problem. A capable system can take an objective, break it into parts, compare routes, use evidence and tools, inspect a result, and revise its next move. It does not have to resemble human thought internally to reproduce some of the work expert reasoning performs externally.

Expertise was scarce not only because experts knew more facts, but because they could structure a difficult problem: identify what mattered, discard weak approaches, notice missing evidence, choose the next step. The best evidence that this transfers comes from the field. In a study of 5,172 customer-support agents, access to a generative AI assistant raised issues resolved per hour by 14% on average and by 34% for the least experienced workers, largely by spreading the practices of top performers to everyone else (Brynjolfsson, Li, Raymond, 2025). That is the pattern that matters for personalization: the reasoning of your best account manager becomes available on every account.

One caution matters in every later chapter. A displayed explanation feels like a judgment you have inspected, but it is not an audit trail. When researchers slipped hints into questions, reasoning models that used the hint mentioned it in their reasoning only 25% (Claude 3.7 Sonnet) to 39% (DeepSeek R1) of the time (Chen et al., 2025). The explanation is an output to evaluate, not a record of why the system acted. That record has to be built separately (Chapter 18 calls it the decision trace).

It decides, but only as well as the system around it

Reasoning produces a view. A business needs a decision: what to do next for this person, across thousands of people, under rules. Models can propose a next action with a rationale for every account overnight. What they cannot do alone is hold the organization's objectives, contact rules, and constraints consistently across 40,000 decisions. That comes from a designed decision layer where deterministic policy stays deterministic (Chapter 14). The capability to decide arrived; the reliability of deciding is something you build.

In the same engine, a model scores each lead against the campaign's criteria and returns a score, a band, a rationale, and a suggested approach, weighing weak and strong signals together the way a good account manager would. Code then checks that the pieces agree: a score of 88 labeled LOW is flagged, not sent. The model supplies judgment; the system supplies consistency.

It writes like a person, for each person

Software can now write in the idiom, length, and tone a moment calls for, cheaply enough to do it separately for every customer. In a 2025 three-party Turing test, GPT-4.5 prompted to adopt a persona was judged to be the human 73% of the time, more often than the real humans it was compared against (Jones and Bergen, 2025, preprint). In marketing, a set of studies found AI-generated images matched or beat human-made ones, and a field test with more than 173,000 impressions reported click-through rates up to 50% higher than stock imagery (Hartmann, Exner, Domdey, 2025).

That does not mean every output passes. It means the old assumption, that volume must reveal a template, no longer holds. The two halves of personalization, forming a useful account of someone and speaking to them as an individual, can now be joined in software. In the engine I built, one account context feeds a seller playbook, an email sequence, and a landing page, and the hypotheses a seller may see never reach the buyer.

It also means the capability cuts both ways. In a 2025 debate experiment, GPT-4 given six demographic facts about its opponent was more persuasive than a human in 64.4% of the pairs where one side was more persuasive; a 2026 author correction adds that its advantage over un-personalized GPT-4 did not reach statistical significance (Salvi et al., 2025). That is one reason this handbook spends a whole part on governance.

It acts: the loop that used to need a person

Running an account well is not one message. It is a loop: observe the customer's world, plan, act, read what comes back, adjust. A good account manager does this over weeks, and every turn needs their attention.

Software can now run a bounded stretch of that loop itself. You give it a goal, particular tools, and limits. It observes, decides, acts, reads the result, and decides again without returning to you between steps. The human moves from driving each step to setting the objective, granting the tools, defining the stopping conditions, and handling the exceptions. Traditional automation acted only along routes written in advance, which is why it was reliable. An agent is useful where the route cannot be fully written beforehand, which is also what makes it harder to predict and control.

In the systems I run, "bounded" is literal. In a self-hosted memory system I built, an outside agent reaches customer memory through a tool interface (MCP) limited to its role, and autonomous agents are never offered deletion. A long job can run on a model provider's own background service, carrying a credential minted for that job alone: three memory tools, revoked when the job ends, and the job cancelled after 24 hours by default. Autonomy is a set of limits written down in advance.

It learns, in a narrower sense than the word suggests

Update is the more accurate word. An agent carries new information from one step into the next without retraining the underlying model. Across sessions, it remembers only what you built a memory to hold. That is not a limitation to wait out. It is where most of the durable value sits (Part IV).

In my systems, learning ships today in two forms. Memory is curated while the system is idle, by passes that only propose changes until an operator promotes them (Chapter 11). And a person uses AI to read across the first cohorts of real output, asking which errors repeat and which fixes belong in code rather than in another instruction. AI that implements and validates its own changes before a person approves them is a design I am pursuing, not a result.

The letter became a conversation

Put reading, writing, and reasoning together in real time and personalization changes shape. The old message was a letter. It was written once, said the same thing to everyone in a segment, and could not respond. If the customer had a doubt the letter did not anticipate, the letter had no answer.

What arrived is a conversation, and a company can hold a different one with each customer. A support exchange can pick up where the last one ended. A website can answer the question this visitor is actually asking rather than the question the page was designed for. An onboarding flow can notice that a user is stuck and ask why, instead of sending day-three email number two. The same shift reaches pages, product screens, and notifications, which can now be composed for one person; Chapter 16 sets out how far that ships. A static message has to anticipate everyone. A conversation only has to answer this person, one turn at a time, and each reply tells the system more about what matters.

A letter is judged once, when it is written. A conversation is judged at every turn, so the rules, the memory, and the checks have to be present at every turn too. Personalization used to be a property of a message. It is becoming a property of a relationship.

The price fell; the bill did not

Now the economics, where most leadership conversations go wrong. The price of a fixed level of capability has collapsed. The Stanford AI Index reports that the cost of querying a model scoring at GPT-3.5's level on the MMLU benchmark fell from $20.00 per million tokens in November 2022 to $0.07 in October 2024, a more than 280-fold reduction (Stanford HAI, AI Index 2025). Epoch AI, tracking the cheapest model to reach fixed thresholds across six benchmarks, found prices falling between 9x and 900x per year depending on the task; GPT-4-level performance on MMLU went from $37.50 to $0.18 per million tokens between March 2023 and February 2025, about 40x a year (Epoch AI, 2025-03-12). Epoch cautions that the fastest drops were the most recent and may not persist.

So why are AI bills rising? Because nobody buys a fixed capability for a fixed task. Agents read more, reason longer, and call tools repeatedly. Anthropic reported that in its own systems, agents typically used about 4x more tokens than chat interactions and multi-agent systems about 15x more, and that token usage alone explained 80% of the performance variance on the evaluation it studied (Anthropic, 2025-06). Better results are bought with more reading.

Two linear-scale charts side by side. Left: the price per million tokens for GPT-3.5-level performance on MMLU fell from $20.00 in November 2022 to $0.07 in October 2024, more than 280-fold cheaper. Right: in Anthropic's own systems, agents used about 4x the tokens of chat and multi-agent systems about 15x. Beneath them, price per token is set by the market, while tokens per decision are set by your architecture.

Figure 5.2. The market sets the price per token; your architecture sets how many tokens each decision consumes.

This is where I part company with the popular reading of "LLMflation," a16z's observation that the cost of a given capability falls about 10x a year (Appenzeller, 2024). The observation is right. The conclusion people draw, that AI cost will take care of itself, is wrong. Price per token is set by the market. Tokens per decision are set by your architecture. A pipeline that researches the same company once per employee, or loads the whole customer history into every call, makes cheap tokens expensive. On a production pipeline I run, per-record research bought the same company's research close to a hundred times in one day for one popular employer. Moving research to the company, cached for seven days, cut that spend and raised the share of records with rich research from 60.3% to 93.5%. The budget stopped paying for answers it already had.

Where the model runs is an architecture choice too. In the self-hosted memory system I built, each function gets its own model: a cheap one for extraction, a strong one only where judgment pays. Bulk loads can go through a provider's batch interface at roughly half the per-token price. A local model removes the per-token bill entirely and trades it for hardware and lower concurrency.

The unit to manage is cost per decision, not price per token. Chapter 12 does that accounting.

How long it can go without you

How far does autonomy go today? The honest answer is a moving line. One useful measure is not how smart a system seems but how long a task it can complete before it needs a person. METR estimates this "time horizon" by comparing an AI system's success on tasks with the time skilled people take to do them. In its original 2025 analysis, the length of task frontier systems could complete with 50% reliability had been doubling roughly every seven months for six years (METR, 2025-03-19). An updated methodology published in January 2026 put the long-run doubling at about 196 days, and at about 89 days for models since 2024; its top measured model at the time, Claude Opus 4.5, had a 50% horizon of about 320 minutes, with a confidence interval from 170 to 729 minutes (METR Time Horizon 1.1, 2026-01-29). By May 2026, METR's tracking page warned that measurements above 16 hours are unreliable with its current task suite, which tells you where the frontier had reached.

Three qualifications keep this honest. The tasks are mostly software engineering, machine learning, and cybersecurity; METR itself describes capability as jagged across domains. The horizon is measured against people with little prior context, not experienced staff who know the account. And the success threshold matters as much as the length. A task completed half the time is impressive in a demonstration and unusable in a process that must work every night; require eight successes in ten and the horizon becomes far shorter. Capability changes what is possible. Reliability decides how the work must be supervised.

Field evidence agrees. In a randomized trial of 16 experienced open-source developers working on 246 real issues, early-2025 AI tools made them 19% slower, while the developers believed they had been sped up by about 20% (METR, 2025-07-10). Gains are real and uneven, and the difference is in where you point the system.

For personalization, full autonomy is not required for the operating model to change. If a system handles the repeatable middle of the account loop and returns only the ambiguous or consequential moments, one person can supervise far more accounts than they could personally work. When I build these systems, keeping them on task is most of the work, and the gain comes from deciding in advance which failures are tolerable and which uncertainty sends the case back to a person.

Ships now, lab demo, forecast

Here is the map I would give a leadership team, with each claim labeled by kind. Dates matter: this is as of September 2026.

Three columns sorting AI capabilities by kind of claim, as of September 2026. Ships now: reading calls, emails, and tickets; a brief for every account; individually written messages; conversational personalization; and bounded agent tasks running for hours, with accuracy depending on checks. Lab demo or early production: agents managing a relationship for weeks, and interfaces and offers composed per person, not yet common practice. Forecast: reliable day-long tasks and per-decision cost that keeps falling, both conditional and not guaranteed.

Figure 5.3. Much of the stack ships now; long, lightly supervised autonomy is still a lab demo or a forecast.

CapabilityStatusWhat the evidence supports
Extracting facts and signals from calls, emails, tickets, documents at customer-base scaleShips nowRoutine in production; accuracy depends on schemas and checks (Chapter 8)
First-pass account research and briefs for every accountShips nowMinutes and cents to dollars per account; quality depends on source reconciliation
Individually written messages at human qualityShips nowFluency is solved for most business writing; factual accuracy is not (see limits)
Conversational personalization in support, onboarding, and webShips nowDeployed widely; continuity depends on shared memory most companies lack
Agents running bounded multi-step tasks for hoursShips now, boundedMeasured on software tasks; business-process reliability varies by design; in my systems, bounded by role-scoped tools, a one-job credential, and a time cap
Agents managing a customer relationship for weeks with little supervisionLab demo / early productionPossible with durable memory and checkpoints (Chapter 13); not common practice
Interfaces and offers composed per person at request timeLab demo / early productionPer-account pages generated before the visit, inside declared zones, run in production in my engine; composing at request time for anonymous visitors is pilot design; measured lift at scale is thin (Chapter 16; Closing chapter)
AI improving its own rules and memoryEarly production, propose-onlyShips as curation that proposes until an operator promotes it and as AI-assisted cohort review; AI implementing and validating its own changes is design direction
Agents reliably completing day-long business tasks at high success ratesForecastInference from the METR trend, if it holds and transfers beyond software; neither is guaranteed
Per-decision cost continuing to fallForecast, conditionalPer-token prices likely keep falling; per-decision cost falls only if architecture controls tokens

Adoption is moving too. The 2026 AI Index reports that generative AI reached 53% population adoption within three years, faster than the personal computer or the internet (Stanford HAI, 2026-04-13). That is observed adoption of the tools, not evidence that companies are personalizing well with them.

What this does not do

A chapter about new capability is most tempting to oversell, so here are the limits, stated plainly.

Fluency is not truth. Models trained through prediction can infer something useful about a customer and confidently invent something false. They can produce a strong analysis and then accept a bad premise, or construct an explanation that sounds complete because sounding complete is part of what they learned. Research on why models hallucinate argues that standard training and evaluation reward guessing over admitting uncertainty (Kalai et al., 2025). On Vectara's grounded summarization benchmark, where the model only has to summarize a document it was given, hallucination rates ranged from about 1.8% for the best models to over 24% for the worst (Vectara leaderboard, snapshot 2026-09-22). Open-ended personalization is harder than summarizing a supplied document. At a million messages, a 2% error rate is 20,000 messages that say something false to a customer about their own business.

Lab accuracy is not production accuracy. Benchmarks use clean inputs. Your inputs include duplicate records, stale titles, and two people with the same name at the same company. A single mistaken identity can contaminate every later step, because the system treats its own earlier conclusion as context. Autonomy compounds both useful work and error.

The persuasion studies are weaker than the headlines. Matz and colleagues reported that messages generated to fit a person's inferred psychology were more influential (Matz et al., 2024), and that finding's robustness has been publicly disputed (PNAS commentary). The Salvi correction above narrows what can be attributed to personalization itself. The honest claim is that personalization shifts odds under some conditions, not that it commands outcomes.

Cheap understanding does not mean you should use all of it. More detail can make a message worse when it is irrelevant, stale, or included to prove the sender found it. The capability to infer is not permission to say what was inferred (Chapter 19).

Nobody knows the slope in 2029. Anyone presenting the time-horizon trend as a schedule is selling something. I plan for continued improvement and design so the system still works if it stalls.

At scale

Across 40,000 accounts and 250,000 contacts, three things change that a single brief hides.

Errors stop being anecdotes and become rates. One wrong detail in Priya's brief she catches on the call. The same error rate across every account is a population of wrong sentences, measured as a population (Chapter 20).

Research multiplies unless it is shared. Five teams running five agents on one customer buy the same understanding five times and reach five conclusions. Write it once to a memory everyone reads (Chapter 10).

The bottleneck moves to supervision. When research is cheap, the scarce resource is human attention for review and exceptions. Design where it goes, or it goes everywhere and nowhere.

Failure story: The Fluent Guess

A composite of a pattern I see often: a team at a company much like Larkspur switched on agent-written renewal briefs for every account. The briefs were excellent to read. One of them told an account manager that the customer was "expanding into Canada," based on a job posting for a remote role that listed Toronto among several cities. The manager opened the call with Canadian pricing. The customer had no such plan and asked, reasonably, where that had come from.

Nothing in the system distinguished what the account had said from what the agent had inferred. The Fluent Guess is the anti-pattern: inference presented with the confidence of evidence. The fix is not a better model. It is a system that records where each claim came from and how strong it is, and lets that strength decide how the claim may be used.

Patterns

1. Evidence-Weighted Voice. Problem: outputs state weak inferences as confidently as verified facts. Forces: readers trust fluent prose; inferences are often the most valuable content. Solution: attach source and confidence to every claim before generation; state only corroborated facts as fact; phrase inferences as hypotheses; fall back to honestly general when evidence is thin. Tradeoffs: fewer striking sentences; more work at extraction.

2. Research Once per Entity. Problem: per-record pipelines re-research the same company or person repeatedly. Forces: the record is the natural unit of work; the entity is the natural unit of knowledge. Solution: key research to the entity, cache it with a freshness window, and have downstream steps read the cache. Tradeoffs: needs identity resolution first (Chapter 9) and an invalidation policy (Chapter 11).

3. Human at the Boundary. Problem: either every step waits for a person, or no step does. Forces: agents handle the repeatable middle well and consequential edges poorly. Solution: automate the middle of the loop; route to a person on low evidence, high consequence, or irreversible action. Tradeoffs: requires explicit thresholds and a review queue people can actually use.

Leader questions

  1. For which decisions would a three-hour human brief change the outcome, and how many of those decisions do we make per year?
  2. What is our cost per decision today, not our price per token, and who owns that number?
  3. When our AI states something about a customer, can we show where it came from?
  4. Which of our AI claims to the board are ships-now, which are demos, and which are forecasts?
  5. Where does human attention go in our current AI workflows, and is it on the consequential cases?

Build checklist

  • Inventory unstructured sources (calls, emails, tickets, documents) and who can read them.
  • Run one account-research pipeline on 50 accounts and have a person grade the briefs.
  • Record source and confidence for every extracted claim.
  • Measure tokens and cost per decision, not per call.
  • Key research to the entity with a freshness window.
  • Define the evidence threshold below which output becomes honestly general.
  • Define which actions always return to a person.
  • Give every agent a role-scoped, expiring credential and a maximum run time.

Metrics to watch

  • Cost per decision (all model, tool, and data costs divided by decisions made).
  • Unsupported-claim rate in a graded sample of outputs.
  • Research reuse rate (share of research reads served from existing entity research).
  • Human touches per hundred decisions, split by reason.

Reader Q&A

Can AI really understand a customer? It can build a useful working account of one from what you give it. Whether that is understanding philosophically does not matter operationally. What matters is whether its claims are true and traceable.

Does it write as well as our best people? For fluency and tone, often yes. For knowing what not to say, and for being right about the customer, only when the system around it supplies accurate memory and checks.

Should we wait for the next model? Waiting for a better model is waiting for the part everyone will rent. The memory, decisions, and governance you build now carry over to every model you use later.

For your AIThis chapter's concepts, patterns and checklists as structured data. Paste it into your assistant.
chapter: 5
title: "The New Era"
concepts:
  - name: Capability Stack
    definition: "Six layers AI now performs for personalization: read, reason, decide, write, act, learn; each matures and fails differently."
  - name: Cost of Understanding
    definition: "The cost to form a usable, specific account of one customer; historically a skilled human hour, now a model run whose cost is set by architecture."
  - name: Ships now / Lab demo / Forecast
    definition: "Labeling discipline for AI capability claims: deployed in production, demonstrated but not common, or projected from trends."
  - name: Letter to Conversation
    definition: "Personalization shifts from a static message judged once to an interactive exchange judged at every turn."
  - name: Time Horizon
    definition: "Length of task (in skilled-human time) an AI completes at a given success rate; METR measures it mainly on software tasks."
  - name: Cost per Decision
    definition: "Total model, tool, and data cost divided by personalized decisions made; the unit to manage instead of price per token."
decision_rules:
  - if: "a claim about a customer is an inference without corroboration"
    then: "phrase it as a hypothesis or question, or fall back to an honestly general statement"
  - if: "a capability is cited in planning"
    then: "label it ships-now, lab demo, or forecast, with a date"
  - if: "research about an entity already exists within its freshness window"
    then: "reuse it; do not re-research per record"
  - if: "an action is irreversible or evidence is below threshold"
    then: "route to a person"
  - if: "the AI bill rises while per-token prices fall"
    then: "measure tokens per decision and find repeated or oversized reads"
  - if: "an agent is granted tools to act"
    then: "limit them to its role and job, give the credential an expiry and the run a time cap, and withhold destructive operations"
  - if: "a model score and its label or band disagree"
    then: "flag the record in code before anything is sent"
assessment_questions:
  - "Which unstructured sources about customers exist, and does any system read them today?"
  - "What is the current cost per personalized decision, and who measures it?"
  - "Can the team trace any AI-stated customer fact to its source?"
  - "Which AI workflows run unattended, and what returns them to a person?"
  - "How many accounts receive human-depth research today, out of how many?"
patterns: [Evidence-Weighted Voice, Research Once per Entity, Human at the Boundary]
anti_patterns: [The Fluent Guess, Per-Token Illusion, Horizon-as-Schedule]
maturity_dimension: understanding
as_of: 2026-09-26

References

  1. Brynjolfsson, E., Li, D., Raymond, L. (2025). "Generative AI at Work." Quarterly Journal of Economics 140(2). https://academic.oup.com/qje/article/140/2/889/7990658
  2. Jones, C. R., Bergen, B. K. (2025). "Large Language Models Pass the Turing Test." arXiv:2503.23674 (preprint). https://arxiv.org/abs/2503.23674
  3. Hartmann, J., Exner, Y., Domdey, S. (2025). "The power of generative marketing." International Journal of Research in Marketing 42(1). https://www.sciencedirect.com/science/article/pii/S0167811624000843
  4. Salvi, F., Horta Ribeiro, M., Gallotti, R., West, R. (2025). "On the conversational persuasiveness of GPT-4." Nature Human Behaviour 9:1645-1653. https://pmc.ncbi.nlm.nih.gov/articles/PMC12367540/ ; Author Correction, 2026-09-03: https://www.nature.com/articles/s41562-026-02588-0
  5. Stanford HAI (2025). AI Index Report 2025. https://hai.stanford.edu/ai-index/2025-ai-index-report
  6. Epoch AI (2025-03-12). "LLM inference prices have fallen rapidly but unequally across tasks." https://epoch.ai/data-insights/llm-inference-price-trends
  7. Anthropic (2025-06). "How we built our multi-agent research system." https://www.anthropic.com/engineering/multi-agent-research-system
  8. Appenzeller, G. (2024-11). "Welcome to LLMflation." a16z. https://a16z.com/llmflation-llm-inference-cost/
  9. METR (2025-03-19). "Measuring AI Ability to Complete Long Tasks." https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
  10. METR (2026-01-29). "Time Horizon 1.1." https://metr.org/blog/2026-1-29-time-horizon-1-1/ ; tracking page (updated 2026-05-08): https://metr.org/time-horizons/
  11. METR (2025-07-10). "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  12. Stanford HAI (2026-04-13). "12 Takeaways from the 2026 AI Index." https://hai.stanford.edu/news/inside-the-ai-index-12-takeaways-from-the-2026-report
  13. Kalai, A. T., Nachum, O., Vempala, S., Zhang, E. (2025). "Why Language Models Hallucinate." arXiv:2509.04664. https://arxiv.org/abs/2509.04664
  14. Vectara. Hallucination Leaderboard (HHEM), snapshot 2026-09-22. https://github.com/vectara/hallucination-leaderboard
  15. Chen, Y., Benton, J., et al. (2025). "Reasoning Models Don't Always Say What They Think." arXiv:2505.05410. https://arxiv.org/abs/2505.05410
  16. Matz, S. C., et al. (2024). "The potential of generative AI for personalized persuasion at scale." Scientific Reports 14:4692. https://www.nature.com/articles/s41598-024-53755-0 ; critique: https://www.pnas.org/doi/10.1073/pnas.2418817121
  17. H. Taheri, first-party system documentation: capabilities reference for a governed personalization engine (2026-08-25; only capabilities marked live are described as built) and release documentation for a self-hosted governed memory system (release 0.8.2, verified 2026-09-16); unpublished. [PRODUCT-NAMING: name the systems here if Hamed decides to]

This chapter is a working draft. If something is wrong or missing, tell me on LinkedIn.

Get chapters by email as they are revised

Prefer LinkedIn? Subscribe to the newsletter there instead.