Skip to main content
Layers

Experience and channels

Customer-facing agents get their own edge for channels, hand-offs, disclosure, and defence, while knowledge, identity, tools, testing, monitoring, and model access stay shared.

Target state

In short: Customer-facing agents get their own front edge, but everything behind it is shared with the rest of the company's agents.

A separate customer-facing edge sits on a shared control plane. The edge is the part of the system that faces the customer. The control plane is the common set of knowledge, identity, tools, tests, monitoring, and model access that every agent in the estate uses. The edge owns what has no internal equivalent. It runs the channel and telephony infrastructure. It measures containment (conversations that never reach a person) and escalation, wired to workforce management. It handles disclosure and consent. It carries brand-safety guardrails hardened against anonymous adversaries, and it keeps the legal-attribution and evidentiary record. Everything else is shared with the rest of the estate: the knowledge corpus, customer identity data, the tool and action layer, evaluation harnesses (the test suites that check an agent), observability, and model access. Duplicating any of them is a defect, not a safeguard. Every statement the agent makes is corporate communication, legally attributable to the operator. The agent grounds its answers in a governed knowledge system, answering from approved company sources it can point to. It refuses rather than improvises on policy. It discloses its nature at first interaction. It escalates to a real human queue and hands over the working state of the conversation. Resolution and repeat contact are the paired targets. Containment is an outcome, never the target. The edge is becoming two-sided. Customer agents answer inbound contact, while the enterprise's own pages and checkout must be readable by the customer's agent.

layer · architecture

A separate edge on a shared control plane

Channels, disclosure, adversarial hardening, and the legal-evidence layer are edge-specific; knowledge, identity, tools, evals, and observability are shared, and duplicating them is a defect.

  1. 01

    System

    Channel edge

    Web, voice, telephony; barge-in and latency budgets

  2. 02

    Control

    Disclosure and consent

    At first interaction; Article 50 live

  3. 03

    Boundary

    Adversarial hardening

    Anonymous users; public injection surface

  4. 04

    Evidence

    Legal evidence layer

    Verbatim records; statements bind the company

  5. 05

    Control

    Shared control plane

    Knowledge, identity, tools, evals, observability, models

  6. 06

    Human

    Staffed escalation queue

    Working-state transfer; never manufactured containment

The separation is justified by legal attribution, not technology.

Diagram description: Customer experience architecture with the channel edge (telephony, disclosure, guardrails, legal evidence) distinct from the shared control plane (knowledge corpus, identity, tool layer, evaluation, observability, model access). The map contains Channel edge: Web, voice, telephony; barge-in and latency budgets; Disclosure and consent: At first interaction; Article 50 live; Adversarial hardening: Anonymous users; public injection surface; Legal evidence layer: Verbatim records; statements bind the company; Shared control plane: Knowledge, identity, tools, evals, observability, models; Staffed escalation queue: Working-state transfer; never manufactured containment. Its connections are channels to plane; disclosure to channels; guard to channels; channels to evidence for every conversation; channels to human for escalation.

Reviewed 2026-08-20Sources:r09-experience-and-channels findings

Mechanisms

The separate edge, itemized

In short: Build separately only what customers uniquely need; share everything else, because the legal exposure, not the technology, justifies the split.

Five things are edge-specific because no internal equivalent exists. First, telephony acoustics and barge-in (a caller interrupting the agent mid-sentence). Second, containment and escalation instrumentation wired to workforce management. Third, disclosure and consent mechanics. Fourth, guardrails built for anonymous, incentivised adversaries. Internal agents face authenticated employees bound by an acceptable-use policy. A customer-facing agent faces anyone: a Chevrolet dealership assistant was manipulated into agreeing to sell a vehicle for $1 and into asserting the offer was legally binding. Fifth, the evidentiary layer, because transcripts are legal evidence. Six things are shared because duplication is a defect: the knowledge corpus, customer identity data, the tool and action layer, evaluation harnesses, observability, and model access. Knowledge is the layer you least want duplicated. Gartner opened an inaugural Customer Service Knowledge Management Magic Quadrant (5 Aug 2026) because knowledge quality is make-or-break.

Market structure supports the split. Gartner tracks four distinct customer experience (CX) markets: contact centre as a service (CCaaS, Sep 2025); Conversational AI (Jul 2026); customer relationship management (CRM) Customer Engagement Center; and customer service knowledge management (Aug 2026). Their vendor rosters barely overlap with internal-agent platforms. Microsoft and ServiceNow appear in none of the 14 Conversational AI or 9 CCaaS positions. Microsoft is pushing the other way: Service Agent reached general availability (GA) in Microsoft 365 (M365) Copilot on 15 Jul 2026, unifying customer service across Dynamics and M365.

The justification for the separate lane is legal, not technical. In Moffatt v Air Canada (2024 BCCRT 149), the British Columbia Civil Resolution Tribunal rejected the argument that a chatbot is a separate legal entity, calling it "a remarkable submission". It found negligent misrepresentation. The defence that the correct information was available elsewhere failed. The Higher Regional Court of Hamm (OLG Hamm, 12 May 2026) held that a chatbot is part of corporate communication, not a third party. Its statements are attributed to the operator like employee statements; liability attaches regardless of where the training data came from; and general disclaimers do not provide sufficient protection. The Munich Regional Court I (LG Muenchen I, 28 May 2026) extended operator liability to AI Overviews about third parties. Customer-facing failures are attributable, contractually binding, and public. Internal failures are contained and recoverable.

Voice architecture

In short: Voice agents live or die on response delay and on handling interruptions; the famous failures were acoustics, not language.

The architecture split is the dividing line. A cascaded pipeline (speech-to-text, then a large language model, then text-to-speech, or STT-LLM-TTS) adds 1 to 3 seconds per turn. Speech-to-speech collapses that delay. Latency bands in circulation among practitioners: under 500 milliseconds reads as natural, 500 to 1,000 milliseconds is acceptable, and over 1,000 milliseconds reads as broken. Treat these as thresholds in circulation, not measured distributions. Barge-in (a caller interrupting the agent while it is speaking) is the hard problem, with no text or internal equivalent. It decomposes into three controls that must agree under telephony constraints. Voice activity detection (VAD) sensitivity decides whether a sound is an interruption. Speech-to-text confidence decides whether it was speech worth acting on. Text-to-speech cancellation stops playback cleanly mid-utterance. The famous failure was acoustics, not language. McDonald's ended its IBM drive-thru partnership in July 2024 over a beamforming and acoustics problem (beamforming: using several microphones to pick out one voice). Taco Bell slowed its rollout after viral failures while reporting more than 2 million successful orders across more than 500 locations [vendor]. That is the signature of a system that works on the cases it was built for and fails visibly outside them. Demand is pull rather than push. In one survey, 97 percent of 400 leaders already use voice technology and 80 percent run some voice agent. Only 21 percent are very satisfied (Deepgram and Opus Research) [vendor]. No independent production accuracy figures for voice agents exist. Published latency distributions come from small vendor samples.

Advertised versus achieved

In short: The vendor's own published test scores far below the vendor's own marketing, and multi-turn conversations are the honest test.

CRMArena-Pro (arXiv 2505.18878, May 2025, 4,280 task instances) was published by Salesforce's own research arm, and it reads against the company's own marketing. The best model scores 58.3 percent on single-turn business-to-consumer (B2C) tasks and falls to 30.0 to 35.1 percent on multi-turn tasks. Knowledge question answering lands around 23 to 36 percent. Confidentiality awareness is 0 to 0.4 percent refusal under standard prompts. It rises to 24 to 62 percent with hardened prompts, at a cost of 1 to 11 points of task success. Contrast the same vendor's production self-report: 84 percent resolution across 380,000 conversations, support headcount down from 9,000 to about 5,000, and costs down 17 percent (Sept 2025) [vendor]. The task mixes and the resolution definitions differ. This guide reports both figures and the gap between them. Two design consequences follow. Multi-turn is the honest benchmark condition, because single-turn overstates capability by roughly half. Confidentiality behaviour is not free: it is bought with hardened prompts at a measured cost in task success, and that cost belongs in the budget.

Measurement architecture

In short: Count problems solved and customers who had to come back, not conversations that simply ended without a human.

False containment is the named metric failure. A session that ends without escalation scores as contained whether the customer was helped or gave up. The corrective is structural, not analytic. Instrument post-contact resolution and repeat-contact rate as a pair, and price cost per resolved issue rather than cost per contact. Circulating vendor pricing of $0.99 to $2.00 per automated resolution, against a $6 to $12 human comparator [vendor], means nothing when the denominator counts abandonments. Customer outcome rises or falls sharply on resolution. In one survey, 74 percent were satisfied with their most recent AI interaction, and above 90 percent when the issue was resolved without further steps. Net Promoter Score (NPS) fell by as much as 70 points on failure. In the US, full resolution followed a failed AI interaction only about half the time (COPC, Jan 2026, 1,000+ consumers, six countries). COPC names three primary causes of AI customer experience failure: measuring deflection rather than resolution, the people-simulation trap (optimising for sounding human), and weak data foundations. Escalation is where instrumentation is weakest. Zendesk's CX Trends 2026 survey covered 11,000+ respondents in 22 countries [vendor]. In it, 52 percent of Chinese respondents experienced context loss during escalation (having to start again after the handover). Only 20 percent of Australians called the handover seamless. And 74 percent find retelling their story frustrating. Preference is barely moving. In Metrigy's survey (503 respondents, fielded Nov 2025), 84.9 percent of consumers prefer human agents. Even when assured of equal resolution, 80.1 percent still do. Only 13 percent prefer AI, up from 11.6 percent. Escalation with working-state transfer is therefore launch scope, not a fast follow. No credible independent escalation-rate benchmark exists as of mid-2026. This guide states the absence as a finding rather than repeating circulating figures.

Disclosure law as a design input

In short: Customers must be told they are talking to an AI; the law now settles whether, so design decides how.

EU AI Act Article 50(1) became enforceable on 2 August 2026. The omnibus that delayed the high-risk regime did not defer it. The duty: inform people that they are interacting with an AI system, at the latest at first interaction, in a clear, distinguishable, and accessibility-compliant manner. The "unless obvious" exception is judged against an average, reasonably well-informed person and is to be interpreted restrictively (Commission guidelines, 20 July 2026). Penalties reach EUR 15 million or 3 percent of worldwide turnover. The obligation applies outside the EU wherever the outputs are used in the EU. Article 50(2), on marking synthetic content, has a grace period to 2 December 2026. The US is a patchwork with three different triggers. California Senate Bill (SB) 1001 (2019) is triggered by commerce. The Utah Artificial Intelligence Policy Act (AIPA, 2024) is triggered by a request and by sector. Colorado House Bill (HB) 26-1263, the Chatbot Safety Act (signed 29 May 2026, effective 1 January 2027), is triggered by status and adds duties. The law forecloses a paradox. A field experiment with more than 6,200 customers found that up-front bot disclosure cut purchase rates by more than 79.7 percent. Undisclosed bots matched proficient human agents (Luo, Tong, Fang and Qu, Marketing Science 38(6), 2019). Yet customers who knew they were interacting with AI reported satisfaction 34 percentage points higher (COPC, Jan 2026). Disclosure destroys persuasion and improves service. Article 50 removes the choice in the EU, so the live design question is how to disclose, not whether. The clearest published control that worked came from an incident. Cursor's support bot (April 2025) invented a single-device login policy to explain logouts that were actually caused by a race condition (a timing bug). Its responses were not labelled as AI, and developers cancelled within hours. The remediation fixed the bug and labelled all AI responses in email support. That adopted the disclosure control fourteen months before Article 50 made it mandatory.

Agentic commerce

In short: Customers' shopping agents now read company pages and buy through new protocols, and the pages that matter most are the least machine-readable.

Three protocols divide the labour. The Agentic Commerce Protocol (ACP; OpenAI and Stripe, Sep 2025; specification version 2026-04-17, beta) covers agent-driven checkout. The Universal Commerce Protocol (UCP) covers the merchant surface. Google and Shopify announced it at the National Retail Federation (NRF) event on 11 January 2026, with Etsy, Wayfair, Target, and Walmart. The Agent Payments Protocol (AP2) covers payment authorisation mandates, the signed spending permissions an agent cannot alter. Machine-readability is the binding constraint. Adobe's first-quarter (Q1) 2026 scorecard puts product detail pages last at 66 percent readable, below returns, contact, and FAQ pages at 80 percent or more [vendor]. The pages that matter most for conversion are the least legible to the customer's agent. Direction reversed within twelve months. AI-referred traffic converted 9 percent worse in March 2025 and 42 percent better in March 2026. Traffic was up 393 percent year on year and time on site up 48 percent (Adobe Analytics, more than 1 trillion visits) [vendor]. Salesforce reported AI and agents influencing $67 billion, 20 percent of global orders, across Cyber Week 2025 [vendor]. Access is contested and currently favours the agent. The Ninth Circuit overturned Amazon's injunction against Perplexity's Comet browser. It held Amazon unlikely to succeed on its claim under the Computer Fraud and Abuse Act (CFAA), because Amazon's own users, not Perplexity, accessed the platform through the agentic browser. Whether agentic access can lawfully be refused remains unsettled.

Design decisions

  • Containment-as-target versus resolution-as-target (CD-20): set a resolution target and let containment be an outcome. The AI-first versus human-first framing, which CD-20 examines, is the wrong axis. Every documented reversal set a containment or headcount target and found that resolution did not follow. Klarna reversed in May 2025 after automating 67 percent of chats (2.3 million in 30 days, the work of 700 full-time equivalent (FTE) staff [vendor]). Its chief executive conceded that cost-cutting had been overweighted. Commonwealth Bank cut 45 roles, citing a voice bot that reduced calls by 2,000 per week. It then admitted volumes were rising, called the redundancy an "error" (21 Aug 2025), and offered the roles back. McDonald's ended the IBM drive-thru partnership in July 2024. Gartner's forecasts converge from four directions (reached via relay: second-hand reports rather than the primary releases). It forecasts 80 percent autonomous resolution of common issues by 2029. It also expects 50 percent of organisations that planned service-workforce cuts to abandon those plans by 2027, and no Fortune 500 company to fully eliminate human customer service by 2028. And it expects half of the staff cutters who cited AI to be rehiring by 2027, while only about a fifth had actually cut staffing. The strongest independent causal evidence sits on the assist side (Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 140(2), May 2025). Across 5,172 support agents, it found 15 percent more issues resolved per hour, the largest gains for novices, and better customer sentiment and agent retention. Nothing on the deflect side has comparable standing.
  • Separate stack versus shared platform for customer-facing agents (CD-24): a separate lane and edge on a shared control plane. The lane is justified legally, not technically. Customer-facing failures are legally attributable, contractually binding, and publicly visible. That warrants different governance, funding gates, and risk appetite, whichever platform the agent runs on. Honest caveat: no published enterprise case measures a deliberately unified customer and internal stack at scale, and none attributes success to deliberate separation either. The verdict rests on market structure and risk profile, not on measured architecture outcomes.

Cross-cutting concerns

Evidence and limits

The incident record is dated. The earlier entries: the Chevrolet $1-vehicle manipulation; Moffatt v Air Canada, decided February 2024 (2024 BCCRT 149); the McDonald's and IBM drive-thru termination, July 2024; the Cursor support-bot incident and remediation, April 2025. The later entries: the Klarna reversal, May 2025; the Commonwealth Bank reversal, August 2025; OLG Hamm, 12 May 2026; LG Muenchen I, 28 May 2026. No Common Vulnerabilities and Exposures (CVE) records, the public catalogue of software security flaws, attach to this layer's record. Rollback is normal rather than exceptional. In a survey of 2,527 senior decision-makers across ten countries, 74 percent had already rolled back or shut down an AI customer-communications agent after a governance failure. The share rose to 81 percent among organisations with mature guardrails; maturity correlates with catching failures, not avoiding them. And 76 percent invest more in trust, security, and compliance than in AI development itself (Sinch, January to February 2026, vendor-commissioned, fielded by an independent institute).

Vendor-published figures are flagged [vendor] inline. They include the 84 percent resolution self-report, the Deepgram and Opus voice survey, the Taco Bell order counts, the Zendesk escalation figures, the Adobe traffic and readability series, the Cyber Week aggregate, and the per-resolution price ranges. Gartner figures reached this guide via relay because the primary press releases were unavailable at research time; confirm the wording before citing onward. The voice latency bands are practitioner thresholds in circulation, not measured distributions. Two readings are the authors' synthesis of the cited cases and studies: that the separate lane is justified legally rather than technically, and that disclosure destroys persuasion while improving service.

Refusals: a widely recycled pairing of two vendors' containment figures was traced to its attributed article, and the figures are not in it. This guide treats the pairing as fabricated in transit and excludes it. Also excluded as unverifiable: circulating figures for voice-agent penetration among large banks, year-on-year voice-deployment growth, an average containment rate, every escalation-rate benchmark, and a pilots-never-reach-production rate. Re-verify on a quarterly cycle: Article 50 enforcement practice and first penalties; the Colorado act's 1 January 2027 effective date; ACP and UCP specification status (both pre-1.0); the agentic-access line of cases after the Ninth Circuit ruling; and whether a credible independent escalation benchmark has appeared.

The research behind this page

On this page