Skip to main content
Layers

Productivity and collaboration

Office assistants and agents: every agent gets a registered identity, a mailbox or meeting seat is a rare separate grant, and rollout follows whether work is solitary or shared.

Target state

In short: Assistants help with solitary work everywhere, every agent has a registered identity, and a mailbox or meeting seat is rare.

Assistants run broadly across solitary work. Coordinated work changes only where teams deliberately agree new norms. Every agent carries an access identity: a service principal in the enterprise directory (the directory's account type for software rather than people) with a recorded human sponsor. It is issued automatically and governed like any workload identity, the platform-issued proof of which piece of software is acting. Presence (a mailbox, a calendar, a meeting-roster seat) is a separate, revocable, per-capability grant, and it is rare. Shadow usage (staff using personal AI accounts for work) is governed through identity and conditional access rather than through blocking. Conditional access means the directory's rules on who may reach what, and from where. The sanctioned path is made easier than the personal account. Interaction mode is configured per surface. Where diversity of ideas matters, the assistant draws out the author's own thinking with questions; it rewrites people's contributions only where diversity does not matter. Oversight controls are judged on measured detection accuracy, never on reviewer confidence. This layer is the highest-volume agent surface in the enterprise. It is where identity governance holds or fails first.

layer · architecture

Access identity by default, presence by exception

The agent identity is a service principal with the sponsor attribute; the user account carrying mailbox and meeting presence is a separate, optional, revocable grant.

  1. 01

    Agent

    Agent service principal

    Mandatory; sponsor attribute; non-revocable creation right, capped

  2. 02

    Control

    Connector permissions

    Surfaced as API permissions; conditional-access targetable

  3. 03

    Control

    Optional user account

    Separate 1:1 child object; admin-granted, revocable

  4. 04

    Boundary

    Presence surfaces

    Mailbox, calendar, licence, HR system, meeting roster

  5. 05

    Evidence

    Nameability

    @mentionable without a user account

Sponsor accountability is an access-identity attribute; presence is a deliberate, revocable decision.

Nameability is not presence; five capability surfaces define the presence tier.

Diagram description: Collaboration identity architecture from mandatory agent service principal with sponsor field, through optional one-to-one user account, to the five presence capability surfaces, with conditional-access-targetable connector permissions. The map contains Agent service principal: Mandatory; sponsor attribute; non-revocable creation right, capped; Connector permissions: Surfaced as API permissions; conditional-access targetable; Optional user account: Separate 1:1 child object; admin-granted, revocable; Presence surfaces: Mailbox, calendar, licence, HR system, meeting roster; Nameability: @mentionable without a user account. Its connections are principal to connector; principal to account for only when acting as a user; account to surfaces; principal to mention. Important boundary: Sponsor accountability is an access-identity attribute; presence is a deliberate, revocable decision.

Reviewed 2026-08-20Sources:r08-productivity-and-collaboration findings

Mechanisms

The identity spine: service principal first, user account by exception

In short: Every agent gets a directory account with a named sponsor; a mailbox is a separate, rare grant.

The agent identity is a service principal in the enterprise directory. It carries the accountability attribute directly: a sponsor field recording "the human user or group that's accountable for an agent". The same field is used, among other purposes, for "contacting a human in case a security incident happens". The object that carries a mailbox and meeting presence is different. It is a separate, optional one-to-one child user account, created only "for interactions where the agent needs to act as a user". The asymmetry between the two permissions is the control surface. Creating agent identities uses a permission that is "automatically granted... and can't be revoked", capped at 250 agent identities per blueprint (the template an agent is created from). Creating the child user account requires an admin-granted permission that can be revoked. From July 2026 every new agent must have an agent identity, with no opt-out, and no user account is created automatically. In the terms of the identity chain, access identity (tier ID2) is mandatory and presence (tier ID3) is a decision. Being nameable is not the same as presence. The display name appears in Teams, Outlook, and admin surfaces at the access tier, and agents can be @mentioned in Teams and Slack without any user account. The presence tier is defined by five capability surfaces: mailbox, calendar, licence consumption, human resources (HR) system participation, meeting-roster seat. Evidence: Microsoft Entra agent-identity documentation, 2025-2026 [vendor].

Shadow AI: govern through identity, not blocking

In short: Blocking AI apps pushes staff to personal accounts; governing through the directory keeps them on sanctioned paths.

Blocking moves usage rather than stopping it. Although 90 percent of organisations block at least one generative AI (genAI) app, 47 percent of genAI users reach tools through personal, unmanaged accounts. The measured exposure: an average organisation sends 18,000 prompts per month to genAI apps. It also records 223 incidents per month of sensitive data sent to AI apps (2,100 in the top quartile), and 54 percent of violations involve regulated data. GenAI users grew 200 percent year on year and prompt volume grew 500 percent. These figures are security-vendor telemetry (measured traffic, not a survey), Jan 2026 [vendor]. The governance mechanics converge on identity. Connector permissions are surfaced as application programming interface (API) permissions that conditional access can target. The policy engine that governs user sign-ins therefore also governs what an agent can reach. Agent identities are deleted automatically when the agent is deleted, so orphaned accounts (identities with no live agent behind them) do not accumulate. Mandatory issuance makes the agent inventory complete by construction: an agent cannot exist without an identity. The control point is the directory and its conditional-access policy set, not the proxy blocklist.

The solitary-vs-coordinated gate

In short: Measured gains appear in solitary tasks such as email; shared work does not change unless the team agrees new norms.

The only large randomised field experiment with telemetry outcomes is Dillon, Jaffe, Immorlica and Stanton, National Bureau of Economic Research (NBER) working paper 33795, revised November 2025. It covered 66 firms, 7,137 knowledge workers, and six months. Users spent two fewer hours per week on email and did less out-of-hours work. The study detected no other shift in the quantity or composition of work. Meetings showed no measurable effect: the bounds rule out anything outside -0.01 to +0.21 hours against a 5.22 hours-per-week mean. Document counts also showed no effect. The authors' explanation is that email is solitary, while shifting meetings or document ownership "requires coordinating with colleagues and agreeing on new norms".

UK government evaluations of the same product show the same pattern: the more rigorous the method, the smaller the effect. Self-reported savings with no comparison group came to 26 minutes per day. A quasi-experimental study (a comparison without random assignment) at the Department for Work and Pensions (DWP) found 19. A randomised study at HM Revenue and Customs (HMRC) found roughly 12. The telemetry-measured effect is confined to email alone. Danish administrative data (NBER working paper 33777) shows precisely estimated zero effects on earnings and hours. The statistical bounds rule out effects above 1 percent, beside 3 percent average time savings.

Observed-task results come from the UK Department for Business and Trade (DBT) evaluation, Aug 2025, with six people per group: small, but the direction is clear. The effect varies by task.

Adjusted hours saved per task ran from drafting 1.3 and summarising 0.8 down to presentations 0.0, image generation -0.5, and scheduling -0.6 (a net loss of time). In 15 percent of data-analysis uses and 36 percent of scheduling uses there was no quality assurance at all. The design consequence (the authors' position, drawn from the evidence above) is a routing gate. Route solitary, single-author work to the horizontal assistant by default. Route coordinated work there only with an explicit decision to change the team's norms attached, because without one the measured effect is null.

Oversight at realistic effort

In short: Reviewers miss most deliberate errors in agent output, so size controls for low review effort and judge them on detection accuracy.

In five-hour coding tasks with more than 100 participants, 94 percent of developers failed to detect deliberate sabotage inserted by an agent. When a safety monitor flagged the malice, 56 percent still accepted the malicious code (Ye and colleagues, arXiv 2606.05647, Jun 2026; preprint). A redesigned interface for reading agent traces (the step-by-step record of a run) cut the time to find errors and raised reported confidence. It did so without meaningfully improving final accuracy (three vendor-research studies, Feb 2026 [vendor]). That is the signature of a governance control that adds an over-reliance problem instead of removing one. The baseline is consistent. A meta-analysis (a study that pools other studies) registered in advance covered 106 studies and 370 effect sizes. It found human-AI combinations performing worse than the best of human or AI alone, with an effect size (Hedges g) of -0.23. The losses concentrated in decision tasks, and review is a decision task. Three design consequences follow. Assume low review effort when sizing controls. Measure oversight on detection accuracy, not on reviewer confidence or speed. Treat any control that raises confidence without raising accuracy as a regression.

Interaction mode as a control

In short: Whether the assistant rewrites people's ideas or draws them out with questions is a setting, and it changes idea diversity.

Whether assistance costs idea diversity is a configuration choice, not a property of the model. In a model-led mode, the large language model (LLM) rewrites people's contributions. It improved quality but reduced idea diversity and the authors' sense of ownership. In a reflective human-led mode, the LLM draws out elaboration through questions. It improved quality while preserving both (486 participants, with a validation study of 640). A related result: LLM rewriting systematically shifted the political values expressed in participants' comments (CHI 2026, the human-computer interaction conference, 465 participants). Default collaborative surfaces to reflective elicitation wherever diversity of contributions matters. Permit model-led rewriting where it does not. Record the mode as reviewable configuration rather than accepting the vendor default.

Design decisions

  • One universal assistant vs many specialised agents (CD-23): this is the wrong axis to decide on. Keep the architectural question separate: whether to run a single agent or an orchestrated multi-agent system is decided in CD-21, which weighs multi-agent orchestration against a single good loop. The deployment question turns on a different variable: whether the target work is solitary or coordinated, and whether the organisation will change the norms of the coordinated work. Horizontal copilots scale because they require no process change. They take 86 percent of horizontal application spend against 10 percent for agents, and they reach 64 percent weekly active users but only 1.14 actions per user per day. That is exactly why their measured value is confined to solitary work. Specialised agents stall on coordination. Only coding has broken out ($4.2 billion of $7.3 billion departmental spend), because that workflow was already tool-mediated and its feedback loop already automated. The evidenced third answer, argued by neither camp, is a long tail of narrow agents built by employees inside the horizontal platform. One firm generated roughly 15,000 of them after a company-wide rollout. Only 16 percent of enterprise deployments qualify as true agents.
  • Presence identity refinement (see the identity chain): accountability belongs at the access tier. That is where the sponsor field already lives, and requiring a mailbox in order to obtain accountability is backwards. Presence also does not buy the norm change it appears to buy. The only controlled study that varied invoked (called on demand) against ambient (always present) architecture is from AIES 2026 (the AI, Ethics and Society conference), with 157 participants. It found the change "did not substantially redistribute trust"; accountability stayed anchored to the human expert. The only randomised controlled trial (RCT) of an AI teammate as a participant compared 16 AI teams against 17 all-human teams (Jul 2026 preprint, student sample). It found the AI was the most talkative and least informative member. Human-to-human responsiveness, belonging, and status were all lower. Grant presence per capability surface, rarely, with a recorded justification.

Cross-cutting concerns

Evidence and limits

No Common Vulnerabilities and Exposures (CVE) records, the public catalogue of software security flaws, attach to this layer's surfaces in the research window. The incident record is legal and exposure-shaped. Two US class actions over meeting recording are active. One is the consolidated Otter.AI action in the Northern District of California. The other is the Granola action filed 30 July 2026 under the California Invasion of Privacy Act (CIPA), which carries $5,000 statutory damages per violation. A July 2026 survey of 500 US workers found 33.4 percent had encountered an AI notetaker. Among those, 34.7 percent were always asked permission and 25.1 percent never were. Across all workers, 18.8 percent had discovered a meeting was recorded without their knowledge. Notetaking is the highest-volume collaboration-agent function: 16.43 percent of all reported uses, and 110 million Meet attendees used automatic notes in a month, 8.5 times the year before [vendor]. No independent field study publishes production accuracy for enterprise meeting summaries. Field evaluations also record assistants surfacing files users should never have had access to. That is the oversharing debt carried over from layer R02, the data platform (content shared more widely than its owners intended), appearing at this layer first.

Evidence statuses to carry. The identity mechanics and the July 2026 mandate come from vendor documentation [vendor] and should be re-verified against tenant behaviour. Shadow-usage figures are security-vendor telemetry [vendor]. Spend composition comes from a single market survey (Dec 2025). The oversight-sabotage study and the AI-teammate RCT are preprints, the latter with a student sample. Observed-task figures rest on six people per group. One refusal: a widely circulated share of companies said to have delayed or cancelled copilot rollouts traces to a vendor security survey whose stated driver is data-protection risk. This guide excludes it as evidence about worker consultation. Open gaps: there is no comparative outcome study of presence-grade versus invoked agents at organisational scale, and no rigorous study of how notetakers change what participants say. Re-verify quarterly: the identity mandate's scope and caps, conditional-access coverage of connector permissions, and the shadow-telemetry baseline.

The research behind this page

On this page