Productivity and collaboration
Office assistants and agents: every agent gets a registered identity, a mailbox or meeting seat is a rare separate grant, and rollout follows whether work is solitary or shared.
Target state
In short: Assistants help with solitary work everywhere, every agent has a registered identity, and a mailbox or meeting seat is rare.
Assistants run broadly across solitary work. Coordinated work changes only where teams deliberately agree new norms. Every agent carries an access identity: a service principal in the enterprise directory (the directory's account type for software rather than people) with a recorded human sponsor. It is issued automatically and governed like any workload identity, the platform-issued proof of which piece of software is acting. Presence (a mailbox, a calendar, a meeting-roster seat) is a separate, revocable, per-capability grant, and it is rare. Shadow usage (staff using personal AI accounts for work) is governed through identity and conditional access rather than through blocking. Conditional access means the directory's rules on who may reach what, and from where. The sanctioned path is made easier than the personal account. Interaction mode is configured per surface. Where diversity of ideas matters, the assistant draws out the author's own thinking with questions; it rewrites people's contributions only where diversity does not matter. Oversight controls are judged on measured detection accuracy, never on reviewer confidence. This layer is the highest-volume agent surface in the enterprise. It is where identity governance holds or fails first.
layer · architecture
Access identity by default, presence by exception
The agent identity is a service principal with the sponsor attribute; the user account carrying mailbox and meeting presence is a separate, optional, revocable grant.
- 01Agent
Agent
Agent service principal
Mandatory; sponsor attribute; non-revocable creation right, capped
- 02Control
Control
Connector permissions
Surfaced as API permissions; conditional-access targetable
- 03Control
Control
Optional user account
Separate 1:1 child object; admin-granted, revocable
- 04Boundary
Boundary
Presence surfaces
Mailbox, calendar, licence, HR system, meeting roster
- 05Evidence
Evidence
Nameability
@mentionable without a user account
Sponsor accountability is an access-identity attribute; presence is a deliberate, revocable decision.
Diagram description: Collaboration identity architecture from mandatory agent service principal with sponsor field, through optional one-to-one user account, to the five presence capability surfaces, with conditional-access-targetable connector permissions. The map contains Agent service principal: Mandatory; sponsor attribute; non-revocable creation right, capped; Connector permissions: Surfaced as API permissions; conditional-access targetable; Optional user account: Separate 1:1 child object; admin-granted, revocable; Presence surfaces: Mailbox, calendar, licence, HR system, meeting roster; Nameability: @mentionable without a user account. Its connections are principal to connector; principal to account for only when acting as a user; account to surfaces; principal to mention. Important boundary: Sponsor accountability is an access-identity attribute; presence is a deliberate, revocable decision.
- Component
- Agent identity plane
- Responsibility
- One service principal per agent, sponsor recorded
- Control it hosts
- Mandatory issuance (no opt-out from July 2026); 250-per-blueprint creation cap; conditional access over connector permissions; deletion with the agent
- Where it runs
- Enterprise directory (Microsoft Entra agent identity and equivalents)
- Component
- Presence tier
- Responsibility
- Optional one-to-one child user account where an agent must act as a user
- Control it hosts
- Admin-granted, revocable creation permission; justification per capability surface
- Where it runs
- Directory plus collaboration suite
- Component
- Horizontal assistant estate
- Responsibility
- Solitary-task assistance: email, summarisation, drafting
- Control it hosts
- Usage telemetry (actions per user per day); task-type routing by measured effect, positive or negative
- Where it runs
- Suite seats across the tenant
- Component
- Agent-building surface
- Responsibility
- The employee-built long tail of narrow agents
- Control it hosts
- Blueprint caps; automatic identity issuance; connector permission scoping
- Where it runs
- Low-code builders inside the horizontal platform
- Component
- Meeting capture
- Responsibility
- Notes and transcripts over meetings
- Control it hosts
- Consent capture and disclosure by default; retention aligned to recording law
- Where it runs
- Meeting platforms and specialist notetakers
- Component
- Shadow-usage telemetry
- Responsibility
- Detect personal-account generative AI (genAI) traffic and sensitive data leaving the organisation
- Control it hosts
- Personal-account detection; data loss prevention (DLP) incidents routed back to sanctioned paths
- Where it runs
- Existing security service edge (SSE) and DLP stack
Mechanisms
The identity spine: service principal first, user account by exception
In short: Every agent gets a directory account with a named sponsor; a mailbox is a separate, rare grant.
The agent identity is a service principal in the enterprise directory. It carries the accountability attribute directly: a sponsor field recording "the human user or group that's accountable for an agent". The same field is used, among other purposes, for "contacting a human in case a security incident happens". The object that carries a mailbox and meeting presence is different. It is a separate, optional one-to-one child user account, created only "for interactions where the agent needs to act as a user". The asymmetry between the two permissions is the control surface. Creating agent identities uses a permission that is "automatically granted... and can't be revoked", capped at 250 agent identities per blueprint (the template an agent is created from). Creating the child user account requires an admin-granted permission that can be revoked. From July 2026 every new agent must have an agent identity, with no opt-out, and no user account is created automatically. In the terms of the identity chain, access identity (tier ID2) is mandatory and presence (tier ID3) is a decision. Being nameable is not the same as presence. The display name appears in Teams, Outlook, and admin surfaces at the access tier, and agents can be @mentioned in Teams and Slack without any user account. The presence tier is defined by five capability surfaces: mailbox, calendar, licence consumption, human resources (HR) system participation, meeting-roster seat. Evidence: Microsoft Entra agent-identity documentation, 2025-2026 [vendor].
Shadow AI: govern through identity, not blocking
In short: Blocking AI apps pushes staff to personal accounts; governing through the directory keeps them on sanctioned paths.
Blocking moves usage rather than stopping it. Although 90 percent of organisations block at least one generative AI (genAI) app, 47 percent of genAI users reach tools through personal, unmanaged accounts. The measured exposure: an average organisation sends 18,000 prompts per month to genAI apps. It also records 223 incidents per month of sensitive data sent to AI apps (2,100 in the top quartile), and 54 percent of violations involve regulated data. GenAI users grew 200 percent year on year and prompt volume grew 500 percent. These figures are security-vendor telemetry (measured traffic, not a survey), Jan 2026 [vendor]. The governance mechanics converge on identity. Connector permissions are surfaced as application programming interface (API) permissions that conditional access can target. The policy engine that governs user sign-ins therefore also governs what an agent can reach. Agent identities are deleted automatically when the agent is deleted, so orphaned accounts (identities with no live agent behind them) do not accumulate. Mandatory issuance makes the agent inventory complete by construction: an agent cannot exist without an identity. The control point is the directory and its conditional-access policy set, not the proxy blocklist.
The solitary-vs-coordinated gate
In short: Measured gains appear in solitary tasks such as email; shared work does not change unless the team agrees new norms.
The only large randomised field experiment with telemetry outcomes is Dillon, Jaffe, Immorlica and Stanton, National Bureau of Economic Research (NBER) working paper 33795, revised November 2025. It covered 66 firms, 7,137 knowledge workers, and six months. Users spent two fewer hours per week on email and did less out-of-hours work. The study detected no other shift in the quantity or composition of work. Meetings showed no measurable effect: the bounds rule out anything outside -0.01 to +0.21 hours against a 5.22 hours-per-week mean. Document counts also showed no effect. The authors' explanation is that email is solitary, while shifting meetings or document ownership "requires coordinating with colleagues and agreeing on new norms".
UK government evaluations of the same product show the same pattern: the more rigorous the method, the smaller the effect. Self-reported savings with no comparison group came to 26 minutes per day. A quasi-experimental study (a comparison without random assignment) at the Department for Work and Pensions (DWP) found 19. A randomised study at HM Revenue and Customs (HMRC) found roughly 12. The telemetry-measured effect is confined to email alone. Danish administrative data (NBER working paper 33777) shows precisely estimated zero effects on earnings and hours. The statistical bounds rule out effects above 1 percent, beside 3 percent average time savings.
Observed-task results come from the UK Department for Business and Trade (DBT) evaluation, Aug 2025, with six people per group: small, but the direction is clear. The effect varies by task.
- Task
- Summarising a report
- Time with vs without (minutes:seconds)
- 12:37 vs 41:34
- Accuracy with vs without
- 4.0 vs 2.5
- Task
- Writing email
- Time with vs without (minutes:seconds)
- About equal
- Accuracy with vs without
- Quality better
- Task
- Building a presentation
- Time with vs without (minutes:seconds)
- Faster
- Accuracy with vs without
- 1.5 vs 5.0
- Task
- Excel data analysis
- Time with vs without (minutes:seconds)
- 25:01 vs 20:33 (slower)
- Accuracy with vs without
- 1.5 vs 2.7 (worse)
Adjusted hours saved per task ran from drafting 1.3 and summarising 0.8 down to presentations 0.0, image generation -0.5, and scheduling -0.6 (a net loss of time). In 15 percent of data-analysis uses and 36 percent of scheduling uses there was no quality assurance at all. The design consequence (the authors' position, drawn from the evidence above) is a routing gate. Route solitary, single-author work to the horizontal assistant by default. Route coordinated work there only with an explicit decision to change the team's norms attached, because without one the measured effect is null.
Oversight at realistic effort
In short: Reviewers miss most deliberate errors in agent output, so size controls for low review effort and judge them on detection accuracy.
In five-hour coding tasks with more than 100 participants, 94 percent of developers failed to detect deliberate sabotage inserted by an agent. When a safety monitor flagged the malice, 56 percent still accepted the malicious code (Ye and colleagues, arXiv 2606.05647, Jun 2026; preprint). A redesigned interface for reading agent traces (the step-by-step record of a run) cut the time to find errors and raised reported confidence. It did so without meaningfully improving final accuracy (three vendor-research studies, Feb 2026 [vendor]). That is the signature of a governance control that adds an over-reliance problem instead of removing one. The baseline is consistent. A meta-analysis (a study that pools other studies) registered in advance covered 106 studies and 370 effect sizes. It found human-AI combinations performing worse than the best of human or AI alone, with an effect size (Hedges g) of -0.23. The losses concentrated in decision tasks, and review is a decision task. Three design consequences follow. Assume low review effort when sizing controls. Measure oversight on detection accuracy, not on reviewer confidence or speed. Treat any control that raises confidence without raising accuracy as a regression.
Interaction mode as a control
In short: Whether the assistant rewrites people's ideas or draws them out with questions is a setting, and it changes idea diversity.
Whether assistance costs idea diversity is a configuration choice, not a property of the model. In a model-led mode, the large language model (LLM) rewrites people's contributions. It improved quality but reduced idea diversity and the authors' sense of ownership. In a reflective human-led mode, the LLM draws out elaboration through questions. It improved quality while preserving both (486 participants, with a validation study of 640). A related result: LLM rewriting systematically shifted the political values expressed in participants' comments (CHI 2026, the human-computer interaction conference, 465 participants). Default collaborative surfaces to reflective elicitation wherever diversity of contributions matters. Permit model-led rewriting where it does not. Record the mode as reviewable configuration rather than accepting the vendor default.
Design decisions
- One universal assistant vs many specialised agents (CD-23): this is the wrong axis to decide on. Keep the architectural question separate: whether to run a single agent or an orchestrated multi-agent system is decided in CD-21, which weighs multi-agent orchestration against a single good loop. The deployment question turns on a different variable: whether the target work is solitary or coordinated, and whether the organisation will change the norms of the coordinated work. Horizontal copilots scale because they require no process change. They take 86 percent of horizontal application spend against 10 percent for agents, and they reach 64 percent weekly active users but only 1.14 actions per user per day. That is exactly why their measured value is confined to solitary work. Specialised agents stall on coordination. Only coding has broken out ($4.2 billion of $7.3 billion departmental spend), because that workflow was already tool-mediated and its feedback loop already automated. The evidenced third answer, argued by neither camp, is a long tail of narrow agents built by employees inside the horizontal platform. One firm generated roughly 15,000 of them after a company-wide rollout. Only 16 percent of enterprise deployments qualify as true agents.
- Presence identity refinement (see the identity chain): accountability belongs at the access tier. That is where the sponsor field already lives, and requiring a mailbox in order to obtain accountability is backwards. Presence also does not buy the norm change it appears to buy. The only controlled study that varied invoked (called on demand) against ambient (always present) architecture is from AIES 2026 (the AI, Ethics and Society conference), with 157 participants. It found the change "did not substantially redistribute trust"; accountability stayed anchored to the human expert. The only randomised controlled trial (RCT) of an AI teammate as a participant compared 16 AI teams against 17 all-human teams (Jul 2026 preprint, student sample). It found the AI was the most talkative and least informative member. Human-to-human responsiveness, belonging, and status were all lower. Grant presence per capability surface, rarely, with a recorded justification.
Cross-cutting concerns
- #
- C1
- Concern
- Identity and access
- Treatment at this layer
- Access identity mandatory and automatic; presence granted per capability with justification; sponsor recorded at the access tier
- #
- C2
- Concern
- Observability
- Treatment at this layer
- Actions per active user per day; task-type outcomes; shadow-usage indicators
- #
- C3
- Concern
- Traceability and audit
- Treatment at this layer
- Meeting artifacts kept with their consent records; retention aligned to recording law
- #
- C4
- Concern
- Grounding
- Treatment at this layer
- Permissions on the collaboration content (files, sites, chats) fixed before deployment, not after
- #
- C5
- Concern
- Impersonation
- Treatment at this layer
- Agent output labelled in shared surfaces; no agent presents as a colleague without disclosure
- #
- C6
- Concern
- Sovereignty
- Treatment at this layer
- Recording consent depends on jurisdiction; transcripts inherit the classification of their content
- #
- C7
- Concern
- Privacy
- Treatment at this layer
- Notetaker consent capture; worker-monitoring boundaries; personal-account leakage
- #
- C8
- Concern
- Safety and oversight
- Treatment at this layer
- Review effort assumed low; controls that raise confidence without raising accuracy treated as harmful
- #
- C9
- Concern
- Cost
- Treatment at this layer
- Usage measured against licence spend; tasks with negative returns identified and stopped
- #
- C10
- Concern
- Resilience
- Treatment at this layer
- Sanctioned paths easy enough to displace shadow usage; a defined degraded mode when assistants are unavailable
Evidence and limits
No Common Vulnerabilities and Exposures (CVE) records, the public catalogue of software security flaws, attach to this layer's surfaces in the research window. The incident record is legal and exposure-shaped. Two US class actions over meeting recording are active. One is the consolidated Otter.AI action in the Northern District of California. The other is the Granola action filed 30 July 2026 under the California Invasion of Privacy Act (CIPA), which carries $5,000 statutory damages per violation. A July 2026 survey of 500 US workers found 33.4 percent had encountered an AI notetaker. Among those, 34.7 percent were always asked permission and 25.1 percent never were. Across all workers, 18.8 percent had discovered a meeting was recorded without their knowledge. Notetaking is the highest-volume collaboration-agent function: 16.43 percent of all reported uses, and 110 million Meet attendees used automatic notes in a month, 8.5 times the year before [vendor]. No independent field study publishes production accuracy for enterprise meeting summaries. Field evaluations also record assistants surfacing files users should never have had access to. That is the oversharing debt carried over from layer R02, the data platform (content shared more widely than its owners intended), appearing at this layer first.
Evidence statuses to carry. The identity mechanics and the July 2026 mandate come from vendor documentation [vendor] and should be re-verified against tenant behaviour. Shadow-usage figures are security-vendor telemetry [vendor]. Spend composition comes from a single market survey (Dec 2025). The oversight-sabotage study and the AI-teammate RCT are preprints, the latter with a student sample. Observed-task figures rest on six people per group. One refusal: a widely circulated share of companies said to have delayed or cancelled copilot rollouts traces to a vendor security survey whose stated driver is data-protection risk. This guide excludes it as evidence about worker consultation. Open gaps: there is no comparative outcome study of presence-grade versus invoked agents at organisational scale, and no rigorous study of how notetakers change what participants say. Re-verify quarterly: the identity mandate's scope and caps, conditional-access coverage of connector permissions, and the shadow-telemetry baseline.
The research behind this page
Agent platform
Buy the runtime and build the harness: the loop around the model gets explicit stop rules and budgets, inspectable state, separate checking, tested release gates, and crash survival.
Experience and channels
Customer-facing agents get their own edge for channels, hand-offs, disclosure, and defence, while knowledge, identity, tools, testing, monitoring, and model access stay shared.