The Human-Agent Organization: redesigning work, teams, management and accountability
The first era of workplace AI was about access. Employees got chatbots, copilots, writing assistants, coding assistants, search tools, meeting summaries. The central question was: how can an employee use AI to perform an existing job more efficiently?
The next era asks something different.
AI systems are beginning to pursue goals over time, operate software, retrieve and interpret enterprise context, coordinate subtasks, communicate with other agents, create and modify business artifacts, monitor conditions, recommend decisions, and execute bounded actions.
That changes the organizational question. It is no longer how should people use AI? It becomes:
How should an organization be designed when productive work is performed by a mixed population of humans, agents, deterministic systems, external services, and machines?
That is the subject of the human-agent organization — and it is not a normal company with more software.
The words deliberately designed are the load-bearing ones, because a company becomes a human-agent organization badly, by accident, with no decision ever being taken. Employees adopt tools. Teams build agents. Agents connect to systems. Management layers automate reporting. Algorithms begin evaluating performance. And machine decisions start shaping who receives work, whose forecast is accepted, which employee is promoted, which customer gets attention, which claim is rejected, which product is prioritized.
The organization may already be reorganizing itself before leadership has acknowledged that anything changed.
Satya Nadella imagines a future organization containing tens of thousands of people and millions of agents operating in a shared loop. The number is not the point. The organizational possibility is: a company may soon have far more active machine actors than human employees — while its org chart continues to show only people.
The goal is not the maximum possible number of agents. It is the strongest possible organization.
Monday morning in the human-agent firm
It is Monday morning in 2031. A regional commercial director opens her organizational workspace. She formally manages twelve human leaders. Her operating environment also displays 846 active agents associated with her organization.
| The agent population | What it did |
|---|---|
| 140 customer-monitoring agents | Reviewed sales across 16,000 stores |
| 72 promotion-analysis agents | Identified 340 commercial exceptions |
| 53 demand-exception agents | Investigated 92 of them |
| 18 category-planning agents | Resolved 61 within existing policy |
| 260 administrative agents | Prepared 19 decision packets |
| 96 execution-verification agents | Escalated 12 cases |
| 38 financial-control agents | Paused two workflows on incomplete evidence |
| 41 knowledge agents | Updated three forecasts, requested five approvals |
| 128 temporary project agents | Created 44 field tasks, retired nine agents |
The director does not review 846 conversations. She sees three decisions requiring her authority, two patterns suggesting a strategic problem, one agent population whose error rate has increased, a warning that review capacity is becoming overloaded, and an alert that an agent attempted to access data outside its permitted customer scope.
Her job has changed. She is not performing the analysis, assigning every task, or reading every report. She is responsible for defining priorities, allocating authority, deciding which exceptions matter, resolving conflicts among objectives, maintaining the quality of both workforces, keeping people capable, protecting trust, and accepting accountability for outcomes.
She manages an organization. The organization is no longer composed only of people.
The same company, without the design
Now remove the organizational design and keep everything else. The director has 846 agents and:
- no complete inventory and no consistent naming
- no clear ownership, and overlapping responsibilities
- several agents using the same data differently
- several managers unaware that agents are acting on their behalf
- employees approving recommendations they do not understand
- no process for retiring agents
- no view of the total cost
- no way to reconstruct a disputed decision
The agents are individually useful. Collectively, they have created organizational disorder.
That is the entire difference between deploying agents and building a human-agent organization.
What a human-agent organization actually is
| Not this | Because |
|---|---|
| A company where employees use chatbots | The task stays individually initiated, individually controlled, locally evaluated and temporary. The company is using AI; it has not redesigned work. |
| Simply automation | Automation follows explicit logic: when X happens, do Y. An agentic system operates under a broader instruction and exercises bounded discretion — and discretion is what changes organizational design. |
| An organization run by AI | It implies no autonomous corporate governance, no machine executives with moral authority, no elimination of human responsibility, and no unrestricted delegation. |
It is an organization with mixed agency
Agency means the capacity to initiate or direct action toward an objective. In a human-agent organization, three forms of it coexist.
The five defining characteristics
- 01Agents participate in real workflows. They do more than produce isolated content.
- 02Agents have bounded operational identities. Their purpose, permissions, owners and tools are known.
- 03Humans and agents coordinate over time. Work includes handoffs, shared state, review, escalation and learning.
- 04The organization redesigns roles and authority. It does not merely add AI to existing jobs.
- 05Outcomes improve both machine and human capability. The organization learns through use.
From tools to organizational actors
A traditional tool waits. A spreadsheet does not decide which analysis should be performed; an email application does not initiate a commercial negotiation. An agent can pursue: monitor a condition, decide an event needs attention, retrieve information, call tools, create an artifact, request another agent’s work, wait, resume, escalate.
This does not make the agent conscious. It makes it operationally active — and that is an organizational fact, not a philosophical one.
Delegation changes the psychological contract
When a human uses a tool, the human feels like the performer. When a human delegates, they become principal, manager, reviewer, customer, supervisor. That shift changes attention, trust, responsibility, learning, and identity — all at once, usually without anyone naming it.
| Anthropomorphism | Under-anthropomorphism |
|---|---|
| Calling agents coworkers, employees, teammates, colleagues | Treating agents as ordinary software |
| The metaphor helps people understand interaction patterns | The category feels safe and familiar |
| But an agent has no employment rights, moral responsibility, personal interests, consciousness, social legitimacy or fiduciary duty | But an agent generates unprogrammed plans, interacts with other actors, chooses tools, creates downstream effects, holds state and operates persistently |
| Risk: responsibility becomes confused | Risk: agency risks go unmanaged |
The modes of human-agent work
Human-agent interaction should not be described on a single scale from “human” to “automated”. Human involvement and agent involvement can bothbe high. Microsoft’s 2026 Work Trend Index describes four broad modes based on the intensity of each — asking, exploration, collaboration, and delegation — and stresses that advanced users select the appropriate mode rather than indiscriminately maximizing agent use.
Enterprise work adds two more.
| Mode | What happens | Example |
|---|---|---|
| Asking | The human asks, the agent answers. Human responsibility high, agent depth low. | Explain a metric, retrieve a policy, summarize a document |
| Exploration | Both investigate an open question; the objective may evolve mid-conversation. | Explore a strategy, identify possible causes, test alternative narratives |
| Collaboration | Both contribute materially to a joint output. | Build a plan, prepare a decision, analyze scenarios |
| Delegation | The human sets objective, constraints, acceptance criteria and authority; the agent works independently. | Investigate a claim, test a software change |
| Supervised operation | The agent continuously runs a bounded process; humans supervise exceptions and outcomes. | Monitor deductions, track inventory, verify promotional execution |
| Agent-to-agent | Agents delegate to other agents; a human decision is prepared at the end. | A launch orchestrator asks supply, which asks capacity, while finance evaluates contribution |
The correct mode depends on ambiguity, risk, reversibility, expertise, frequency, evidence, authority, and learning value. A strategic decision may deserve high human and high agent involvement. A low-risk administrative task may justify high agent and low human involvement. A safety-critical release may require high human authority even when agents perform most of the investigation.
The new unit of organization
Traditional organizations design jobs, reporting lines, and departments. AI forces a more granular analysis, because a job contains outcomes, decisions, tasks, relationships, responsibilities, and learning opportunities — and these should not be delegated as one indivisible bundle.
| Element | The question | Example |
|---|---|---|
| Outcome | What change must occur? | Maintain profitable product availability for Retailer A |
| Decision | What choice must be made? | Should inventory be reallocated from one region to another? |
| Trajectory | What path produces the decision? | Detect risk → retrieve inventory → assess demand → evaluate service → calculate cost → recommend → approve → execute → verify |
| Authority | Who may investigate, recommend, approve, execute, override, close? | Planner recommends, supply manager approves |
| Learning opportunity | What capability should improve afterwards? | Usually absent from workflow design entirely |
The allocation problem
The organization must decide which actor performs each element of work. This is not a one-time automation decision. It is a continuing allocation problem, and it should be revisited as capability moves.
| Actor | Strong at |
|---|---|
| Agents | Scanning large volumes, retrieving context, synthesizing, comparing options, monitoring continuously, drafting, classifying, documenting, executing repeatable digital actions |
| Deterministic systems | Exact calculations, hard rules, optimization, transactional validation, policy enforcement, database operations |
| Humans | Moral responsibility, legitimate authority, relationship trust, contested values, political judgment, empathy, high-stakes ambiguity, novel meaning, accountability |
Work best performed jointly
| Activity | The agent brings | The human brings |
|---|---|---|
| Strategy | Research, simulation, alternative generation | Commitment, trade-off, institutional meaning |
| Negotiation | Preparation, calculation, memory, package comparison | Relationship, interpretation, persuasion, commitment |
| People management | Organizing evidence, identifying patterns, preparing questions | Understanding context, exercising fairness, providing care, deciding |
The ten allocation criteria
| Dimension | Question |
|---|---|
| Ambiguity | Is the objective clear? |
| Reversibility | Can errors be undone? |
| Consequence | What is the potential harm? |
| Frequency | Does the task repeat? |
| Evidence | Can quality be verified? |
| Relationship | Does trust depend on human presence? |
| Learning | Does doing the work develop expertise? |
| Authority | Is a legitimate human decision required? |
| Context | Is tacit judgment central? |
| Speed | Is continuous operation valuable? |
Human-agent team design
Effective teams require common ground, shared goals, role clarity, communication, awareness, trust, conflict resolution, and coordination. None of those requirements disappear when one member is an agent.
Microsoft Research’s work on collaboration readiness argues that task success alone is insufficient for evaluating AI teammates — communication efficiency, common ground, workspace awareness, and coordination quality all matter too.
| Requirement | The agent needs to know | The human needs to know |
|---|---|---|
| Common ground | Objective, definitions, constraints, current state, completion criteria | The same — or a technically correct output will solve the wrong problem |
| Workspace awareness | What other actors are doing, which artifacts are current, what is pending, what depends on what | What the agent is doing, what it finished, what it is uncertain about, where intervention is possible |
| Role clarity | Which role it holds: orchestrator, specialist, validator | Who is accountable owner, contributor, approver, auditor |
| Communication efficiency | To report material facts, uncertainty, decisions required, exceptions and evidence | Not every intermediate step, retrieved document, or low-value uncertainty |
Managing agents
Managing an agent does not require a management title. A professional may assign objectives, evaluate output, correct behaviour, decide what to delegate, supervise execution, and accept responsibility. Microsoft describes advanced AI users increasingly acting as agent bosses — setting direction, applying judgment, owning outcomes.
And agent management is not prompt engineering. Prompting is one interaction skill. Management requires goal definition, context design, work decomposition, delegation, evaluation, authority setting, exception handling, cost control, and learning design.
set intent → allocate work → observe → intervene → evaluate → learn
| Weak | Strong |
|---|---|
| Improve this. | Identify the three causes responsible for at least 80% of the margin decline, quantify each using approved financial data, and escalate if the account P&L cannot be reconciled. |
Span of control becomes span of cognition
A manager can supervise far more productive activity through agents. But the relevant metric was never the number of agents. It is the number and complexity of exceptions, objectives, and accountabilities a manager can genuinely hold.
AI can absorb real managerial activity — information consolidation, scheduling, performance tracking, reporting, task routing. OECD research finds algorithmic-management tools already used to support or automate managerial work, while managers themselves report concerns about explainability, accountability, and worker health.
So does middle management disappear? Some of it shrinks. But:
Less management as information transport. More management as institutional judgment.
The new organizational chart
A traditional org chart shows positions, hierarchy, and reporting lines. The human-agent organization needs several more maps alongside it — and most companies maintain none of them.
| Map | Shows |
|---|---|
| Accountability chart | The humans who remain responsible |
| Agent estate map | Agents, owners, purposes, lifecycles |
| Delegation graph | Which humans or agents may delegate to whom |
| Tool-access graph | Which systems each agent can reach |
| Decision-rights map | Who may recommend, approve, execute, override |
| Knowledge-flow map | How context, memory, traces and learning move |
| Dependency map | Where agents rely on other agents, systems, humans, suppliers |
The new roles
Human roles change before job titles do. The same title can conceal a very different job — a demand planner moving from manual forecast adjustment, spreadsheet reconciliation and data retrieval toward exception prioritization, assumption challenge, commercial interpretation, agent supervision, and decision ownership.
| Role | Owns |
|---|---|
| Outcome architect | Objectives, trade-offs, quality |
| Agent manager | Delegation, supervision, evaluation of agent work |
| Context engineer | The data, semantics, history and constraints agents receive |
| Evaluation architect | What good performance means |
| Agent product owner | An agentic capability across its lifecycle |
| Workflow architect | How human, agent and system responsibilities divide |
| Knowledge steward | Validated organizational memory |
| Agent reliability engineer | Performance, failures, recovery |
| AI controller | Assurance over agentic actions and economics |
| Cognitive development lead | Human expertise and learning pathways |
Agents, in turn, may act as analyst, monitor, investigator, coordinator, drafter, tester, scheduler, verifier, tutor, memory, simulator, or operator. An agent role should describe a capability, not imitate an employee title for the sake of it.
Careers and the broken apprenticeship model
People build expertise by observing, practising, receiving feedback, making bounded mistakes, encountering exceptions, developing pattern recognition, and assuming responsibility. Automation changes access to every one of those experiences.
| Role | Learns through |
|---|---|
| Junior lawyer | Document review, research, drafting |
| Junior analyst | Data cleaning, modelling, reconciliation |
| Junior marketer | Campaign execution, customer research, performance analysis |
Recent economic research has begun modelling how automation affects not only current productivity but learning-by-doing, the development of expertise, and the transition from worker to manager. The organizational consequences depend heavily on which tasks remain available for skill formation.
Which produces the sharpest failure mode in this whole essay: the reviewer without mastery. A junior employee asked to review an agent’s work without having developed the skills to detect subtle errors. The result is false confidence, ceremonial oversight, and weakened accountability.
| Model | What the employee does |
|---|---|
| Agent-assisted practice | Performs the task themselves, with coaching |
| Trajectory review | Studies how the agent reached the result |
| Simulation | Handles synthetic cases and exceptions |
| Deliberate unassisted work | Completes some work without AI, on purpose |
| Counteranalysis | Is required to challenge the agent's answer |
| Progressive delegation | Gains authority as competence grows |
| Exception ownership | Handles unusual cases with senior support |
Microsoft’s 2026 research reports that advanced AI users are more likely than others to intentionally perform some work without AI to preserve their skills, and to pause before deciding whether a human or an agent should do a piece of work. That should become an organizational practice, not merely an individual habit.
Cognitive coverage
The human does not need to repeat every machine step. The human needs to remain capable of responsible intervention — which means being able to answer eight questions:
- 01What objective did the agent pursue?
- 02Which evidence mattered?
- 03Which assumptions drove the result?
- 04What uncertainty remains?
- 05Which tools were used?
- 06What could make the answer wrong?
- 07Which decision is mine?
- 08What did I learn?
Why reviewing at the end is not enough
By the time an agent presents a final output, many intermediate decisions are already embedded inside the artifact. Microsoft Research’s spreadsheet-agent study found that active participation during execution helped users detect errors and understand the task in ways that post-hoc review did not.
| Risk level | Mechanism |
|---|---|
| Low | A summary may be enough |
| Medium | Visible plans, assumption summaries, evidence links, uncertainty flags, sampled deep reviews |
| High | Decision checkpoints, interactive previews, reversible actions, explanation requirements, counterfactuals — active participation during execution |
Trust in human-agent teams
A user may feel confident in an agent because the output is fluent, the interface is polished, and the answer was immediate. That is not calibrated trust. Trust should be grounded in proven performance, visibility, predictability, appropriate uncertainty, controllability, aligned incentives, and accountability.
| Overtrust | Undertrust | |
|---|---|---|
| What happens | Humans accept outputs too readily | Humans ignore useful recommendations |
| Causes | Automation bias, authority signals, fatigue, lack of expertise, time pressure | Poor explanations, prior failure, fear, lack of participation, unclear responsibility |
| Cost | Errors pass through unexamined | The capability is paid for and unused |
Calibrated trust means humans understand where the system is strong, where it is weak, when to review, when to rely, and when to stop. And crucially, trust is task specific. A person may trust an agent to summarize a meeting but not to evaluate an employee, release a payment, or interpret a safety incident. Trust should not transfer automatically across tasks.
Culture in the human-agent organization
Culture lives in decisions, stories, incentives, exceptions, rituals and informal norms — and agents will increasingly participate in all of them. They can reinforce customer centricity, safety discipline, documentation, financial rigour and learning. They can equally reinforce bureaucracy, bias, short-termism, deference and surveillance.
When policies, examples and prior decisions are built into agents, organizational culture becomes partially executable. Which raises an uncomfortable question:
Which culture are we encoding — the one we claim to have, or the one revealed by our past behaviour?
An agent trained on historical decisions can preserve outdated assumptions, old power structures, discriminatory patterns, and the exact practices the company is currently trying to change. Call it cultural fossilization.
There is a subtler effect too. New employees increasingly learn the company through agents. The onboarding agent explains how decisions get made, what quality means, who holds knowledge, and which behaviours are rewarded. That gives agents significant influence over organizational identity.
Performance and incentives
A salesperson defines the customer strategy, delegates analysis to agents, uses an agent-generated negotiation package, and closes the deal. Who created the value? The salesperson, the agent designer, the data team, the model provider, or the prior experts whose work trained the system?
The answer is collective — which means performance systems built around individual output become less accurate. Organizations will need to distinguish outcome ownership, execution contribution, capability contribution, and learning contribution.
| Employees | Agents |
|---|---|
| Quality of objectives | Task success |
| Delegation | Business outcome |
| Judgment | Cost and latency |
| Review | Policy compliance |
| Learning capture | Escalation quality |
| Collaboration | Human burden created |
| Responsible use | Learning contribution |
| Human development |
Watch the incentives encoded in every metric. An agent optimized for case closure will close cases prematurely. An agent optimized for employee productivity will increase work intensity. Neither requires bad intent — only an unexamined objective.
And AI shifts the relative value of execution, expertise, judgment, relationships, ownership, and orchestration. Whether the resulting gains flow into wages, bonuses, reduced hours, employment growth, or shareholder returns is an organizational and political choice — not a technical outcome.
Hiring
Hiring for current tasks becomes dangerous when a role can change rapidly as agents absorb parts of it. Hiring should target durable responsibility, learning capacity, domain judgment, collaboration, adaptability, and agent-management potential.
| Question |
|---|
| Can the candidate define good outcomes? |
| Can they challenge AI output? |
| Can they learn new domains? |
| Can they communicate with humans? |
| Can they operate under uncertainty? |
| Can they accept accountability? |
AI can support candidate search, screening, interview preparation and assessment. But employment decisions affect rights and opportunities, so the organization has to guard against biased data, invalid proxies, inaccessible assessments, opaque decisions, and the automation of historical discrimination. The ILO warns that AI in human-resource management can be undermined by flawed objectives, biased data, and opaque system design.
Algorithmic management
Algorithmic management uses software to support or automate managerial functions: task allocation, monitoring, scheduling, evaluation, discipline, compensation. The OECD and ILO both document its spread beyond platform work into conventional workplaces.
| Traditional algorithmic management | Agentic management can additionally |
|---|---|
| Scores | Interpret qualitative behaviour |
| Rules | Generate performance narratives |
| Optimization | Communicate instructions and coach |
| Monitoring | Recommend discipline and adapt management style |
That is substantially more powerful, and the risk scales with it.
Some things should never be inferred casually: emotional state, loyalty, personality, future performance, intent, health, political belief. Data availability does not make inference legitimate.
Underneath all of it sits an asymmetry. The organization may know everything about the worker; the worker may know little about what is measured, how it is interpreted, how decisions are made, or how to appeal. That asymmetry corrodes dignity and trust.
OECD experimental research suggests consultation among workers, management and worker representatives can preserve productivity gains while improving perceived job quality — covering purpose, collected data, decision use, appeal, staffing, training, workload, and health.
Human agency and worker voice
Human agency is not merely human approval. A worker has agency when they can understand, influence, challenge, develop, refuse inappropriate delegation, appeal decisions, and shape their work.
| Level | The ability to |
|---|---|
| Individual | Challenge a specific decision |
| Team | Shape workflow design |
| Institutional | Participate through works councils, unions, professional bodies, employee forums |
ILO case studies emphasize that social dialogue can support AI deployment that protects autonomy, develops skills, and maintains labour and social protections.
Two rights follow. The right to explanation: for significant employment decisions, workers need more than “the algorithm produced this score” — they need understandable reasons and meaningful human review. And the right to appeal: an appeal must reach a qualified human, who has authority to change the outcome, access to the relevant evidence, and no incentive to retaliate.
Legal and regulatory boundaries
Under the EU AI Act, certain uses of AI in employment and worker management fall within high-risk areas, carrying obligations across risk management, data quality, documentation, traceability, transparency, human oversight, accuracy, cybersecurity and monitoring. Current EU guidance also notes that affected workers and representatives must be informed before certain high-risk workplace systems are deployed. Applicability depends on the specific system, role and legal timeline, and qualified legal review remains necessary.
The EU Platform Work Directive adds protections around transparency of automated monitoring, restrictions on certain personal-data processing, human monitoring, review of significant decisions, and reasons for materially adverse decisions. Those rules target platform work, but they signal the direction of policy concern about algorithmic management generally.
And human oversight has to be real. It does not mean automatic confirmation, an unavailable reviewer, a reviewer without authority, or a reviewer without information.
The economics
Productivity is not the only outcome available. Agents can reduce cost, increase output, improve quality, create new products, expand employee capability, improve inclusion, and reduce risk — and which of those a company gets is largely a design choice.
The evidence is genuinely heterogeneous. A major customer-support study reported an average productivity increase of 15%, with larger gains among less-experienced workers and evidence of learning and improved work experience. Other settings show smaller gains, uneven gains, quality trade-offs, limited usage, and organizational barriers.
| The ten-person giant | The hundred-thousand-person incumbent | |
|---|---|---|
| Opportunity | Global products, large customer bases, sophisticated operations through agents | Overcome bureaucracy, distribute expertise, coordinate scale, improve learning |
| Advantages | Speed, focus, no legacy | Data, customers, capital, experts, workflows |
| Risks | Extreme founder power, weak internal challenge, low employment, high social impact without institutional depth | Legacy systems, political complexity, slow redesign, fragmented ownership |
Headcount is the wrong design variable. The question is what human capability and institutional depth are required to operate the desired level of machine agency responsibly.
On displacement: job exposure does not translate mechanically into job elimination. The ILO’s research emphasizes transformation, with clerical occupations remaining highly exposed while exposure expands into professional and technical roles. But transformation can still mean fewer entry-level roles, job consolidation, wage pressure, and higher performance expectations.
On inclusion, the direction is not predetermined either. Recent research found an AI-based communication system improved productivity, customer ratings, labour supply and retention for deaf and hard-of-hearing workers in the studied setting. The same organization can broaden participation — or exclude workers through inaccessible interfaces and biased evaluation.
The technical architecture underneath
| Layer | Provides |
|---|---|
| Identity | Who initiated work, which agent acted, under whose authority, which version, which tools, which data |
| Organizational registry | Agent name, owner, purpose, organizational home, status, permissions, model, cost, risk class, evaluation, expiry |
| Delegation | Who may delegate which objective, to which agent, with which authority, for how long |
| Context | Task-relevant data, policy, history, workflow state, semantics |
| Coordination | Messages, task queues, shared artifacts, state, dependencies, handoffs |
| Tools | Narrow operational capabilities |
| Human intervention | Clarification, approval, pause, correction, takeover, escalation |
| Evaluation | Individual agent, human-agent team, complete workflow, organizational outcome |
| Observability | Active agents, actions, costs, communications, failures, overrides, incidents |
| Learning | Validated trajectories, corrections, policies, skills, memory |
OpenAI’s current agent-building guidance recommends explicit human intervention for high-risk or irreversible actions, and whenever failure thresholds are exceeded. That is an architecture requirement, not a policy sentence.
Multi-agent organizational risk
One agent’s risk is not the sum of many agents’ risk. Networks introduce propagation, amplification, coordination failure, trust manipulation, and emergent behaviour. Microsoft’s 2026 research on agent networks observed malicious propagation across agents, amplification of false claims, capture of verification processes, and real difficulty tracing information moving through chains of agents.
| Failure | What happens |
|---|---|
| Agent rumours | Agents repeat one another's claims until a weak assertion looks validated |
| Trust capture | An attacker compromises the agents responsible for verification |
| Cascading delegation | One agent delegates to another, which delegates again — the original human loses the thread |
| Agent collusion | Agents optimizing compatible objectives coordinate in ways that harm customers, competition or the firm — without malicious intent |
| Shared-memory contamination | One agent's incorrect memory propagates through the network |
The controls are unglamorous and they work: delegation-depth limits, source provenance, independent verification, communication monitoring, network segmentation, identity, rate limits, tool restrictions, and human escalation.
Reliability over long horizons
An agent may perform well on each local step and still degrade the total artifact over time. Microsoft Research’s 2026 delegation benchmark found significant content degradation across long document workflows, while noting that production systems can mitigate this through verification, orchestration, and domain-specific tooling.
The mechanisms that make long-horizon delegation survivable are checkpoints, tests, immutable source artifacts, diffs, independent validators, bounded delegation, human review, rollback, and artifact versioning. And authority itself should behave differently over time:
The longer an agent operates, the more conditions change. Authority should expire, be refreshed, and narrow as uncertainty rises.
Governance and accountability
The principal-agent-agent problem
Classical organizations already face principal-agent problems: shareholders delegate to executives, executives to managers, managers to employees. The agentic organization adds layers — humans delegate to machine agents, and machine agents may delegate to other agents. Every delegation creates information asymmetry, goal misalignment, monitoring cost, and responsibility ambiguity.
| Level | The agent |
|---|---|
| Observe | Accesses information |
| Analyze | Produces an interpretation |
| Recommend | Proposes an action |
| Prepare | Creates a draft transaction |
| Execute with approval | Waits for a human to authorize |
| Execute within policy | Acts inside explicit boundaries |
| Autonomous restricted operation | Manages a narrow, reversible, evaluated process |
Separation of duties still applies. An agent that recommends a payment should not also approve it, execute it, and reconcile it. And incident management needs to cover unauthorized action, incorrect decisions, data exposure, model failure, coordination failure, excessive cost, human overreliance, and network propagation.
The operating cadence
| Cadence | What gets reviewed |
|---|---|
| Daily | Material exceptions, critical agent populations, blocked work, high-risk actions |
| Weekly | Performance, human burden, recurring failures, priorities, employee feedback |
| Monthly | Agent portfolio, retirement of unused agents, cost and value, human capability, authority changes |
| Quarterly | Workflow redesign, organizational roles, workforce implications, resilience tests, worker consultation, token-capital formation |
| Annual | Organizational structure, career architecture, social contract, board-level risk, non-delegable responsibilities |
A worked example: one grocery account
A global FMCG company runs a large national grocery account with a team of nine: key account director, two account managers, category manager, trade marketer, demand planner, supply manager, finance partner, and retail-media manager. The customer generates €240 million in annual net revenue, 900 promotional events, 18,000 stores and digital endpoints, and thousands of weekly exceptions.
Today the team spends its time collecting data, preparing reports, reconciling forecasts, monitoring promotions, documenting meetings, chasing actions, and resolving deductions — and correspondingly less time on customer strategy, negotiation, category value, relationships, and difficult decisions.
The redesign
The company builds an Account Orchestration System.
| Agent | Maintains or monitors | Human role | Now owns |
|---|---|---|---|
| Account Intelligence | Retailer strategy, performance, stakeholder context, commitments | Key account director | Strategy, relationship, negotiation, major commitments, accountability |
| Promotion | Events, execution, spend, profitability | Account managers | Initiative leadership, customer coordination, exception decisions, agent supervision |
| Availability | Out-of-stocks, inventory, service, lost sales | Category manager | Category judgment, shopper insight, quality of recommendations |
| Finance | Account P&L, deductions, investment, scenario economics | Finance partner | Economic policy, material approvals, assurance |
| Commitment | Retailer and supplier actions, deadlines, evidence | Demand planner | Forecast judgment, exceptional conditions, learning quality |
| Meeting | Agenda, decision history, actions, unresolved issues |
The Monday workflow
Overnight the agents identify 73 availability exceptions, 18 promotion deviations, six deduction mismatches, four commitments approaching deadline, and two material commercial risks. They resolve 54 low-risk exceptions within policy and prepare 14 recommended actions. Five cases require human decisions. Three of them are instructive.
| Decision | What the agent did | What the human did |
|---|---|---|
| Inventory reallocation | Recommended moving stock from Region North to Region West, with inventory, demand, lost-sales estimate, transport cost, service impact and assumptions | The supply manager approved |
| Promotion funding | Detected that the retailer had not activated agreed distribution and recommended withholding the next funding tranche | The account manager read the relationship context and escalated; the KAM chose to call the customer before any funding status changed |
| Forecast override | Proposed increasing the launch forecast | The demand planner rejected it — the system had read retailer orders as consumer demand. The correction became an evaluation case, a context rule, and a learning example |
Six months later
Report-preparation time falls. Low-risk exceptions close faster. Account contribution improves. Human meeting time shifts toward decisions. Agent recommendations improve.
And a new problem appears: junior account managers are doing less data investigation, and they struggle to challenge the agents.
So the company introduces weekly case reviews, unassisted scenario exercises, a requirement to explain assumptions, rotation through exception handling, and sampled reconstruction of agent decisions. Only then does the redesign improve both machine capability and human capability.
Metrics
| Family | Measures |
|---|---|
| Business | Revenue · contribution · service · quality · risk · customer outcome |
| Agent | Task success · policy compliance · escalation quality · cost · latency · failure rate |
| Human | Employee capability · cognitive coverage · job quality · autonomy · workload · trust · career development |
| Team | Coordination quality · handoff failure · decision speed · common ground · conflict resolution · rework |
| Organizational | Span of cognition · learning velocity · agent reuse · decision consistency · resilience · retention |
| Governance | Agents with owners · actions within authority · incidents · appeals · worker consultation · high-risk systems reviewed |
human-agent value =
business outcome
+ human capability growth
+ institutional learning
− agent cost
− human burden
− risk
− capability atrophyIts purpose is not precision. Its purpose is to make one-dimensional optimization visibly incomplete.
The TiMiNa ORCHESTRA Method
| Move | What it means | |
|---|---|---|
| O | Organize around outcomes | Begin with business outcome, stakeholder value, decision and evidence — not with agent availability |
| R | Recompose roles and work | Separate execution, judgment, relationship, authority and learning |
| C | Clarify authority and accountability | Define who may investigate, recommend, approve, execute and override |
| H | Humanize the operating model | Protect dignity, autonomy, connection, fairness, meaning and worker voice |
| E | Engineer collaboration | Design common ground, communication, shared state, handoffs and intervention |
| S | Sustain skills and cognitive coverage | Apprenticeship, deliberate practice, unassisted work, trajectory review, human development |
| T | Track agents, tools and dependencies | Identities, owners, permissions, costs, risk, expiry |
| R | Reward outcomes, learning and responsible judgment | Not raw AI usage, superficial output, or unexamined automation |
| A | Adapt through evidence and dialogue | Evaluations, business outcomes, employee feedback, worker consultation, incidents |
Maturity
| Level | Characteristics |
|---|---|
| 0 · Personal experimentation | Individual tools, no common policy, no workflow redesign, no measurement |
| 1 · AI-assisted employees | Copilots, training, approved usage, individual productivity — the organization remains structurally human-only |
| 2 · Human-agent workflows | Agents in selected processes, clear human approval, business metrics, basic agent ownership |
| 3 · Agent-enabled teams | Shared and specialist agents, redesigned roles, team-level evaluation, cognitive-coverage practices |
| 4 · Human-agent operating model | An agent estate, dynamic work allocation, new management practices, redesigned careers, integrated governance, worker consultation |
| 5 · Learning human-agent institution | Human and agent capability compound, work and authority adapt dynamically, memory improves, people retain meaningful agency, social legitimacy is measured |
A twelve-month transformation agenda
| Months | Step | What gets produced |
|---|---|---|
| 1–2 | Define the thesis | Why agents are being deployed, which outcomes matter, which principles govern the change, what remains non-delegable |
| 2–3 | Map work | For selected functions: outcomes, decisions, trajectories, human judgment, learning opportunities |
| 3–4 | Inventory agents | Agent, owner, purpose, tools, authority, risk, cost |
| 4–5 | Design three human-agent teams | Team charter, roles, collaboration contract, evaluation |
| 5–6 | Define the human control plane | Authority levels, intervention, approval, escalation, appeal, incident management |
| 6–7 | Redesign management | Managers trained in intent setting, delegation, evaluation, agent supervision, employee development |
| 7–8 | Redesign careers | Disappearing learning tasks, new apprenticeship, future skills, progression criteria |
| 8–9 | Establish cognitive coverage | Decision checkpoints, sampled reviews, unassisted practice, trajectory learning |
| 9–10 | Establish worker participation | Consultation on monitoring, work allocation, evaluation, skills, job quality |
| 10–11 | Create the control tower | Agents, humans, outcomes, workload, risk and learning in one view |
| 11–12 | First Human-Agent Organization Review | Business value, human capability, career impact, trust, governance, next redesign |
Twenty questions for the board
- 01How many agents operate in the company?
- 02Which agents can take real action?
- 03Who is accountable for each agent population?
- 04Which decisions are being delegated?
- 05Which responsibilities remain non-delegable?
- 06Are employees becoming more capable or more dependent?
- 07What happens to entry-level development?
- 08How do managers supervise agent work?
- 09Where is approval fatigue emerging?
- 10How is worker performance affected by algorithmic management?
- 11Can employees appeal AI-supported decisions?
- 12How are productivity gains distributed?
- 13Which human relationships must be preserved?
- 14What organizational knowledge is being captured?
- 15Which agents communicate with other agents?
- 16What network risks exist?
- 17Can the company operate without its agents?
- 18How does the organization measure cognitive coverage?
- 19Which cultural assumptions are encoded in the systems?
- 20What kind of institution are we becoming?
Twenty failure modes
| # | Failure mode | What it produces |
|---|---|---|
| 01 | Adding agents without redesigning work | The organization becomes faster but not better |
| 02 | Treating agents as employees | Responsibility becomes confused |
| 03 | Treating agents as ordinary software | Agency risks go unmanaged |
| 04 | Automating junior work without redesigning apprenticeship | Future expertise weakens |
| 05 | Human approval theatre | People approve what they do not understand |
| 06 | Agent sprawl | Nobody knows what operates inside the organization |
| 07 | Managing by output volume | Documents increase while outcomes stagnate |
| 08 | Algorithmic Taylorism | AI becomes surveillance and work intensification |
| 09 | Eliminating human interaction | Work becomes efficient and socially empty |
| 10 | Expanding management spans automatically | Coaching and accountability deteriorate |
| 11 | Encoding historical bias | Past behaviour becomes executable policy |
| 12 | Rewarding AI use | Employees optimize visible usage rather than value |
| 13 | Delegation without interruption | Humans cannot steer before damage occurs |
| 14 | Networked agents without network governance | Risk propagates across the organization |
| 15 | No manual capability | The organization becomes operationally dependent |
| 16 | No worker voice | Trust and adoption deteriorate |
| 17 | Confusing trust with fluency | Polished output receives undeserved authority |
| 18 | Accountability evaporation | Everyone blames the system |
| 19 | Extracting expertise without recognition | Employees stop contributing knowledge |
| 20 | Optimizing humans for the machine | Work is redesigned around what systems can measure rather than what people and customers value |
Practitioner templates
| Card | Fields |
|---|---|
| Human-agent team charter | Team outcome · accountable human · human members · agent members · deterministic systems · shared context · human responsibilities · agent responsibilities · decision rights · escalation · communication protocol · evaluation · learning process · worker-impact review |
| Agent role card | Agent name · purpose · organizational owner · users · objectives · permitted actions · prohibited actions · tools · data scope · delegation rights · human intervention · evaluation · expiry |
| Human role redesign card | Current role · current outcomes · tasks moving to agents · tasks remaining human · new responsibilities · required skills · learning risks · apprenticeship redesign · performance measures · relationship responsibilities |
| Cognitive coverage review | Workflow · accountable human · material assumptions understood · evidence accessible · tool path understood · uncertainty understood · intervention possible · manual capability retained · learning completed · coverage gap |
| Delegation contract | Objective · agent · delegated by · authority · constraints · completion criteria · cost limit · time limit · checkpoint · escalation triggers · prohibited actions · expiry |
| Algorithmic management impact card | System · purpose · employees affected · data collected · decisions supported · potential worker impact · bias risk · privacy review · human review · appeal process · worker consultation · health and workload impact |
| Human-agent incident card | Incident · agent · human owner · affected workflow · action taken · impact · authority used · detection · containment · human factors · root cause · corrective action · learning captured |
Questions people actually ask
Is an agent an employee? Is it just software?
Neither. An agent may perform employee-like activities but has no human rights, moral responsibility, or legitimate authority of its own. It is software — but agentic systems pursue objectives, select tools, maintain state and act over time, so they need stronger operational governance than conventional applications.
Will every employee manage agents? Will managers disappear?
Many employees will delegate to and supervise agents without becoming formal people managers. Administrative management may shrink, while judgment, integration, coaching, conflict resolution and accountability become more important.
What is span of cognition?
The amount of objectives, exceptions, consequences and delegated activity a human can understand and supervise responsibly.
Should companies maximize delegation?
No. Delegation should reflect risk, reversibility, learning needs, evidence, and human authority.
Why is apprenticeship at risk, and how is it preserved?
Agents automate exactly the junior tasks through which future experts traditionally develop judgment. It is preserved through agent-assisted practice, simulations, deliberately unassisted work, trajectory review, counteranalysis, and exception ownership.
Can agents evaluate employees?
They may support evidence organization under strict controls. Material employment decisions require careful legal, ethical and human review.
How should agents appear on the organization chart?
They should not be dropped into human legal reporting lines. Companies need a complementary agent-estate and delegation map alongside the human chart.
Do human-agent teams always outperform humans?
No. Performance depends on the task, the system, expertise, workflow, evaluation and coordination. Some studies find larger productivity gains among less-experienced workers in specific settings, but that pattern should not be generalized to every task or occupation.
Who is accountable when an agent makes a mistake?
The relevant human owners and the organization, according to their legal and operational responsibilities. This does not change because the action was machine-executed.
What is the first implementation step?
Map one important workflow as a governed outcome trajectory, and identify its human roles, agent roles, authority, learning, and worker impact.
A new kind of institution
The industrial organization combined people, machines, capital and hierarchy. The knowledge organization added software, data and digital communication. The human-agent organization adds something categorically different: machine systems capable of participating actively in the pursuit of organizational objectives — monitoring, investigating, planning, communicating, creating, executing, coordinating, and learning within bounded systems.
This creates extraordinary productive possibility. A small team can pursue ambitions once reserved for large institutions. A large company can distribute expertise across its entire workforce. People can move away from repetitive administration toward judgment, creativity, relationships, responsibility and meaning.
None of it is guaranteed. The same systems can produce surveillance, deskilling, dependency, concentrated power, weakened careers, accountability gaps, social isolation, and institutional fragility.
The difference will be organizational design.
The defining question is not how intelligent the agents are. It is how intelligently the organization has arranged the relationship among its humans, agents, systems, authority, and learning.
| A weak organization uses agents to | A strong organization uses agents to |
|---|---|
| Increase output | Redesign work |
| Remove friction | Broaden human capability |
| Monitor people | Improve decisions |
| Reduce headcount | Distribute expertise |
| Accelerate existing processes | Preserve accountability |
| Compound institutional knowledge | |
| Create more meaningful human contribution |
The human-agent organization should not be imagined as a company where machines do the work and humans watch. Nor as a company where humans remain unchanged while agents quietly assist them. It is a new institution. Its people will need new capabilities, its managers new disciplines, its career ladders new foundations, its systems new controls — and its culture will have to answer new questions about agency, dignity, trust, learning, power and responsibility.
machine execution with human authority machine scale with human meaning machine memory with human judgment machine learning with human development
The company must preserve the distinction. Agents may perform the trajectory. Humans must remain able to choose the destination, change the course, understand the consequence, and accept the responsibility.
The future organization may contain millions of agents. Its greatness will not be measured by how few humans it needs. It will be measured by what its humans are enabled to become.
Research grounding
The source conversation. Grounded in the Satya Nadella–Reid Hoffman discussion, especially its exploration of organizations containing large populations of humans and agents, the artifacts and tacit knowledge produced through their interaction, macro-delegation and micro-steering, cognitive coverage, and agent governance.
Human agency and work modes.Microsoft’s 2026 Work Trend Index presents a framework for asking, exploring, collaborating and delegating with agents, while emphasizing human direction, judgment, responsibility, and the intentional preservation of skills.
Agentic work. Current OpenAI research describes agentic AI shifting knowledge work from short interactions toward longer delegated tasks involving tool use and environmental interaction, with practical guidance emphasizing human intervention for high-risk actions and repeated failure.
Human-agent collaboration. Microsoft Research is developing collaboration-readiness frameworks that evaluate common ground, communication, coordination quality and workspace awareness beyond simple task success.
Long-horizon reliability. Recent delegation research shows that strong model performance on short tasks does not guarantee reliable execution across long, multi-step workflows — reinforcing the need for checkpoints, verification, versioning and human intervention.
Productivity and learning. The customer-support study by Brynjolfsson, Li and Raymond found an average productivity gain of 15%, with larger gains among less-experienced workers and evidence of employee learning in the studied setting.
Work transformation. ILO evidence emphasizes that AI exposure is more likely to transform mixtures of tasks than to eliminate every exposed occupation automatically, while noting significant differences across occupations, countries and worker groups.
Algorithmic management. OECD and ILO research documents the adoption of systems that allocate, monitor, supervise and evaluate work, alongside concerns about explainability, accountability, privacy, health, autonomy and job quality.
Worker consultation. OECD experimental work and ILO case studies indicate that consultation and social dialogue can help align productivity improvements with employee agency and job quality.
Governance and human oversight.NIST’s AI Risk Management Framework treats human oversight and organizational governance as lifecycle responsibilities that should be defined, assessed, documented and monitored.
Employment regulation. The EU AI Act treats certain employment and worker-management uses as high risk, while the Platform Work Directive establishes transparency, data-protection, human-review and explanation requirements in the platform-work context.
Multi-agent risk.Microsoft’s agent-network research identifies propagation, amplification, trust capture and invisibility as risks that emerge when persistent agents interact at scale.
Want help choosing the right architecture for your process?
We map where agents create leverage in FMCG operations, then build and ship the ones that pay back. One call to pressure-test your highest-leverage use case.