Dominic Chiappe · People, capability & transformation

Thinking about how organisations perform in an AI-enabled world

AyEye — Workforce Management ·

When everyone can build the workflow, everyone can change the organisation

Microsoft is putting natural-language app creation, persistent delegation, managed runtime and metered spend into one work surface. The workforce consequence is operational authorship: deciding who may turn local judgement into a persistent system that changes how other people work.

Editorial systems diagram showing an individual intention becoming a built and published work system that coordinates people and agents, with an explicit renewal or expiry loop
The consequential moment is not creation but publication: when a local judgement becomes a persistent environment in which other people and machines must work.

Workforce intelligence · Issue 11

AyEye — Workforce Management

Human systems in the agentic enterprise

Saturday, 26 September 2026

For most of office history, changing how work worked required a project, a budget and at least one meeting that should have been an email. Microsoft now says a person will be able to describe an app, tracker, dashboard or workflow in ordinary language, run it inside the company tenant and hand recurring work to a persistent agent with its own identity, memory, computer and workspace.

The intriguing question is not whether everybody becomes a developer. It is whether everybody becomes a small-scale organisation designer.


The executive brief

  • Microsoft has announced its largest Copilot redesign to date. A new Code surface is intended to let non-developers create purpose-built apps and workflows; Autopilot is a persistent agent that can continue work without a prompt; and a managed runtime can host what people build inside the company environment. Most elements are preview or forthcoming, not established outcomes.
  • Agentic work is acquiring its own cost-allocation system. Microsoft will combine fixed user licences with usage-based billing for long-running work, model controls for groups and APIs for spending policies. A budget is becoming part of an agent’s practical mandate.
  • Assurance is moving into the build loop. Microsoft’s open-source run-assert-eval skill attempts to discover risks, measure them, generate runtime policy and rerun the same evaluation. Its billing-support example reduced observed cross-customer violations, but it remains a worked example and human review is still required.
  • The UK Ministry of Justice offers a public-sector counterexample to casual creation. It reports more than 1.5 million probation meetings transcribed, but places scaling inside a Justice AI Unit, portfolio visibility, ethics rules, testing and ongoing monitoring.
  • Physical autonomy is showing the same separation between operation and oversight. Kodiak says a truck has completed the 219-mile Lancaster-to-Houston route without a human touching the wheel and plans unsupervised long-haul service by year-end. Intervention-free movement still depends on hubs, validation, fleet operations and a safety case.
  • The workforce issue is authoring power. When a local employee can turn an intention into a persistent system that routes work, spends money and changes other people’s environment, deployment authority becomes a form of management authority.

When everyone can build the workflow, everyone can change the organisation

ORIGINAL SYNTHESIS · Confidence: medium-high · Horizon: 6–24 months

The next wave of “citizen development” is not simply easier software creation. It is the distribution of operational authorship: the power to turn a local judgement into a repeatable rule that other people and machines encounter as work.

Signal one: software creation enters the ordinary work surface

REPORTED FACT. On 25 September, Microsoft announced a new Copilot organised around Home, Code and Autopilot. Code is intended to let a person describe an app, tracker, dashboard, automation or workflow in natural language. Microsoft says the system can choose an approach and build persistent widgets, interactive dashboards or cloud-hosted internal applications that can be shared with a team.

Code uses technology related to GitHub Copilot, runs in a sandbox and can be hosted inside the organisation’s tenant. Microsoft Copilot Managed Runtime, now in preview, is intended to let IT govern hosted applications while employees connect them to live data. Home and Code are due to begin rolling out through the Frontier programme; broader availability is promised later. These are product claims and plans, not evidence that ordinary employees will build dependable production systems. Microsoft announcement, 25 September

ANALYSIS. The unit of knowledge work may be expanding from a file to a functioning system. A spreadsheet describes a rota; a small application allocates it. A memo proposes an approval rule; a workflow applies it. A presentation recommends a customer sequence; an agent pursues it after the meeting ends.

That shift matters because software is not neutral stationery. It makes some actions easy, others difficult and still others invisible. The person who creates a local workflow is deciding what counts as a case, which evidence is requested, where an exception goes and when somebody else is interrupted.

Signal two: delegation becomes persistent and economically metered

REPORTED FACT. Microsoft’s renamed Autopilot—previously Scout—is described as a cloud-hosted agent with its own identity, memory, computer and workspace. A user gives it a name, role and goal; it can watch channels, follow up on threads, run recurring work and resume projects days later. Microsoft’s example is an agent managing a supplier-review process, including the schedule, meetings and stakeholder follow-up.

The same announcement separates everyday Copilot use from long-running agentic work. Quick assistance remains covered by a user subscription, while Cowork, Code, Autopilot and frontier models use usage-based billing. Administrators can set spending policies, route credit requests through approval workflows, restrict model families for groups and compare task outcomes with cost.

ANALYSIS. Identity says which machine actor is present. Memory gives it continuity. A runtime gives its creations somewhere to operate. Usage billing gives activity an economic boundary. Put together, those are not the attributes of a clever document tool. They are the beginnings of a local operating unit.

The budget is especially revealing. A person with permission to create a recurring agent is not only delegating labour. They are creating a stream of future expenditure and attention. Cost control therefore belongs beside authority, not in a monthly cloud bill discovered after the work has already changed.

Signal three: the control loop is also becoming generative

REPORTED FACT. Microsoft released an open-source skill called run-assert-eval on 24 September. From a prompt in Visual Studio Code, it combines threat modelling, evaluation, runtime-policy generation and a repeated test intended to determine whether the policy fixed the measured failure without destroying useful behaviour.

In Microsoft’s worked billing-support example, the baseline agent exposed another customer’s data in 12 of 40 applicable conversations. After a deterministic account-matching policy was added before and after tool calls, the governed run recorded two violations in 34 applicable conversations. Across the reported splits, permissible-behaviour violations fell to zero. The samples are small, the judge is automated and the results come from the tool’s maker. Microsoft explicitly says generated policy is not automatic approval: a person must review the policy, intervention point and wiring. Microsoft run-assert-eval, 24 September

ANALYSIS. This closes part of the distance between “somebody should check this” and an enforceable control. It also creates a recursion problem. If building, threat discovery, test design and policy drafting all become easier, the scarce activity becomes accepting the institutional consequence: deciding which risk matters, who may approve the control, what evidence is sufficient and who owns the system when its original creator moves on.

Signal four: scale still needs an institution

REPORTED FACT. The UK Ministry of Justice said on 24 September that its Justice Transcribe tool had processed more than 1.5 million probation meetings. It describes a Justice AI Unit, a chief AI officer, cross-functional steering and risk groups, an internal portfolio tracker and requirements for accuracy, bias, fairness, security and reliability testing before wide deployment, with ongoing monitoring. It also says future work will involve appropriate engagement, including trade unions where relevant. Ministry of Justice update, 24 September

ANALYSIS. The public sector example is not proof that every control works. It is a useful contrast. Ease of creation does not remove the need for a body that can see the portfolio, retain learning across projects and connect technical behaviour to professional duty, public legitimacy and employee voice.

The operational-authorship stack

LayerNewly easy activityOrganisational right at stakeEvidence that should survive
DescribeTurn an intention into a specificationWho may define the problem and affected group?Purpose, assumptions and exclusions
BuildGenerate an app, workflow or agentWho may encode a work rule?Creator, data, model, tests and version
RunHost it inside the enterpriseWho may make the rule operational?Users, permissions, dependencies and expiry
DelegateLet a persistent agent continue without promptsWho may initiate future acts?Objective, boundaries, exceptions and stop right
SpendConsume models, compute and attentionWho may commit recurring resources?Budget, outcome, unit cost and owner
AssureGenerate tests and draft controlsWho decides that evidence is good enough?Frozen test, human approval and residual failure
RetireEnd or replace the local systemWho protects people from orphaned rules?Dependencies, records, remedy and shutdown

What to notice: technical permission to create is only one layer. Operational authorship includes the rights to make a rule live, let it persist, fund its activity, certify its safety and remove it cleanly.

ORIGINAL SYNTHESIS. These developments point to a new management capability: operational authorship—the legitimate ability to translate local judgement into a persistent work system.

Organisations have distributed authorship before. Spreadsheet macros, low-code tools and departmental databases are old acquaintances, occasionally discovered under a desk like a startled hedgehog. What changes now is the combination. A single work surface can help a person design software, connect live enterprise context, assign a persistent agent, host the result, meter its expenditure and generate parts of its control system.

The resulting artefact is more than an app and less than a department. AyEye calls it a micro-institution: a small, executable arrangement of purpose, rules, resources and authority that can coordinate work beyond its author’s immediate attention.

Some micro-institutions will be splendidly modest: a tracker that keeps a handover from being forgotten. Others will influence hiring queues, supplier decisions, customer treatment or employee schedules. The design challenge is to preserve the first without pretending the second is merely personal productivity.

Unexpected connection

Natural-language software + persistent delegation + usage-based finance + automated assurance + public-sector portfolio governance

These look like separate product and policy developments. Together they create a production system for new ways of working. The employee supplies intent; a model generates structure; a runtime gives it persistence; finance grants fuel; assurance shapes its boundary; the organisation inherits the consequence.

The conclusion is surprising but practical: AI literacy is becoming less about knowing what to ask a model and more about knowing when an answer has become an institution.

PROVOCATION

Your organisation chart will matter less than your publishing permissions

A manager with five direct reports may soon be able to create systems that route the work of hundreds of colleagues, customers and agents. A junior specialist may encode an exception rule that travels further than an executive instruction. Formal hierarchy will still allocate accountability, pay and status; operational power will increasingly follow who may publish executable rules into shared work. If leaders govern titles but ignore publishing rights, the real organisation will grow beside the visible one.

What if we are right?

Opportunity. Frontline knowledge could become reusable capability without waiting for a central transformation programme. People closest to a recurring problem could build a small solution, test it with affected colleagues and improve it while the context is still fresh. The organisation could gain thousands of local experiments and retain the best as shared practice.

Organisational consequence. Job design would distinguish using an AI tool from authoring a work system. Identity, HCM, finance, architecture and risk would need a shared record of who may build, deploy, fund and retire micro-institutions. Recognition and progression would include the capability a person safely creates for others, not only the tasks they personally complete.

Likely horizon. Local apps and long-running agents are entering preview now. Within 6–12 months, early adopters will have enough employee-built systems to expose ownership and retirement problems. A portable work-authoring record across HCM, identity, portfolio and cost systems is plausible within 18–36 months.

What would prove us wrong?

The thesis weakens if Code-like systems remain novelty tools, if security constraints keep them isolated from consequential workflows, or if employees consistently build only personal utilities that disappear when a session ends. It also weakens if central platforms can automatically contain access, spend, quality and lifecycle without adding material human governance.

Microsoft’s announcement is a product roadmap, not usage evidence. The Ministry of Justice reports activity and governance structures, not independently verified service outcomes. Automated evaluation may reduce assurance work without becoming trustworthy enough for release decisions.

The concept fails if “operational authorship” becomes a grand title for every spreadsheet or an excuse to slow useful local improvement. The disconfirming test is straightforward: if employee-built applications and agents do not begin changing other people’s work, consuming persistent resources or surviving their creators, ordinary end-user computing controls may be enough.

Optimistic possibility: the organisation can learn at the edge

For decades, improvement has often travelled through a narrow pipe. A frontline worker spots a better route; a project queue translates it months later; context drains away between the two. Natural-language building can widen that pipe.

The constructive future is not uncontrolled shadow IT. It is legitimate local invention: people receive a bounded right to improve work, affected colleagues can challenge the design, evidence travels with the artefact and successful local practice can be adopted elsewhere. That could make organisational learning faster and more democratic—provided authorship earns responsibility without trapping the author as permanent unpaid support.


Citizen development is becoming citizen management

WORK DESIGN · IMPLEMENTATION SIGNAL

REPORTED FACT. Microsoft says Code will let people create shared internal apps from ordinary language, while specialised Skills in Word and Excel can distribute reusable role-specific workflows.

ANALYSIS. The phrase “citizen developer” frames the person as a software maker. “Citizen manager” is closer to the organisational consequence. The artefact can define a sequence, assign a queue, request evidence and escalate an exception. Those are management acts even when no manager performs them live.

Organisations should avoid solving this by classifying every builder as a manager. The better move is to recognise work-rule authorship as a capability with a scope. A finance analyst might be trusted to build tools for their own team but not to publish a credit-decision workflow. A service adviser might encode a handover checklist but not change who receives priority.

The important boundary is the affected population. Personal utility can be governed lightly. A system that changes another person’s options needs review proportionate to its reach and consequence.


A persistent agent creates a new kind of unfinished work

AGENT OPERATIONS · CONTINUITY

Autopilot’s appeal is that work continues while a person sleeps or turns to something else. That continuity also means decisions can accumulate while the original context changes.

ANALYSIS. The traditional handover has a natural social moment: one person tells another what remains open. A cloud agent may simply continue. The organisation therefore needs a machine handover state: what objective is active, which assumptions remain true, what the agent has promised, how much it has spent and what must be re-confirmed after silence, absence or organisational change.

Persistence should expire by default when the principal changes role, a supplier review ends or the affected team reorganises. Otherwise the enterprise will acquire zombie workflows: still technically authorised, still spending, still following goals that no current owner would write down.


FinOps is becoming part of workforce design

ECONOMICS · OPERATING MODEL

REPORTED FACT. Microsoft is introducing separate usage-based billing for long-running agentic work, group-level model controls, API-managed spending policies and outcome-versus-cost views.

ANALYSIS. Workforce planning has traditionally allocated people and money in different systems. Agentic work makes them inseparable at task level. A model choice changes cost, speed and quality; a persistent agent converts a managerial objective into recurring consumption; an employee can request more credits much as a manager requests overtime or contractor budget.

This is not a reason to treat agents as employees. It is a reason to connect machine expenditure to the human work system it alters. The meaningful unit is not tokens, seats or “digital workers”. It is the cost of achieving an objective, including human briefing, exception handling, assurance and repair.


An automatically generated control still needs a constitutional moment

ASSURANCE · HUMAN CONTROL

Run-assert-eval is constructive precisely because it does not stop at finding a failure. It tries to turn observed behaviour into an enforceable rule and retest the same cases.

ANALYSIS. The danger is not that machines help write policy. It is that policy generation makes a contested judgement look like plumbing. Choosing the critical risk, defining permissible behaviour and accepting residual failure are decisions about whose interests count.

The constitutional moment is the point at which a named human body says: this rule may now constrain work. Automation can supply evidence, draft the control and keep the comparison honest. It cannot determine legitimate trade-offs on behalf of people who were never represented.


Intervention-free does not mean institution-free

PHYSICAL AI · OUTSIDE-IN ANALYSIS

REPORTED FACT. Kodiak said on 25 September that its autonomous truck had completed the 219-mile route from its Lancaster, Texas hub to Houston without a human touching the wheel, including surface streets, and that it plans unsupervised long-haul service on the lane by the end of 2026. The company is still conducting daily preparation runs and closed-course validation. These are supplier claims and plans, not an independent safety finding. Kodiak announcement, 25 September

ANALYSIS. The absent driver is the visible change; the newly designed institution is the deeper one. A hub must prepare and release the vehicle. A safety case must justify operation. Remote, maintenance and recovery roles must exist even if they do not touch the wheel on a successful run.

The same principle applies to office agents. “No human intervention” describes a normal execution path. It says nothing about who created the path, validates it, absorbs exceptions or restores service. Workforce plans should count that surrounding institution before booking the saving.


The shadow organisation may be made of helpful little apps

TENUOUS BUT PLAUSIBLE · Confidence: medium-low · Horizon: 12–36 months

There is no evidence yet that Copilot Code has created a parallel management structure. The product is not broadly available. The causal chain is nevertheless credible:

Natural-language building reduces creation cost → more employees publish local workflows → some workflows persist and shape colleagues’ choices → operational influence follows artefacts rather than reporting lines → the application portfolio becomes a partial shadow organisation chart.

SPECULATION. Leaders may discover that the most influential person in a process is not its formal owner but the employee who maintains the small system everybody has learned to obey.

What to watch: whether employee-built apps gain shared users; whether they survive job moves; whether performance reviews reward their creation; whether colleagues can contest a rule; and whether HCM, identity and application inventories can identify the human steward without treating them as solely liable for every outcome.


Noise filter: “everyone can build” is not an operating model

Creation demos show the happy path from prompt to useful artefact. Enterprise value begins later: somebody must decide whether the problem deserved automation, whether the data is fit, whether affected people consent to the new route and whether the system should still exist six months later.

Counting generated apps will reward proliferation. Counting active users will reward compulsory dependence. Count the objective improved, the exceptions created, the human time released and the capability retained.


Operating-model implication

Create a work-authoring licence, not a universal ban

ScopeDefault freedomAdditional thresholdNamed steward
Personal aidBuild and discard within own data and tasksNo effect on another person’s rights or recordsUser
Team utilityShare a reversible tool with informed colleaguesNamed owner, visible limitations and expiryTeam process owner
Operational workflowDeploy after evidence and affected-user reviewTests, spend boundary, exception route and rollbackBusiness and technology owners
Consequential decisionNo local publication right by defaultIndependent assurance, legal basis, remedy and executive authorityAccountable institution

The point is not a new bureaucracy for every useful widget. It is a graduated right to publish work rules, sized to affected population, persistence, expenditure and consequence.


Human control watch

Assistance: a person uses AI to create or improve their own work. Control depends on data boundaries, inspectability and freedom to stop using the artefact.

Shared operation: an employee-built system changes how a team works. Control depends on colleague visibility, contestability, version ownership and an expiry route.

Delegated institution: a persistent agent or application routes work, spends resources or affects people beyond its creator. Control depends on publication authority, independent assurance, consequence limits and organisational remedy.

Today’s shift: human control is moving upstream from approving each machine action to deciding who may publish the system that makes future actions normal.


Capability-model update

Gaining valueUnder pressure
Operational-authorship stewardJob title as the map of organisational influence
Micro-institution product ownerCitizen development measured by app count
Agent-cost architectSoftware seats separated from workforce economics
Machine-handover designerPersistent work without re-confirmation
Constitutional assurance leadGenerated policy mistaken for legitimate policy
Local-practice curatorFrontline invention left as unpaid support work

1

ONE THING

IF I WERE TO DO ONE THING NOW

Issue one work-authoring licence

◇ INTENT ─── □ BUILD ─── │ PUBLISH RIGHT │ ─── ▶ RUN ─── ○ EXPIRE

A local invention becomes an organisational act at the moment other people must work inside it.

Return to the workflow whose jurisdiction boundary you set in the previous exercise—or choose one employee-built app, automation or agent already used by a team—and issue a one-page licence this week naming who may change it, how many people it may affect, what data and spend it may consume, which evidence is required before a new version goes live and the date on which its authority expires unless somebody deliberately renews it. Do not create an enterprise citizen-development policy. Test one licence with the builder, one affected colleague and the business owner; the conversation will reveal whether your organisation can distinguish helpful local invention from the quiet publication of new work rules, while giving responsible builders a legitimate route to improve the system.


Mental-model update

The permeable enterprise has acquired earned authority, an executable constitution, a capability border, a memory for failure, a limit on consequence, an accountability surface, a visible management layer, a rehearsal space and negotiated jurisdiction.

Today it acquires publishing discipline.

The organisation is no longer changed only through structure, policy and large systems. It is also changed by small executable artefacts that can spread local judgement through shared work. Elastic autonomy therefore requires a legitimate path from invention to institution—and an equally clear path back out.

The emerging North Star is a permeable enterprise in which people can improve work from the edge, machines can carry that improvement at scale and no helpful little system becomes permanent government by accident.

Questions for the executive table

  1. Which employees can already publish workflows that change colleagues’ choices, and does their formal authority reflect that reach?
  2. Where would an employee-built agent continue working after its creator changed role?
  3. Can your finance data connect agent spend to an objective, a human owner and the exception work around it?
  4. Who may approve an automatically generated runtime policy, and which affected people are represented in that decision?
  5. Which local inventions deserve to become shared organisational capability—and how will their creators be recognised rather than trapped as permanent support?

Evidence note. Microsoft describes products, previews, worked examples and intended availability; those claims do not establish adoption, reliability or business outcomes. Its run-assert-eval results use small samples and automated judges. The Ministry of Justice reports its own activity and governance. Kodiak reports its own intervention-free runs and planned launch; independent safety evidence has not been assessed here. World Economic Forum and ManpowerGroup material was reviewed as workforce context but was not used as a primary basis for the lead thesis.

The concepts operational authorship, micro-institution, work-authoring licence and publishing discipline are original AyEye analysis, not claims made by the cited sources.