When everyone can build the workflow, everyone can change the organisation
Microsoft is putting natural-language app creation, persistent delegation, managed runtime and metered spend into one work surface. The workforce consequence is operational authorship: deciding who may turn local judgement into a persistent system that changes how other people work.
Workforce intelligence · Issue 11
AyEye — Workforce Management
Human systems in the agentic enterprise
Saturday, 26 September 2026
For most of office history, changing how work worked required a project, a budget and at least one meeting that should have been an email. Microsoft now says a person will be able to describe an app, tracker, dashboard or workflow in ordinary language, run it inside the company tenant and hand recurring work to a persistent agent with its own identity, memory, computer and workspace.
The intriguing question is not whether everybody becomes a developer. It is whether everybody becomes a small-scale organisation designer.
The executive brief
- Microsoft has announced its largest Copilot redesign to date. A new Code surface is intended to let non-developers create purpose-built apps and workflows; Autopilot is a persistent agent that can continue work without a prompt; and a managed runtime can host what people build inside the company environment. Most elements are preview or forthcoming, not established outcomes.
- Agentic work is acquiring its own cost-allocation system. Microsoft will combine fixed user licences with usage-based billing for long-running work, model controls for groups and APIs for spending policies. A budget is becoming part of an agent’s practical mandate.
- Assurance is moving into the build loop. Microsoft’s open-source run-assert-eval skill attempts to discover risks, measure them, generate runtime policy and rerun the same evaluation. Its billing-support example reduced observed cross-customer violations, but it remains a worked example and human review is still required.
- The UK Ministry of Justice offers a public-sector counterexample to casual creation. It reports more than 1.5 million probation meetings transcribed, but places scaling inside a Justice AI Unit, portfolio visibility, ethics rules, testing and ongoing monitoring.
- Physical autonomy is showing the same separation between operation and oversight. Kodiak says a truck has completed the 219-mile Lancaster-to-Houston route without a human touching the wheel and plans unsupervised long-haul service by year-end. Intervention-free movement still depends on hubs, validation, fleet operations and a safety case.
- The workforce issue is authoring power. When a local employee can turn an intention into a persistent system that routes work, spends money and changes other people’s environment, deployment authority becomes a form of management authority.
When everyone can build the workflow, everyone can change the organisation
ORIGINAL SYNTHESIS · Confidence: medium-high · Horizon: 6–24 months
The next wave of “citizen development” is not simply easier software creation. It is the distribution of operational authorship: the power to turn a local judgement into a repeatable rule that other people and machines encounter as work.
Signal one: software creation enters the ordinary work surface
REPORTED FACT. On 25 September, Microsoft announced a new Copilot organised around Home, Code and Autopilot. Code is intended to let a person describe an app, tracker, dashboard, automation or workflow in natural language. Microsoft says the system can choose an approach and build persistent widgets, interactive dashboards or cloud-hosted internal applications that can be shared with a team.
Code uses technology related to GitHub Copilot, runs in a sandbox and can be hosted inside the organisation’s tenant. Microsoft Copilot Managed Runtime, now in preview, is intended to let IT govern hosted applications while employees connect them to live data. Home and Code are due to begin rolling out through the Frontier programme; broader availability is promised later. These are product claims and plans, not evidence that ordinary employees will build dependable production systems. Microsoft announcement, 25 September
ANALYSIS. The unit of knowledge work may be expanding from a file to a functioning system. A spreadsheet describes a rota; a small application allocates it. A memo proposes an approval rule; a workflow applies it. A presentation recommends a customer sequence; an agent pursues it after the meeting ends.
That shift matters because software is not neutral stationery. It makes some actions easy, others difficult and still others invisible. The person who creates a local workflow is deciding what counts as a case, which evidence is requested, where an exception goes and when somebody else is interrupted.
Signal two: delegation becomes persistent and economically metered
REPORTED FACT. Microsoft’s renamed Autopilot—previously Scout—is described as a cloud-hosted agent with its own identity, memory, computer and workspace. A user gives it a name, role and goal; it can watch channels, follow up on threads, run recurring work and resume projects days later. Microsoft’s example is an agent managing a supplier-review process, including the schedule, meetings and stakeholder follow-up.
The same announcement separates everyday Copilot use from long-running agentic work. Quick assistance remains covered by a user subscription, while Cowork, Code, Autopilot and frontier models use usage-based billing. Administrators can set spending policies, route credit requests through approval workflows, restrict model families for groups and compare task outcomes with cost.
ANALYSIS. Identity says which machine actor is present. Memory gives it continuity. A runtime gives its creations somewhere to operate. Usage billing gives activity an economic boundary. Put together, those are not the attributes of a clever document tool. They are the beginnings of a local operating unit.
The budget is especially revealing. A person with permission to create a recurring agent is not only delegating labour. They are creating a stream of future expenditure and attention. Cost control therefore belongs beside authority, not in a monthly cloud bill discovered after the work has already changed.
Signal three: the control loop is also becoming generative
REPORTED FACT. Microsoft released an open-source skill called run-assert-eval on 24 September. From a prompt in Visual Studio Code, it combines threat modelling, evaluation, runtime-policy generation and a repeated test intended to determine whether the policy fixed the measured failure without destroying useful behaviour.
In Microsoft’s worked billing-support example, the baseline agent exposed another customer’s data in 12 of 40 applicable conversations. After a deterministic account-matching policy was added before and after tool calls, the governed run recorded two violations in 34 applicable conversations. Across the reported splits, permissible-behaviour violations fell to zero. The samples are small, the judge is automated and the results come from the tool’s maker. Microsoft explicitly says generated policy is not automatic approval: a person must review the policy, intervention point and wiring. Microsoft run-assert-eval, 24 September
ANALYSIS. This closes part of the distance between “somebody should check this” and an enforceable control. It also creates a recursion problem. If building, threat discovery, test design and policy drafting all become easier, the scarce activity becomes accepting the institutional consequence: deciding which risk matters, who may approve the control, what evidence is sufficient and who owns the system when its original creator moves on.
Signal four: scale still needs an institution
REPORTED FACT. The UK Ministry of Justice said on 24 September that its Justice Transcribe tool had processed more than 1.5 million probation meetings. It describes a Justice AI Unit, a chief AI officer, cross-functional steering and risk groups, an internal portfolio tracker and requirements for accuracy, bias, fairness, security and reliability testing before wide deployment, with ongoing monitoring. It also says future work will involve appropriate engagement, including trade unions where relevant. Ministry of Justice update, 24 September
ANALYSIS. The public sector example is not proof that every control works. It is a useful contrast. Ease of creation does not remove the need for a body that can see the portfolio, retain learning across projects and connect technical behaviour to professional duty, public legitimacy and employee voice.
The operational-authorship stack
| Layer | Newly easy activity | Organisational right at stake | Evidence that should survive |
|---|---|---|---|
| Describe | Turn an intention into a specification | Who may define the problem and affected group? | Purpose, assumptions and exclusions |
| Build | Generate an app, workflow or agent | Who may encode a work rule? | Creator, data, model, tests and version |
| Run | Host it inside the enterprise | Who may make the rule operational? | Users, permissions, dependencies and expiry |
| Delegate | Let a persistent agent continue without prompts | Who may initiate future acts? | Objective, boundaries, exceptions and stop right |
| Spend | Consume models, compute and attention | Who may commit recurring resources? | Budget, outcome, unit cost and owner |
| Assure | Generate tests and draft controls | Who decides that evidence is good enough? | Frozen test, human approval and residual failure |
| Retire | End or replace the local system | Who protects people from orphaned rules? | Dependencies, records, remedy and shutdown |
What to notice: technical permission to create is only one layer. Operational authorship includes the rights to make a rule live, let it persist, fund its activity, certify its safety and remove it cleanly.
ORIGINAL SYNTHESIS. These developments point to a new management capability: operational authorship—the legitimate ability to translate local judgement into a persistent work system.
Organisations have distributed authorship before. Spreadsheet macros, low-code tools and departmental databases are old acquaintances, occasionally discovered under a desk like a startled hedgehog. What changes now is the combination. A single work surface can help a person design software, connect live enterprise context, assign a persistent agent, host the result, meter its expenditure and generate parts of its control system.
The resulting artefact is more than an app and less than a department. AyEye calls it a micro-institution: a small, executable arrangement of purpose, rules, resources and authority that can coordinate work beyond its author’s immediate attention.
Some micro-institutions will be splendidly modest: a tracker that keeps a handover from being forgotten. Others will influence hiring queues, supplier decisions, customer treatment or employee schedules. The design challenge is to preserve the first without pretending the second is merely personal productivity.
Unexpected connection
Natural-language software + persistent delegation + usage-based finance + automated assurance + public-sector portfolio governance
These look like separate product and policy developments. Together they create a production system for new ways of working. The employee supplies intent; a model generates structure; a runtime gives it persistence; finance grants fuel; assurance shapes its boundary; the organisation inherits the consequence.
The conclusion is surprising but practical: AI literacy is becoming less about knowing what to ask a model and more about knowing when an answer has become an institution.
PROVOCATION
Your organisation chart will matter less than your publishing permissions
A manager with five direct reports may soon be able to create systems that route the work of hundreds of colleagues, customers and agents. A junior specialist may encode an exception rule that travels further than an executive instruction. Formal hierarchy will still allocate accountability, pay and status; operational power will increasingly follow who may publish executable rules into shared work. If leaders govern titles but ignore publishing rights, the real organisation will grow beside the visible one.
What if we are right?
Opportunity. Frontline knowledge could become reusable capability without waiting for a central transformation programme. People closest to a recurring problem could build a small solution, test it with affected colleagues and improve it while the context is still fresh. The organisation could gain thousands of local experiments and retain the best as shared practice.
Organisational consequence. Job design would distinguish using an AI tool from authoring a work system. Identity, HCM, finance, architecture and risk would need a shared record of who may build, deploy, fund and retire micro-institutions. Recognition and progression would include the capability a person safely creates for others, not only the tasks they personally complete.
Likely horizon. Local apps and long-running agents are entering preview now. Within 6–12 months, early adopters will have enough employee-built systems to expose ownership and retirement problems. A portable work-authoring record across HCM, identity, portfolio and cost systems is plausible within 18–36 months.
What would prove us wrong?
The thesis weakens if Code-like systems remain novelty tools, if security constraints keep them isolated from consequential workflows, or if employees consistently build only personal utilities that disappear when a session ends. It also weakens if central platforms can automatically contain access, spend, quality and lifecycle without adding material human governance.
Microsoft’s announcement is a product roadmap, not usage evidence. The Ministry of Justice reports activity and governance structures, not independently verified service outcomes. Automated evaluation may reduce assurance work without becoming trustworthy enough for release decisions.
The concept fails if “operational authorship” becomes a grand title for every spreadsheet or an excuse to slow useful local improvement. The disconfirming test is straightforward: if employee-built applications and agents do not begin changing other people’s work, consuming persistent resources or surviving their creators, ordinary end-user computing controls may be enough.
Optimistic possibility: the organisation can learn at the edge
For decades, improvement has often travelled through a narrow pipe. A frontline worker spots a better route; a project queue translates it months later; context drains away between the two. Natural-language building can widen that pipe.
The constructive future is not uncontrolled shadow IT. It is legitimate local invention: people receive a bounded right to improve work, affected colleagues can challenge the design, evidence travels with the artefact and successful local practice can be adopted elsewhere. That could make organisational learning faster and more democratic—provided authorship earns responsibility without trapping the author as permanent unpaid support.
Citizen development is becoming citizen management
WORK DESIGN · IMPLEMENTATION SIGNAL
REPORTED FACT. Microsoft says Code will let people create shared internal apps from ordinary language, while specialised Skills in Word and Excel can distribute reusable role-specific workflows.
ANALYSIS. The phrase “citizen developer” frames the person as a software maker. “Citizen manager” is closer to the organisational consequence. The artefact can define a sequence, assign a queue, request evidence and escalate an exception. Those are management acts even when no manager performs them live.
Organisations should avoid solving this by classifying every builder as a manager. The better move is to recognise work-rule authorship as a capability with a scope. A finance analyst might be trusted to build tools for their own team but not to publish a credit-decision workflow. A service adviser might encode a handover checklist but not change who receives priority.
The important boundary is the affected population. Personal utility can be governed lightly. A system that changes another person’s options needs review proportionate to its reach and consequence.
A persistent agent creates a new kind of unfinished work
AGENT OPERATIONS · CONTINUITY
Autopilot’s appeal is that work continues while a person sleeps or turns to something else. That continuity also means decisions can accumulate while the original context changes.
ANALYSIS. The traditional handover has a natural social moment: one person tells another what remains open. A cloud agent may simply continue. The organisation therefore needs a machine handover state: what objective is active, which assumptions remain true, what the agent has promised, how much it has spent and what must be re-confirmed after silence, absence or organisational change.
Persistence should expire by default when the principal changes role, a supplier review ends or the affected team reorganises. Otherwise the enterprise will acquire zombie workflows: still technically authorised, still spending, still following goals that no current owner would write down.
FinOps is becoming part of workforce design
ECONOMICS · OPERATING MODEL
REPORTED FACT. Microsoft is introducing separate usage-based billing for long-running agentic work, group-level model controls, API-managed spending policies and outcome-versus-cost views.
ANALYSIS. Workforce planning has traditionally allocated people and money in different systems. Agentic work makes them inseparable at task level. A model choice changes cost, speed and quality; a persistent agent converts a managerial objective into recurring consumption; an employee can request more credits much as a manager requests overtime or contractor budget.
This is not a reason to treat agents as employees. It is a reason to connect machine expenditure to the human work system it alters. The meaningful unit is not tokens, seats or “digital workers”. It is the cost of achieving an objective, including human briefing, exception handling, assurance and repair.
An automatically generated control still needs a constitutional moment
ASSURANCE · HUMAN CONTROL
Run-assert-eval is constructive precisely because it does not stop at finding a failure. It tries to turn observed behaviour into an enforceable rule and retest the same cases.
ANALYSIS. The danger is not that machines help write policy. It is that policy generation makes a contested judgement look like plumbing. Choosing the critical risk, defining permissible behaviour and accepting residual failure are decisions about whose interests count.
The constitutional moment is the point at which a named human body says: this rule may now constrain work. Automation can supply evidence, draft the control and keep the comparison honest. It cannot determine legitimate trade-offs on behalf of people who were never represented.
Intervention-free does not mean institution-free
PHYSICAL AI · OUTSIDE-IN ANALYSIS
REPORTED FACT. Kodiak said on 25 September that its autonomous truck had completed the 219-mile route from its Lancaster, Texas hub to Houston without a human touching the wheel, including surface streets, and that it plans unsupervised long-haul service on the lane by the end of 2026. The company is still conducting daily preparation runs and closed-course validation. These are supplier claims and plans, not an independent safety finding. Kodiak announcement, 25 September
ANALYSIS. The absent driver is the visible change; the newly designed institution is the deeper one. A hub must prepare and release the vehicle. A safety case must justify operation. Remote, maintenance and recovery roles must exist even if they do not touch the wheel on a successful run.
The same principle applies to office agents. “No human intervention” describes a normal execution path. It says nothing about who created the path, validates it, absorbs exceptions or restores service. Workforce plans should count that surrounding institution before booking the saving.
The shadow organisation may be made of helpful little apps
TENUOUS BUT PLAUSIBLE · Confidence: medium-low · Horizon: 12–36 months
There is no evidence yet that Copilot Code has created a parallel management structure. The product is not broadly available. The causal chain is nevertheless credible:
Natural-language building reduces creation cost → more employees publish local workflows → some workflows persist and shape colleagues’ choices → operational influence follows artefacts rather than reporting lines → the application portfolio becomes a partial shadow organisation chart.
SPECULATION. Leaders may discover that the most influential person in a process is not its formal owner but the employee who maintains the small system everybody has learned to obey.
What to watch: whether employee-built apps gain shared users; whether they survive job moves; whether performance reviews reward their creation; whether colleagues can contest a rule; and whether HCM, identity and application inventories can identify the human steward without treating them as solely liable for every outcome.
Noise filter: “everyone can build” is not an operating model
Creation demos show the happy path from prompt to useful artefact. Enterprise value begins later: somebody must decide whether the problem deserved automation, whether the data is fit, whether affected people consent to the new route and whether the system should still exist six months later.
Counting generated apps will reward proliferation. Counting active users will reward compulsory dependence. Count the objective improved, the exceptions created, the human time released and the capability retained.
Operating-model implication
Create a work-authoring licence, not a universal ban
| Scope | Default freedom | Additional threshold | Named steward |
|---|---|---|---|
| Personal aid | Build and discard within own data and tasks | No effect on another person’s rights or records | User |
| Team utility | Share a reversible tool with informed colleagues | Named owner, visible limitations and expiry | Team process owner |
| Operational workflow | Deploy after evidence and affected-user review | Tests, spend boundary, exception route and rollback | Business and technology owners |
| Consequential decision | No local publication right by default | Independent assurance, legal basis, remedy and executive authority | Accountable institution |
The point is not a new bureaucracy for every useful widget. It is a graduated right to publish work rules, sized to affected population, persistence, expenditure and consequence.
Human control watch
Assistance: a person uses AI to create or improve their own work. Control depends on data boundaries, inspectability and freedom to stop using the artefact.
Shared operation: an employee-built system changes how a team works. Control depends on colleague visibility, contestability, version ownership and an expiry route.
Delegated institution: a persistent agent or application routes work, spends resources or affects people beyond its creator. Control depends on publication authority, independent assurance, consequence limits and organisational remedy.
Today’s shift: human control is moving upstream from approving each machine action to deciding who may publish the system that makes future actions normal.
Capability-model update
| Gaining value | Under pressure |
|---|---|
| Operational-authorship steward | Job title as the map of organisational influence |
| Micro-institution product owner | Citizen development measured by app count |
| Agent-cost architect | Software seats separated from workforce economics |
| Machine-handover designer | Persistent work without re-confirmation |
| Constitutional assurance lead | Generated policy mistaken for legitimate policy |
| Local-practice curator | Frontline invention left as unpaid support work |
ONE THING
IF I WERE TO DO ONE THING NOW
Issue one work-authoring licence
◇ INTENT ─── □ BUILD ─── │ PUBLISH RIGHT │ ─── ▶ RUN ─── ○ EXPIRE
Return to the workflow whose jurisdiction boundary you set in the previous exercise—or choose one employee-built app, automation or agent already used by a team—and issue a one-page licence this week naming who may change it, how many people it may affect, what data and spend it may consume, which evidence is required before a new version goes live and the date on which its authority expires unless somebody deliberately renews it. Do not create an enterprise citizen-development policy. Test one licence with the builder, one affected colleague and the business owner; the conversation will reveal whether your organisation can distinguish helpful local invention from the quiet publication of new work rules, while giving responsible builders a legitimate route to improve the system.
Mental-model update
The permeable enterprise has acquired earned authority, an executable constitution, a capability border, a memory for failure, a limit on consequence, an accountability surface, a visible management layer, a rehearsal space and negotiated jurisdiction.
Today it acquires publishing discipline.
The organisation is no longer changed only through structure, policy and large systems. It is also changed by small executable artefacts that can spread local judgement through shared work. Elastic autonomy therefore requires a legitimate path from invention to institution—and an equally clear path back out.
The emerging North Star is a permeable enterprise in which people can improve work from the edge, machines can carry that improvement at scale and no helpful little system becomes permanent government by accident.
Questions for the executive table
- Which employees can already publish workflows that change colleagues’ choices, and does their formal authority reflect that reach?
- Where would an employee-built agent continue working after its creator changed role?
- Can your finance data connect agent spend to an objective, a human owner and the exception work around it?
- Who may approve an automatically generated runtime policy, and which affected people are represented in that decision?
- Which local inventions deserve to become shared organisational capability—and how will their creators be recognised rather than trapped as permanent support?
Evidence note. Microsoft describes products, previews, worked examples and intended availability; those claims do not establish adoption, reliability or business outcomes. Its run-assert-eval results use small samples and automated judges. The Ministry of Justice reports its own activity and governance. Kodiak reports its own intervention-free runs and planned launch; independent safety evidence has not been assessed here. World Economic Forum and ManpowerGroup material was reviewed as workforce context but was not used as a primary basis for the lead thesis.
The concepts operational authorship, micro-institution, work-authoring licence and publishing discipline are original AyEye analysis, not claims made by the cited sources.
