Stop testing agents in empty rooms
New research shows that curated task tests can materially overstate workplace performance, while self-organising agents derive capability from shared organisational infrastructure. Enterprises should assure the configured human–agent institution, not merely the model.
Workforce intelligence · Issue 9
AyEye — Workforce Management
Human systems in the agentic enterprise
Thursday, 24 September 2026
We test AI agents as though work happens in a freshly painted room: one task, the right files and nothing awkward left in the cupboard. Real organisations contain history, partial permissions, competing work, missing context and other actors. New research shows that this difference is not scenery. It changes whether the agent succeeds. The next unit of AI assurance is therefore not the agent alone, but the organisation it enters.
The executive brief
- A new workplace benchmark finds that curated task context materially overstates performance. In WorkWorlds, giving agents the full workplace visible to a role rather than a preselected file set reduced evidence access from 90.4% to 74.5% and criterion pass from 79.4% to 68.2%. Most of the loss occurred before the agent found enough evidence.
- Microsoft researchers have scaled self-organising software work to 1,024 agents. Agensh replaces a central orchestrator with shared workspace, messaging and context. Coordination, verification and role specialisation emerge through infrastructure—but only in constrained coding experiments so far.
- Anthropic says 950 agents helped surface an unusual enzyme system. Workers, supervisor agents, a shared record and human laboratory scientists formed one discovery system. The biological function remains unknown, but the work shows why “the model did it” is a poor account of capability.
- Security operations are being rebuilt as a shared environment for people and agents. Microsoft’s new ISOC design combines signals, context and actuators; Cisco’s vendor-sponsored survey says 51% of respondents already allow agentic systems to act in production networks.
- Agent evaluation is moving closer to professional certification. Casepoint’s legal-review agents are sampled, scored and approved by users before a full run. This is still a supplier implementation, but it points towards authority earned inside a particular workflow rather than inherited from a generic benchmark.
- Training availability is not the same as learning capacity. Workera reports that employer AI training has expanded sharply while 56% of surveyed employees receive no working time to build the skills. A capable organisation needs protected human learning as well as machine evaluation.
ORIGINAL SYNTHESIS · Confidence: medium-high · Horizon: 6–24 months
Stop testing agents in empty rooms
An agent’s performance is partly produced by the workplace around it: what the role can see, how work is claimed, which history persists, who verifies progress and how exceptions return to people. That makes organisational design a component of technical capability.
Signal one: a tidy task can hide the hardest part of work
REPORTED FACT. WorkWorlds, revised on 23 September, is an evaluation infrastructure built around a persistent synthetic organisation rather than a collection of isolated tasks.
A WorkWorld contains employees, roles, permissions, files, records, systems, communications and prior events. Researchers first fix a date, organisational revision and employee seat; only then do they add an assignment. The agent sees what that role could see at that moment, while the grading material remains separate.
Across 192 matched evaluations, agents given a task-curated file set accessed sufficient evidence 90.4% of the time and passed 79.4% of criteria. In the full role-visible workplace, evidence access fell to 74.5% and criterion pass to 68.2%. Conditional on reaching the evidence, performance changed much less. The main loss came from finding the right evidence inside organisational reality.
The study uses synthetic organisations, eight measured tasks in its primary pharmaceutical world and a limited set of agents. It does not establish how any deployed enterprise agent will perform. It does expose a benchmark design problem: selecting “relevant context” in advance quietly completes part of the job. WorkWorlds paper, 23 September
ANALYSIS. Organisations often evaluate an agent with a good prompt, approved documents and a known answer. The test asks whether the agent can reason after somebody else has already navigated the organisation.
Real work begins earlier. Which policy version applies? What can this role legitimately see? Did a customer exception change the case yesterday? Is the decisive evidence in a message, system record or colleague’s memory? Context discovery is not preparatory fuss. It is work.
Signal two: management can be encoded as a cooperation protocol
REPORTED FACT. Microsoft Research authors introduced Agensh on 22 September, a multi-agent harness designed to scale without a central orchestrator assigning every sub-task.
Agents gather context, claim work, act, share findings, verify results and merge progress asynchronously. A shared workspace records proposed, active and completed work; a messaging layer supports coordination; shared context retains reusable findings and intentions.
Using the same model and six-hour budget on the five hardest ProgramBench software tasks, scaling from one to 128 agents increased mean test-pass rate from 19.31% to 28.78%, a 49% relative improvement. On a pandoc reproduction task, the final pass rate rose from 33.89% with one agent to 55.06% with 1,024. Recorded trajectories showed peer coordination, integration work, standardised routines and specialised roles emerging as the system grew.
These are software-reproduction experiments without internet access, not evidence that 1,024 agents can self-organise a company. The absolute pass rates also remind us that scale did not make the work reliable. Agensh technical report, 22 September
ANALYSIS. Yesterday’s edition made the human labour of briefing, routing and checking agents visible. Agensh adds a changed implication: some of that management can move into the rules and shared infrastructure through which agents coordinate.
No central machine manager allocated every act. Yet the system was not managerless. Its designers decided how work could be claimed, how progress became visible, how conflicts were detected, when verification occurred and how contributions were merged.
This is protocol management: managerial decisions expressed as executable conditions for cooperation rather than repeated instructions from a supervisor.
Signal three: a scientific result belongs to a mixed institution
REPORTED FACT. Anthropic reported on 23 September that a campaign involving roughly 950 Claude agents searched genomic data and surfaced a previously uncharacterised family of array-associated reverse transcriptases.
A human-written research brief was divided into stages. Worker agents proposed and executed analyses; supervisor agents reviewed plans and results and could open new tasks; a shared record preserved findings. Human scientists then inspected the result and performed laboratory experiments. Anthropic says the agents noticed an unusual repeated DNA structure beside a known enzyme without being specifically told to search for it.
The system’s biological function remains unknown and further experiments are under way. Anthropic created and reports on the research, so claims of novelty and autonomy require independent scientific scrutiny. The technical report is nevertheless unusually specific about the division of work. Anthropic, 23 September · Technical report
ANALYSIS. The discovery cannot be sensibly attributed to one “AI scientist”. Capability came from the research brief, agent roles, databases, tools, task queue, common memory, review protocol and physical laboratory.
The human scientists did not merely approve a final answer. They built the institution in which a machine observation could become evidence.
Signal four: operations platforms are becoming shared workplaces
REPORTED FACT. Microsoft announced an integrated security operations centre design in Defender on 23 September. Its architecture joins signals and sensors, shared context and actuators so people and agents can investigate and protect the same environment rather than operate through disconnected tools.
Microsoft describes a change from practitioners stitching together information across products towards directing outcomes inside a common operating foundation. This is a supplier account of intended capability, not an independent result. Microsoft Security, 23 September
Separately, Cisco published a survey conducted by Omdia of 1,000 IT and network-operations leaders at organisations with at least 500 employees. It says 51% already use agentic AI that acts in production, 82% are comfortable allowing at least some production changes without prior approval and 24% are comfortable with fully autonomous operation. Yet 69% require detailed explanations and 36% say full tracing, rationale and post-action audit are the minimum acceptable standard.
The study was sponsored by Cisco and reflects reported attitudes, not audited deployments. Its value is directional: production autonomy is arriving inside complex, persistent environments where context and observability are part of the capability. Cisco and Omdia, 23 September
The organisational-competence stack
| Layer | What must be designed | What an isolated task test misses | Evidence of fitness |
|---|---|---|---|
| World | History, data, systems, roles and permissions | Finding the right evidence without privileged curation | Role-visible retrieval, provenance and access boundaries |
| Assignment | Objective, trade-offs, limits and success criteria | Ambiguity and competing organisational purposes | Brief quality, clarification and refusal behaviour |
| Cooperation | Claim, hand-off, communication, conflict and merge rules | Coordination cost and duplicated or incompatible work | Shared state, contribution trace and integration quality |
| Challenge | Verification, intervention, escalation and remedy | Whether somebody can detect and alter a bad trajectory | Counterexamples, stop rights and recovery performance |
| Learning | Persistent lessons, revised authority and human development | Whether the system improves—or merely repeats the demo | Fewer repeated failures and stronger residual capability |
ORIGINAL SYNTHESIS. These developments suggest that an enterprise agent has no stable workplace competence in isolation. Its effective capability is relational: the product of the model, the organisational state it can reach, the role it occupies, the protocols through which it cooperates and the human institution that challenges and repairs its work.
Call this organisational competence: the demonstrated ability of a human–agent system to achieve an objective inside a particular institution without violating its permissions, losing its history or exhausting its capacity to notice and recover.
This changes evaluation. A model benchmark can tell us something about reasoning. A task test can tell us whether one assignment succeeds with selected context. Neither establishes whether the agent can work inside a changing company.
It also changes organisation design. WorkWorlds shows that organisational mess reduces apparent competence. Agensh shows that carefully designed shared infrastructure can produce cooperation without constant central allocation. The science campaign shows that discovery depends on role design, common memory and laboratory verification. Security platforms show the same pattern moving into production.
The organisational implication is not that companies should simulate everything. It is that the workplace is now part of the system under test.
Unexpected connection
Organisation theory + multi-agent software engineering + genome mining + security operations
WorkWorlds explicitly borrows from organisation and routine theory: a workplace persists while particular performances come and go. Agensh turns coordination routines into machine-executable infrastructure. Anthropic uses a role system and common memory to search biology. Microsoft builds a common operational world for defenders and agents.
Together they make an unexpected claim: organisation design is becoming part of the AI stack. Roles, routines, shared memory and intervention rights are no longer merely the social context around the technology. They help determine what the technology can do.
PROVOCATION
Stop certifying agents. Certify the organisation they enter.
A generic agent can pass every laboratory test and still fail because it cannot find the operative policy, sees the wrong history, collides with another agent or hands an exception to nobody. For consequential work, approval should attach to a configured human–agent institution: this version, in this role, with these permissions, routines, colleagues, monitors and remedies. The agent matters. The arrangement deserves the licence.
What if we are right?
Opportunity. Enterprises could rehearse mixed work before exposing employees, customers or live systems to it. Persistent test organisations would reveal missing context, brittle hand-offs and unowned exceptions while changes remain cheap. Smaller models might outperform expensive ones when placed in a better-designed institution.
Organisational consequence. Model evaluation, process design, identity, HCM, enterprise architecture and operational assurance would share a common test object: an objective inside a role-visible organisational world. Leaders would buy less reassurance from one-off demos and more evidence about retrieval, cooperation, intervention and recovery.
Likely horizon. A role-sized test world can be assembled now. Persistent synthetic organisations, cross-agent protocol tests and configuration-specific certification are plausible within 12–24 months, first in software, security, legal review, finance and regulated operations.
What would prove us wrong?
The thesis weakens if most valuable enterprise tasks remain self-contained; if retrieval and access systems make organisational context nearly frictionless; or if individual-agent benchmark performance reliably predicts production outcomes across companies.
Agensh may not transfer beyond software, where tests and merge rules make progress unusually legible. WorkWorlds is small and synthetic. Anthropic’s scientific campaign came from one laboratory and its biological result is incomplete. Platform integration can also centralise errors rather than create competence.
The model fails if “certify the organisation” becomes an excuse for exhaustive digital twins, slow committees or permanent approval. The practical test is comparative: does a persistent, role-visible evaluation predict real exceptions and outcomes better than the current curated task test?
Optimistic possibility: organisations can practise before people pay the price
We normally discover work-design errors after launch, when a frontline employee becomes the integration layer or a customer meets the exception nobody modelled.
A small persistent test workplace creates a more generous possibility. Teams can observe how authority, context and hand-offs behave before judging either the worker or the agent. Employees can contribute real exceptions without surrendering live data. Managers can improve the system rather than counsel people to cope with it.
The prize is not a perfect simulation. It is the ability to learn about the organisation while the consequences are still fictional.
SCIENCE & DISCOVERY · OUTSIDE-IN ANALYSIS
Scientific apprenticeship moves towards choosing what deserves a laboratory
Anthropic’s campaign produced approximately 200,000 candidate sequences before one unusual pattern became the focus of human experimental work. That ratio matters more to workforce design than the headline “AI discovery”.
ANALYSIS. When agents make search and hypothesis generation abundant, scarcity moves into selection, experimental capacity and interpretation. Junior scientists may encounter more plausible findings than they can personally reproduce. Senior judgement can become a queue.
A healthy research organisation should not let the machine monopolise anomaly detection while novices only receive polished candidates. Apprenticeship needs protected access to the messy middle: failed searches, ambiguous signals, contradictory evidence and the decision not to spend laboratory time.
The new capability is not simply prompt engineering. It is experimental triage: deciding which machine-generated observation deserves scarce physical proof and explaining why the rejected alternatives do not.
PROFESSIONAL WORK · IMPLEMENTATION SIGNAL
Legal agents are beginning to earn authority through sampling
REPORTED FACT. Casepoint announced two agents for legal and government document review on 23 September. The company says a Relevance Agent and Issue Coding Agent use multiple models to review documents, cross-check decisions and produce recall and precision measures.
Before a full run, users review a sample and decide whether the configuration meets their threshold. Actions remain inside existing permission boundaries and audit trails. These are supplier claims; no independent comparison or customer outcome was published. Casepoint, 23 September
ANALYSIS. The useful pattern is not “AI does legal review”. It is staged authority. Performance is tested on work drawn from the local matter, with a human accepting the threshold before scope expands.
That resembles probation more than generic certification: competence is demonstrated inside the case, under its vocabulary, permissions and error costs. The missing evidence is longitudinal—whether the agent remains fit as the matter, data and legal strategy change.
TENUOUS BUT PLAUSIBLE · Confidence: medium-low · Horizon: 12–36 months
Agents may need a digital probationary workplace
WorkWorlds is a research infrastructure, not an enterprise deployment standard. Casepoint’s sample approval is product-specific. But together they suggest a cautious new institution.
SPECULATION. Before receiving live authority, an agent could spend time in a digital probationary workplace: a persistent, role-visible environment containing representative history, permissions, colleagues, interruptions and exceptions. It would be assessed not only on completed tasks but on how it finds evidence, shares state, handles conflict, escalates uncertainty and responds when the organisation changes.
The causal chain is plausible:
Persistent organisational simulation makes context and cooperation testable → local failure modes appear before production → authority expands on evidence rather than vendor category → human reviewers focus on changed conditions and residual uncertainty.
This could become bureaucratic theatre, and synthetic worlds may omit exactly the social signals that matter. What to watch is whether configuration-specific rehearsal predicts incidents better than task benchmarks, whether affected employees help design the cases and whether probation ends with a clear decision to authorise, narrow or reject.
LEARNING & CAPACITY · WORKFORCE SIGNAL
Training without time is organisational make-believe
REPORTED FACT. Workera surveyed 1,000 salaried US employees in large organisations in July and reported on 23 September that 58% now receive employer-provided AI training, up from 25% in its March 2025 comparison. Yet 56% said their employer provides no working time to build AI skills.
The survey was commissioned by a skills platform, uses self-report and is limited to the United States. It should not be treated as a global prevalence measure. Workera, 23 September
ANALYSIS. If agent competence depends on organisational context, human learning must occur in context too. A course can explain a tool; it cannot replace time to test a real workflow, find weak evidence, negotiate a hand-off and learn when to stop.
Protected practice is not a benefit offered after transformation. It is part of the organisation’s supervisory capital. A company that allocates compute to machine rehearsal but no work time to human learning is improving only half of the system.
NOISE FILTER
One thousand agents are not one thousand colleagues
Agensh’s scale is striking, but the agents are instances of the same model working for hours inside a purpose-built software environment. They do not have livelihoods, rights, careers or independent organisational membership.
The experiment supports a claim about concurrent machine cooperation, not a headcount comparison or a general law that more agents always produce more intelligence. Final performance remained imperfect, and coordination infrastructure carried much of the value.
Count the work, conflicts, verification and outcomes. Leave the imaginary staff photograph alone.
Operating-model implication
Assure the configured work system
| Assurance question | Evidence to retain | Decision produced | Primary steward |
|---|---|---|---|
| Can the role find the right evidence? | Role-visible sources, retrieval path and provenance | Improve context, access or task placement | Knowledge, identity and process owners |
| Can contributors cooperate without hidden collision? | Claims, hand-offs, messages, conflicts and merges | Change protocol, capacity or role boundaries | Operations and architecture |
| Can somebody challenge the trajectory? | Tests, uncertainty, intervention and escalation trace | Expand, narrow or pause authority | Business owner and independent assurance |
| Does the system learn safely? | Persistent lessons, changed configuration and regression tests | Re-authorise the arrangement | Joint human–agent learning review |
| Are people becoming more capable? | Practice time, unaided judgement and recovery skill | Redesign learning or preserve manual experience | Manager, profession lead and HR |
The object under assurance is neither the model nor the workflow document. It is the configured work system: objective, workplace state, people, agents, permissions, cooperation rules and remedy.
Human control watch
Assistance: a person performs the task with AI support. Control depends on whether the role can see the relevant organisational evidence and understand why it was selected.
Delegated execution: an agent works within a persistent role and hands results to people or other agents. Control depends on role-visible context, protocol design, sampling and clear exception ownership.
Autonomous cooperation: agents claim, divide, verify and merge work without continuous assignment. Control moves into the shared world, cooperation rules, monitors, consequence limits and authority to alter the protocol.
Today’s shift: the human need not allocate every task to retain control. But people must govern the institution through which allocation, evidence, challenge and recovery occur.
Capability-model update
| Gaining value | Under pressure |
|---|---|
| Organisational-context engineer | Curated prompt as proof of workplace readiness |
| Cooperation-protocol designer | Central orchestration of every machine action |
| Role-visible evaluation architect | Generic benchmark as deployment licence |
| Experimental-triage scientist | Hypothesis generation as the sole research bottleneck |
| Digital-probation steward | One-off pilot in a clean environment |
| Protected-practice designer | Training availability without learning time |
ONE THING
IF I WERE TO DO ONE THING NOW
Build one role-visible test world
□ ROLE ─── ◇ HISTORY ─── ⌁ PERMISSIONS ─── △ EXCEPTION ─── ✓ AUTHORISE?
Take the agent-assisted workflow whose management loop you timed in the previous exercise—or another live candidate—and assemble one small, reusable test world this week: the role’s actual permissions, three to five representative artefacts, one prior event that changes the current answer, one awkward exception and the acceptance and stop criteria. Run the same agent through it without pointing to the relevant evidence, then review the trace with the frontline user and business owner: what it found, missed, shared, assumed and escalated. Do not build a digital twin or procurement benchmark. Preserve this single world as a regression test. It converts the newly visible management work into an organisational fitness check and creates evidence for whether authority should expand, narrow or wait.
Mental-model update
The permeable enterprise has acquired earned authority, an executable constitution, a capability border, a memory for failure, a limit on consequence, an accountability surface and a visible management layer.
Today it acquires a rehearsal space.
Elastic autonomy does not depend only on finding a better agent. It depends on building an organisation in which useful context is reachable, cooperation is legible, challenge has authority and both people and machines can learn before consequences become real.
The emerging North Star is a permeable enterprise that can test itself: capability may enter and self-organise, but the configured institution must demonstrate that it can turn machine action into responsible work.
Questions for the executive table
- Which agent passed a demonstration only because somebody quietly selected the right context first?
- What managerial decisions are now encoded in your claim, routing, verification and merge protocols—and who may change them?
- Can you reproduce the role-visible world in which a consequential agent was approved six months ago?
- Where will human specialists learn to reject attractive machine hypotheses before scarce operational or laboratory capacity is committed?
- Do employees receive protected time to improve the mixed system, or only training content and higher output expectations?
Evidence note. WorkWorlds and Agensh are research results in constrained, partly synthetic settings. Anthropic reports on its own agents and laboratory; the enzyme system’s function is unresolved. Microsoft, Cisco, Casepoint and Workera describe their own products or sponsored research. Cisco and Workera survey findings are self-reported and should not be treated as audited prevalence or performance.
The concepts organisational competence, protocol management, organisational-competence stack and digital probationary workplace are original AyEye analysis, not claims made by the cited sources.
