# Stop testing agents in empty rooms

New research shows that curated task tests can materially overstate workplace performance, while self-organising agents derive capability from shared organisational infrastructure. Enterprises should assure the configured human–agent institution, not merely the model.

**AyEye — Workforce Management · 2026-09-24**

Workforce intelligence · Issue 9

# AyEye — Workforce Management

Human systems in the agentic enterprise

Thursday, 24 September 2026

We test AI agents as though work happens in a freshly painted room: one task, the right files and nothing awkward left in the cupboard. Real organisations contain history, partial permissions, competing work, missing context and other actors. New research shows that this difference is not scenery. It changes whether the agent succeeds. The next unit of AI assurance is therefore not the agent alone, but the organisation it enters.

## The executive brief

- A new workplace benchmark finds that curated task context materially overstates performance. In WorkWorlds, giving agents the full workplace visible to a role rather than a preselected file set reduced evidence access from 90.4% to 74.5% and criterion pass from 79.4% to 68.2%. Most of the loss occurred before the agent found enough evidence.

- Microsoft researchers have scaled self-organising software work to 1,024 agents. Agensh replaces a central orchestrator with shared workspace, messaging and context. Coordination, verification and role specialisation emerge through infrastructure—but only in constrained coding experiments so far.

- Anthropic says 950 agents helped surface an unusual enzyme system. Workers, supervisor agents, a shared record and human laboratory scientists formed one discovery system. The biological function remains unknown, but the work shows why “the model did it” is a poor account of capability.

- Security operations are being rebuilt as a shared environment for people and agents. Microsoft’s new ISOC design combines signals, context and actuators; Cisco’s vendor-sponsored survey says 51% of respondents already allow agentic systems to act in production networks.

- Agent evaluation is moving closer to professional certification. Casepoint’s legal-review agents are sampled, scored and approved by users before a full run. This is still a supplier implementation, but it points towards authority earned inside a particular workflow rather than inherited from a generic benchmark.

- Training availability is not the same as learning capacity. Workera reports that employer AI training has expanded sharply while 56% of surveyed employees receive no working time to build the skills. A capable organisation needs protected human learning as well as machine evaluation.

ORIGINAL SYNTHESIS · Confidence: medium-high · Horizon: 6–24 months

## Stop testing agents in empty rooms

An agent’s performance is partly produced by the workplace around it: what the role can see, how work is claimed, which history persists, who verifies progress and how exceptions return to people. That makes organisational design a component of technical capability.

### Signal one: a tidy task can hide the hardest part of work

REPORTED FACT. WorkWorlds, revised on 23 September, is an evaluation infrastructure built around a persistent synthetic organisation rather than a collection of isolated tasks.

A WorkWorld contains employees, roles, permissions, files, records, systems, communications and prior events. Researchers first fix a date, organisational revision and employee seat; only then do they add an assignment. The agent sees what that role could see at that moment, while the grading material remains separate.

Across 192 matched evaluations, agents given a task-curated file set accessed sufficient evidence 90.4% of the time and passed 79.4% of criteria. In the full role-visible workplace, evidence access fell to 74.5% and criterion pass to 68.2%. Conditional on reaching the evidence, performance changed much less. The main loss came from finding the right evidence inside organisational reality.

The study uses synthetic organisations, eight measured tasks in its primary pharmaceutical world and a limited set of agents. It does not establish how any deployed enterprise agent will perform. It does expose a benchmark design problem: selecting “relevant context” in advance quietly completes part of the job. [WorkWorlds paper, 23 September](https://arxiv.org/html/2609.23806v2)

ANALYSIS. Organisations often evaluate an agent with a good prompt, approved documents and a known answer. The test asks whether the agent can reason after somebody else has already navigated the organisation.

Real work begins earlier. Which policy version applies? What can this role legitimately see? Did a customer exception change the case yesterday? Is the decisive evidence in a message, system record or colleague’s memory? Context discovery is not preparatory fuss. It is work.

### Signal two: management can be encoded as a cooperation protocol

REPORTED FACT. Microsoft Research authors introduced Agensh on 22 September, a multi-agent harness designed to scale without a central orchestrator assigning every sub-task.

Agents gather context, claim work, act, share findings, verify results and merge progress asynchronously. A shared workspace records proposed, active and completed work; a messaging layer supports coordination; shared context retains reusable findings and intentions.

Using the same model and six-hour budget on the five hardest ProgramBench software tasks, scaling from one to 128 agents increased mean test-pass rate from 19.31% to 28.78%, a 49% relative improvement. On a pandoc reproduction task, the final pass rate rose from 33.89% with one agent to 55.06% with 1,024. Recorded trajectories showed peer coordination, integration work, standardised routines and specialised roles emerging as the system grew.

These are software-reproduction experiments without internet access, not evidence that 1,024 agents can self-organise a company. The absolute pass rates also remind us that scale did not make the work reliable. [Agensh technical report, 22 September](https://arxiv.org/html/2609.26781v1)

ANALYSIS. Yesterday’s edition made the human labour of briefing, routing and checking agents visible. Agensh adds a changed implication: some of that management can move into the rules and shared infrastructure through which agents coordinate.

No central machine manager allocated every act. Yet the system was not managerless. Its designers decided how work could be claimed, how progress became visible, how conflicts were detected, when verification occurred and how contributions were merged.

This is protocol management: managerial decisions expressed as executable conditions for cooperation rather than repeated instructions from a supervisor.

### Signal three: a scientific result belongs to a mixed institution

REPORTED FACT. Anthropic reported on 23 September that a campaign involving roughly 950 Claude agents searched genomic data and surfaced a previously uncharacterised family of array-associated reverse transcriptases.

A human-written research brief was divided into stages. Worker agents proposed and executed analyses; supervisor agents reviewed plans and results and could open new tasks; a shared record preserved findings. Human scientists then inspected the result and performed laboratory experiments. Anthropic says the agents noticed an unusual repeated DNA structure beside a known enzyme without being specifically told to search for it.

The system’s biological function remains unknown and further experiments are under way. Anthropic created and reports on the research, so claims of novelty and autonomy require independent scientific scrutiny. The technical report is nevertheless unusually specific about the division of work. [Anthropic, 23 September](https://www.anthropic.com/news/claude-discovers-novel-enzyme-system) · [Technical report](https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf)

ANALYSIS. The discovery cannot be sensibly attributed to one “AI scientist”. Capability came from the research brief, agent roles, databases, tools, task queue, common memory, review protocol and physical laboratory.

The human scientists did not merely approve a final answer. They built the institution in which a machine observation could become evidence.

### Signal four: operations platforms are becoming shared workplaces

REPORTED FACT. Microsoft announced an integrated security operations centre design in Defender on 23 September. Its architecture joins signals and sensors, shared context and actuators so people and agents can investigate and protect the same environment rather than operate through disconnected tools.

Microsoft describes a change from practitioners stitching together information across products towards directing outcomes inside a common operating foundation. This is a supplier account of intended capability, not an independent result. [Microsoft Security, 23 September](https://www.microsoft.com/en-us/security/blog/2026/09/23/reimagining-the-soc-for-the-agentic-era-in-microsoft-defender/)

Separately, Cisco published a survey conducted by Omdia of 1,000 IT and network-operations leaders at organisations with at least 500 employees. It says 51% already use agentic AI that acts in production, 82% are comfortable allowing at least some production changes without prior approval and 24% are comfortable with fully autonomous operation. Yet 69% require detailed explanations and 36% say full tracing, rationale and post-action audit are the minimum acceptable standard.

The study was sponsored by Cisco and reflects reported attitudes, not audited deployments. Its value is directional: production autonomy is arriving inside complex, persistent environments where context and observability are part of the capability. [Cisco and Omdia, 23 September](https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2026/m09/cisco-ai-research-agenticops-scaling-quickly-in-the-enterprise.html)

### The organisational-competence stack

LayerWhat must be designedWhat an isolated task test missesEvidence of fitness

WorldHistory, data, systems, roles and permissionsFinding the right evidence without privileged curationRole-visible retrieval, provenance and access boundaries
AssignmentObjective, trade-offs, limits and success criteriaAmbiguity and competing organisational purposesBrief quality, clarification and refusal behaviour
CooperationClaim, hand-off, communication, conflict and merge rulesCoordination cost and duplicated or incompatible workShared state, contribution trace and integration quality
ChallengeVerification, intervention, escalation and remedyWhether somebody can detect and alter a bad trajectoryCounterexamples, stop rights and recovery performance
LearningPersistent lessons, revised authority and human developmentWhether the system improves—or merely repeats the demoFewer repeated failures and stronger residual capability

Agent capability becomes organisational competence only when the surrounding world, cooperation protocol and challenge system work together.

ORIGINAL SYNTHESIS. These developments suggest that an enterprise agent has no stable workplace competence in isolation. Its effective capability is relational: the product of the model, the organisational state it can reach, the role it occupies, the protocols through which it cooperates and the human institution that challenges and repairs its work.

Call this organisational competence: the demonstrated ability of a human–agent system to achieve an objective inside a particular institution without violating its permissions, losing its history or exhausting its capacity to notice and recover.

This changes evaluation. A model benchmark can tell us something about reasoning. A task test can tell us whether one assignment succeeds with selected context. Neither establishes whether the agent can work inside a changing company.

It also changes organisation design. WorkWorlds shows that organisational mess reduces apparent competence. Agensh shows that carefully designed shared infrastructure can produce cooperation without constant central allocation. The science campaign shows that discovery depends on role design, common memory and laboratory verification. Security platforms show the same pattern moving into production.

The organisational implication is not that companies should simulate everything. It is that the workplace is now part of the system under test.

### Unexpected connection

Organisation theory + multi-agent software engineering + genome mining + security operations

WorkWorlds explicitly borrows from organisation and routine theory: a workplace persists while particular performances come and go. Agensh turns coordination routines into machine-executable infrastructure. Anthropic uses a role system and common memory to search biology. Microsoft builds a common operational world for defenders and agents.

Together they make an unexpected claim: organisation design is becoming part of the AI stack. Roles, routines, shared memory and intervention rights are no longer merely the social context around the technology. They help determine what the technology can do.

PROVOCATION

### Stop certifying agents. Certify the organisation they enter.

A generic agent can pass every laboratory test and still fail because it cannot find the operative policy, sees the wrong history, collides with another agent or hands an exception to nobody. For consequential work, approval should attach to a configured human–agent institution: this version, in this role, with these permissions, routines, colleagues, monitors and remedies. The agent matters. The arrangement deserves the licence.

### What if we are right?

Opportunity. Enterprises could rehearse mixed work before exposing employees, customers or live systems to it. Persistent test organisations would reveal missing context, brittle hand-offs and unowned exceptions while changes remain cheap. Smaller models might outperform expensive ones when placed in a better-designed institution.

Organisational consequence. Model evaluation, process design, identity, HCM, enterprise architecture and operational assurance would share a common test object: an objective inside a role-visible organisational world. Leaders would buy less reassurance from one-off demos and more evidence about retrieval, cooperation, intervention and recovery.

Likely horizon. A role-sized test world can be assembled now. Persistent synthetic organisations, cross-agent protocol tests and configuration-specific certification are plausible within 12–24 months, first in software, security, legal review, finance and regulated operations.

### What would prove us wrong?

The thesis weakens if most valuable enterprise tasks remain self-contained; if retrieval and access systems make organisational context nearly frictionless; or if individual-agent benchmark performance reliably predicts production outcomes across companies.

Agensh may not transfer beyond software, where tests and merge rules make progress unusually legible. WorkWorlds is small and synthetic. Anthropic’s scientific campaign came from one laboratory and its biological result is incomplete. Platform integration can also centralise errors rather than create competence.

The model fails if “certify the organisation” becomes an excuse for exhaustive digital twins, slow committees or permanent approval. The practical test is comparative: does a persistent, role-visible evaluation predict real exceptions and outcomes better than the current curated task test?

### Optimistic possibility: organisations can practise before people pay the price

We normally discover work-design errors after launch, when a frontline employee becomes the integration layer or a customer meets the exception nobody modelled.

A small persistent test workplace creates a more generous possibility. Teams can observe how authority, context and hand-offs behave before judging either the worker or the agent. Employees can contribute real exceptions without surrendering live data. Managers can improve the system rather than counsel people to cope with it.

The prize is not a perfect simulation. It is the ability to learn about the organisation while the consequences are still fictional.

SCIENCE & DISCOVERY · OUTSIDE-IN ANALYSIS

## Scientific apprenticeship moves towards choosing what deserves a laboratory

Anthropic’s campaign produced approximately 200,000 candidate sequences before one unusual pattern became the focus of human experimental work. That ratio matters more to workforce design than the headline “AI discovery”.

ANALYSIS. When agents make search and hypothesis generation abundant, scarcity moves into selection, experimental capacity and interpretation. Junior scientists may encounter more plausible findings than they can personally reproduce. Senior judgement can become a queue.

A healthy research organisation should not let the machine monopolise anomaly detection while novices only receive polished candidates. Apprenticeship needs protected access to the messy middle: failed searches, ambiguous signals, contradictory evidence and the decision not to spend laboratory time.

The new capability is not simply prompt engineering. It is experimental triage: deciding which machine-generated observation deserves scarce physical proof and explaining why the rejected alternatives do not.

PROFESSIONAL WORK · IMPLEMENTATION SIGNAL

## Legal agents are beginning to earn authority through sampling

REPORTED FACT. Casepoint announced two agents for legal and government document review on 23 September. The company says a Relevance Agent and Issue Coding Agent use multiple models to review documents, cross-check decisions and produce recall and precision measures.

Before a full run, users review a sample and decide whether the configuration meets their threshold. Actions remain inside existing permission boundaries and audit trails. These are supplier claims; no independent comparison or customer outcome was published. [Casepoint, 23 September](https://www.prnewswire.com/news-releases/casepoint-expands-casepoint-iq-with-first-purpose-built-ai-agents-302887139.html)

ANALYSIS. The useful pattern is not “AI does legal review”. It is staged authority. Performance is tested on work drawn from the local matter, with a human accepting the threshold before scope expands.

That resembles probation more than generic certification: competence is demonstrated inside the case, under its vocabulary, permissions and error costs. The missing evidence is longitudinal—whether the agent remains fit as the matter, data and legal strategy change.

TENUOUS BUT PLAUSIBLE · Confidence: medium-low · Horizon: 12–36 months

## Agents may need a digital probationary workplace

WorkWorlds is a research infrastructure, not an enterprise deployment standard. Casepoint’s sample approval is product-specific. But together they suggest a cautious new institution.

SPECULATION. Before receiving live authority, an agent could spend time in a digital probationary workplace: a persistent, role-visible environment containing representative history, permissions, colleagues, interruptions and exceptions. It would be assessed not only on completed tasks but on how it finds evidence, shares state, handles conflict, escalates uncertainty and responds when the organisation changes.

The causal chain is plausible:

Persistent organisational simulation makes context and cooperation testable → local failure modes appear before production → authority expands on evidence rather than vendor category → human reviewers focus on changed conditions and residual uncertainty.

This could become bureaucratic theatre, and synthetic worlds may omit exactly the social signals that matter. What to watch is whether configuration-specific rehearsal predicts incidents better than task benchmarks, whether affected employees help design the cases and whether probation ends with a clear decision to authorise, narrow or reject.

LEARNING & CAPACITY · WORKFORCE SIGNAL

## Training without time is organisational make-believe

REPORTED FACT. Workera surveyed 1,000 salaried US employees in large organisations in July and reported on 23 September that 58% now receive employer-provided AI training, up from 25% in its March 2025 comparison. Yet 56% said their employer provides no working time to build AI skills.

The survey was commissioned by a skills platform, uses self-report and is limited to the United States. It should not be treated as a global prevalence measure. [Workera, 23 September](https://www.prnewswire.com/news-releases/ai-training-more-than-doubled-this-year-but-56-of-employees-report-no-time-at-work-to-build-the-skills-workera-research-finds-302887120.html)

ANALYSIS. If agent competence depends on organisational context, human learning must occur in context too. A course can explain a tool; it cannot replace time to test a real workflow, find weak evidence, negotiate a hand-off and learn when to stop.

Protected practice is not a benefit offered after transformation. It is part of the organisation’s supervisory capital. A company that allocates compute to machine rehearsal but no work time to human learning is improving only half of the system.

NOISE FILTER

## One thousand agents are not one thousand colleagues

Agensh’s scale is striking, but the agents are instances of the same model working for hours inside a purpose-built software environment. They do not have livelihoods, rights, careers or independent organisational membership.

The experiment supports a claim about concurrent machine cooperation, not a headcount comparison or a general law that more agents always produce more intelligence. Final performance remained imperfect, and coordination infrastructure carried much of the value.

Count the work, conflicts, verification and outcomes. Leave the imaginary staff photograph alone.

## Operating-model implication

### Assure the configured work system

Assurance questionEvidence to retainDecision producedPrimary steward

Can the role find the right evidence?Role-visible sources, retrieval path and provenanceImprove context, access or task placementKnowledge, identity and process owners
Can contributors cooperate without hidden collision?Claims, hand-offs, messages, conflicts and mergesChange protocol, capacity or role boundariesOperations and architecture
Can somebody challenge the trajectory?Tests, uncertainty, intervention and escalation traceExpand, narrow or pause authorityBusiness owner and independent assurance
Does the system learn safely?Persistent lessons, changed configuration and regression testsRe-authorise the arrangementJoint human–agent learning review
Are people becoming more capable?Practice time, unaided judgement and recovery skillRedesign learning or preserve manual experienceManager, profession lead and HR

The object under assurance is neither the model nor the workflow document. It is the configured work system: objective, workplace state, people, agents, permissions, cooperation rules and remedy.

## Human control watch

Assistance: a person performs the task with AI support. Control depends on whether the role can see the relevant organisational evidence and understand why it was selected.

Delegated execution: an agent works within a persistent role and hands results to people or other agents. Control depends on role-visible context, protocol design, sampling and clear exception ownership.

Autonomous cooperation: agents claim, divide, verify and merge work without continuous assignment. Control moves into the shared world, cooperation rules, monitors, consequence limits and authority to alter the protocol.

Today’s shift: the human need not allocate every task to retain control. But people must govern the institution through which allocation, evidence, challenge and recovery occur.

## Capability-model update

Gaining valueUnder pressure

Organisational-context engineerCurated prompt as proof of workplace readiness
Cooperation-protocol designerCentral orchestration of every machine action
Role-visible evaluation architectGeneric benchmark as deployment licence
Experimental-triage scientistHypothesis generation as the sole research bottleneck
Digital-probation stewardOne-off pilot in a clean environment
Protected-practice designerTraining availability without learning time

1

ONE THING

## IF I WERE TO DO ONE THING NOW

### Build one role-visible test world

□ ROLE ─── ◇ HISTORY ─── ⌁ PERMISSIONS ─── △ EXCEPTION ─── ✓ AUTHORISE?

Test whether the work system can find, cooperate and recover—not merely whether a model can answer.

Take the agent-assisted workflow whose management loop you timed in the previous exercise—or another live candidate—and assemble one small, reusable test world this week: the role’s actual permissions, three to five representative artefacts, one prior event that changes the current answer, one awkward exception and the acceptance and stop criteria. Run the same agent through it without pointing to the relevant evidence, then review the trace with the frontline user and business owner: what it found, missed, shared, assumed and escalated. Do not build a digital twin or procurement benchmark. Preserve this single world as a regression test. It converts the newly visible management work into an organisational fitness check and creates evidence for whether authority should expand, narrow or wait.

## Mental-model update

The permeable enterprise has acquired earned authority, an executable constitution, a capability border, a memory for failure, a limit on consequence, an accountability surface and a visible management layer.

Today it acquires a rehearsal space.

Elastic autonomy does not depend only on finding a better agent. It depends on building an organisation in which useful context is reachable, cooperation is legible, challenge has authority and both people and machines can learn before consequences become real.

The emerging North Star is a permeable enterprise that can test itself: capability may enter and self-organise, but the configured institution must demonstrate that it can turn machine action into responsible work.

## Questions for the executive table

- Which agent passed a demonstration only because somebody quietly selected the right context first?

- What managerial decisions are now encoded in your claim, routing, verification and merge protocols—and who may change them?

- Can you reproduce the role-visible world in which a consequential agent was approved six months ago?

- Where will human specialists learn to reject attractive machine hypotheses before scarce operational or laboratory capacity is committed?

- Do employees receive protected time to improve the mixed system, or only training content and higher output expectations?

Evidence note. WorkWorlds and Agensh are research results in constrained, partly synthetic settings. Anthropic reports on its own agents and laboratory; the enzyme system’s function is unresolved. Microsoft, Cisco, Casepoint and Workera describe their own products or sponsored research. Cisco and Workera survey findings are self-reported and should not be treated as audited prevalence or performance.

The concepts organisational competence, protocol management, organisational-competence stack and digital probationary workplace are original AyEye analysis, not claims made by the cited sources.

Canonical: https://www.lecxie.com/publications/ayeye-workforce/2026-09-24.html
