# The human work left behind is being turned into a score

As routine call-centre work contracts, enterprise agents are acquiring persistent roles and one employer is using AI to score every sensitive customer call. The risk is a relational paradox: machines remove standard work, people inherit the hardest exceptions and then those human encounters are judged as though they were standard production.

**AyEye — Workforce Management · 2026-10-09**

The more routine work a machine absorbs, the more human work is likely to involve distress, ambiguity, exception and trust. Those are not the leftovers of a process. They are often the point at which an institution becomes human.

Yet a new form of management is arriving at exactly that boundary. One employer is using AI to score every sensitive customer call across more than 50 criteria. At the same time, global call-centre employment is contracting as enterprise agents acquire persistent identities, memories and roles. The danger is not simply that machines replace people. It is that they remove the standard work, leave people with the hardest conversations and then judge that human remainder as though it were standard production.

## The executive brief

- AI monitoring has crossed from sampling work to continuously interpreting it. The Guardian reports that Co-op Legal Services records and scores every call made by some probate advisers, using more than 50 criteria in performance analysis. The Co-op says the system supports quality, coaching and consistently empathetic service and does not make decisions.

- The entry point into office work is already narrowing. Revelio Labs finds global call-centre headcount has fallen year on year for eight consecutive quarters. In 12 middle-income countries, entry into the occupation has declined much faster than exit; only 10.8% of observed leavers reached three higher-paid, growing occupational groups.

- Enterprise agents are gaining role-like persistence. Google has announced a universal Gemini agent that can hold memory, respond to events, create sub-agents, use shared skills and tools and operate through its own identity as a team or role-based coworker.

- The public legitimacy gap is large. In an ICO survey of 2,157 UK adults, only 17% regarded AI workplace monitoring or performance evaluation as acceptable, while 62% considered it unacceptable.

- The executive issue is not whether conversations may be analysed. It is whether an inference about empathy is allowed to become a fact about a person without a declared purpose, evidence of validity, contextual review and a practical right to challenge it.

## The human work left behind is being turned into a score

ORIGINAL SYNTHESIS · Confidence: medium-high · Horizon: now to 24 months.

### Signal one: a machine is scoring empathy in conversations about death

REPORTED FACT. On 8 October, the Guardian reported that Co-op Legal Services is using an OpenAI model to record, analyse and give percentage scores to every phone call made by some advisers handling probate, wills and estates. The model assesses more than 50 aspects of a conversation. Managers use the results in analysing performance, and the system is also understood to identify ways to improve sales performance.

A worker described the monitoring as oppressive and said it created distrust. The Co-op said it did not recognise that criticism, that advisers were highly engaged and that the system is a support tool rather than a decision-maker. It said reviewing all conversations helps leaders coach colleagues and provide consistently expert and empathetic guidance, while human judgement and accountability remain central. This is reporting about one implementation, not evidence that automated scoring is accurate or that the reported employee experience is universal. [The Guardian, 8 October 2026](https://www.theguardian.com/technology/2026/oct/08/co-op-legal-services-ai-customer-phone-call-surveillance)

ANALYSIS. Quality assurance traditionally samples a small amount of work because human attention is scarce. AI makes total observation economically possible. That changes the nature of the control. A sample asks whether practice is broadly safe and effective. Continuous interpretation can define what normal performance is, which words count as care and which deviations deserve attention.

The distinction between “support” and “decision” is therefore necessary but incomplete. A score can shape coaching, reputation, confidence, promotion and the cases a manager assigns without ever issuing a formal decision. The managerial effect can be real even when the system is technically advisory.

### Signal two: the routine service ladder is shrinking

REPORTED FACT. Revelio Labs published an analysis on 6 October of global call-centre employment and career transitions. It reports that the only eight year-on-year declines in 70 quarter-end observations since 2009 are the most recent eight quarters, with the latest the steepest. The pattern began in high-income countries and spread down the income ladder.

The authors explicitly do not observe which firms replaced people with chatbots and say the decline could also reflect post-pandemic correction or cost pressure. They argue that the divergence from other office work after ChatGPT makes an AI contribution difficult to dismiss.

Across 12 middle-income countries, entry into call-centre work fell from 25.1 to 17.2 per 100 workers since November 2022; exit also fell, but less, from 22.3 to 18.5. The occupation is contracting mainly because fewer people enter it, not because separations suddenly accelerated. Software and data, customer success and technical support grew and paid more, but only 10.8% of observed call-centre leavers entered those three groups. The analysis is observational and based on Revelio's workforce data; it does not establish causation or represent every labour market. [Revelio Labs, 6 October 2026](https://www.reveliolabs.com/news/ai-and-work/after-a-decade-of-growth-global-call-center-employment-is-shrinking)

ANALYSIS. The change is quiet because it happens through a missing invitation. A young person is not dismissed from the first office job; the opening never appears. A service operation does not announce the removal of a career ladder; it hires fewer people at the bottom and keeps more complex work for those already inside.

That matters to measurement. As standard queries move to machines, human conversations become a progressively biased sample: more unusual, more emotional, more likely to involve a failure elsewhere in the system. Comparing those calls with historic averages or uniform scripts can punish precisely the people absorbing the organisation's hardest exceptions.

### Signal three: the agent is acquiring a role while the person acquires an exception queue

REPORTED FACT. Google Cloud announced a universal Gemini agent on 8 October. It says the agent can receive objectives rather than step-by-step instructions, operate across business applications, run persistently in the cloud, respond to events and create temporary sub-agents with their own identities. A coworker agent can have a defined role, email address, calendar, storage and company-directory presence.

Google also describes registries for reusable tools and skills, four kinds of memory and procedural knowledge that the agent can write for itself. It says customers have already built thousands of custom agents, including more than 1,000 at Orange Spain and more than 10,000 at SOMPO. These are supplier descriptions and customer claims; the announcement does not provide independent evidence about net employment, error rates or the distribution of benefits. [Google Cloud, 8 October 2026](https://cloud.google.com/blog/products/ai-machine-learning/welcome-to-gemini-at-work-2026)

ANALYSIS. This is more than another assistant. The platform creates a machine participant with continuity, role, memory and access to organisational process. As it handles the repeatable centre of work, people are likely to receive a different portfolio: judgement calls, distressed customers, disputed facts and events that do not fit the procedural memory.

Organisations must therefore evaluate human service against the case mix people now receive, not the case mix that existed before automation. Otherwise, an agent improves the easy denominator while the employee appears to deteriorate against a harder numerator.

### Signal four: legitimacy is already weaker than technical possibility

REPORTED FACT. The Information Commissioner's Office surveyed 2,157 UK adults in March 2026. Seventeen per cent found AI use in workplace monitoring or performance evaluation acceptable; 62% found it unacceptable, including 37% who considered it completely unacceptable. The survey measures public attitudes, not the legality or effectiveness of a particular system. [ICO AI Omnibus findings, March 2026](https://cy.ico.org.uk/media2/zeunavqg/ai-omnibus-summary-of-findings-2026.pdf)

The ICO's worker-monitoring guidance says employers must define the purpose, minimise data and ensure that information is accurate rather than misleading. It warns that analytical tools can make incorrect inferences and says workers should be able to see, explain and challenge monitoring results, particularly when they inform performance reviews. [ICO worker-monitoring guidance](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/employment/monitoring-workers/data-protection-and-monitoring-workers/)

ANALYSIS. Trust will not be repaired by explaining that the model merely advises a manager. Workers understand that managerial attention is scarce. What an automated system chooses to surface, repeat and compare can shape the manager's view long before a formal decision.

### The unexpected connection

Continuous empathy scoring + a shrinking service-entry ladder + persistent role-based agents + low public acceptance

These developments appear to concern quality assurance, labour economics, enterprise software and privacy. Together they reveal a change in the composition of human work. Automation is taking the work most easily expressed as a rule. Humans are inheriting the work created when rules meet life. At the same time, machines are being asked to standardise how that remaining humanity should sound.

ORIGINAL SYNTHESIS. Call this relational compression: a rich interaction is progressively reduced from lived circumstances, to a transcript, to observable markers, to a score, to a judgement about a worker. Each step can be useful. Each also discards context.

What exists in the workWhat the system retainsWhat can disappear

A person with a need
Grief, uncertainty, history and an outcome that matters.A recorded interaction linked to a case.What happened before the call and what the person could not articulate.
A conversation
Pacing, silence, repair, professional judgement and mutual adaptation.Audio, transcript and behavioural features.Why deviation from a script was appropriate.
An interpretation
A view about empathy, clarity, risk or effectiveness.Criteria, classifications and confidence.Disagreement about what good care required in this case.
A management action
Coaching, case allocation, recognition or formal review.A score, comparison and audit trail.The point at which an inference became consequential.

PROVOCATION: The more human the work becomes, the less defensible a universal productivity score becomes.

An empathy score may detect missed disclosures, interruptions or patterns worth reviewing. It cannot settle what empathy was required, whether the employee inherited an impossible policy or whether the customer needed speed rather than warmth. If the measure becomes a target, employees learn to produce the audible tokens of care. The organisation may improve the score while making the conversation feel less human.

### What if we are right?

Opportunity. Continuous analysis could find genuine coaching needs and systemic failure much earlier than occasional sampling. It could show that a product, policy or agent repeatedly hands distressed customers to people without adequate authority. Used as an inquiry signal, the technology can make invisible emotional load visible and improve both service and work.

Organisational consequence. Service analytics would stop at the boundary between evidence and judgement. Process owners would examine case mix and upstream causes. Frontline employees would be able to see and contest interpretations. Managers would remain responsible for conclusions and for correcting the system that created the difficult interaction.

Horizon. A challenge right and case-mix review can be introduced now. Reliable organisation-wide measures of relational work may take far longer, and some qualities may remain unsuitable for individual scoring.

### What would prove us wrong?

The argument weakens if continuous conversation analysis produces demonstrably more consistent and fair coaching than human sampling; if scores remain low-stakes diagnostic signals; if workers understand, influence and trust the criteria; and if independent validation shows that scores remain accurate across accents, cultures, case types and levels of customer distress.

It also weakens if automation does not change the difficulty of cases reaching people. The practical test is longitudinal: compare case mix, score distribution, worker wellbeing, customer outcomes and successful challenges before and after automation. If harder cases do not concentrate around humans and the measures do not alter behaviour, relational compression is not the binding concern.

### A constructive possibility: the metric becomes a question

The humane alternative is not to stop learning from conversations. It is to reverse the direction of authority. Instead of saying, “The score shows that you lacked empathy,” the manager asks, “The system noticed this pattern; what was happening, and what should we change?”

That small reversal turns monitoring into joint diagnosis. A worker can identify the policy that forced an unhelpful answer, the agent hand-off that lost context or the script that sounded uncaring. The organisation learns about itself rather than merely ranking the person who faced the customer.

## The entry ladder is disappearing before the bridge exists

LABOUR MARKET · OUTSIDE-IN ANALYSIS.

The Revelio evidence does not show mass redundancy. It may reveal something strategically harder to notice: erosion by non-entry. An employer can preserve current service levels while gradually removing the jobs through which people learned how office work, customers and institutions function.

The growing adjacent occupations offer a possible destination, but the transition data suggest that a destination is not a bridge. Customer success requires commercial judgement; technical support requires system knowledge; software and data require capabilities that a call-centre worker may not be given the time or evidence to develop.

This extends the [capability corridor](https://www.lecxie.com/publications/ayeye-workforce/2026-10-06.html). Organisations should not wait for a disappearing role to produce a redundancy pool. They can use live work now to create supervised passages into technical support, service design, quality investigation and agent operations. The question is not whether the old job should be preserved forever. It is whether the institution will preserve an accessible first step into the better work it is creating.

## A universal agent is becoming part of the management system

ENTERPRISE TECHNOLOGY · OPERATING SIGNAL.

Google's announcement gives an agent many attributes organisations traditionally use to make work governable: a role, an identity, permissions, memory, tools, skills and an audit trail. It can exist for a person, a team or a function and keep working after a laptop closes.

The missing object is not another org-chart box. It is a shared account of how the agent changes the human portfolio around it. Which routine cases disappear? Which exceptions accumulate? Who inherits emotional load? Which employee becomes the unofficial translator, corrector or conscience for the machine?

An agent register that records only owner, permissions and cost will miss this redistribution. The [Workforce Context Architecture](https://www.lecxie.com/publications/ayeye-workforce/2026-09-30.html) should connect the agent to the tasks it absorbs, the cases it escalates, the roles it changes and the outcomes for customers and workers. Otherwise, the technology estate will know what the agent can access while the organisation remains blind to what it has made people carry.

## The empathy score could create synthetic emotional labour

TENUOUS BUT PLAUSIBLE · Confidence: medium-low · Horizon: 12–36 months.

HYPOTHESIS. If relational qualities become individual performance metrics, workers may learn to manufacture machine-recognisable empathy: approved phrases, predictable pauses and visible acknowledgements optimised for the monitor rather than the person.

The causal chain is plausible but unproven:

Routine cases move to agents → human cases become harder and more emotional → automated scoring expands because human review cannot cover every call → scores enter coaching and performance routines → employees adapt behaviour to what the model recognises → measured empathy rises while authentic discretion narrows.

What to watch: whether scores affect pay, promotion or case allocation; whether workers receive the criteria; whether challenges change the record; whether customer outcomes improve independently of the score; and whether language becomes more uniform after measurement begins.

## Do not confuse total observation with complete understanding

NOISE FILTER.

Reviewing every call can produce more evidence than sampling. It does not produce omniscience. A transcript may be incomplete, a model may misread an accent, a criterion may reward the wrong behaviour and a manager may treat a percentage as more certain than it is.

Nor should criticism of monitoring become a defence of unreviewable work. Customers deserve safe, lawful and humane service. Employees deserve useful feedback. The discipline is to make the evidence proportionate to the consequence: automated detection may be broad; adverse judgement should become narrower, more contextual and more contestable.

## Operating-model implication

### Separate service assurance from judgement about the person

LayerPrimary questionControl that must exist

PurposeWhat customer or regulatory outcome justifies analysis?One declared use; new uses require a fresh decision.
SignalWhat can the model reliably notice?Validation by accent, case type, channel and severity.
ContextWhat made this interaction unusually difficult?Case-mix and upstream-system evidence beside the score.
JudgementWho decides what the evidence means?A trained manager who can reject the inference and records why.
ChallengeCan the worker inspect and correct the account?A prompt, non-retaliatory route that can change the record.
LearningDoes the pattern point beyond the individual?Routine review of policy, workflow, agent hand-offs and workload.

The design principle is simple: use broad machine attention to find questions, then increase human context as the consequence for a person rises.

## Human control watch

Assistance: AI summarises a conversation or highlights a possible coaching moment. The employee and manager can inspect the underlying evidence, and no individual score persists by default.

Delegated execution: AI classifies every interaction, creates comparative scores and routes exceptions. Human control depends on validated criteria, case-mix adjustment, a visible challenge route and managers who have time and authority to disagree.

Autonomous control: scores alter work allocation, rewards, discipline or employment opportunity without meaningful review. A human name on the process is not sufficient if the decision is practically determined upstream by the model and target.

Today's shift: human control is moving from occasional observation of work to governance of the interpretation layer between work and managerial judgement.

## Capability-model update

Gaining valueUnder pressure

Relational-work designerUniform script as a proxy for care
Case-mix analystRaw score comparisons across unequal work
Monitoring-rights stewardBlanket consent in an employment relationship
Agent-to-human load mapperAutomation benefit measured without exception transfer
Inference challenge facilitatorAdvisory score treated as neutral fact
Accessible career-bridge designerReskilling offer without a route into live work

## IF I WERE TO DO ONE THING NOW

### Contest one score

1 · ONE THING

Take one AI-generated quality or performance score already used in a live workflow this week and review one real case with the person whose work it describes, their manager and the process owner. Ask what the score retained, what context it lost and whether a reasonable disagreement can change the record or the action. Do not launch a monitoring programme or redesign the scorecard. Resolve one contested interpretation and document the route others can use. This builds on the previous capability-claim annotation by testing whether evidence remains challengeable when it becomes managerial judgement.

## Mental-model update

The permeable enterprise has acquired an architecture for context, inherited digital capability, a boundary for operational authorship, a corridor for human transition and a distinction among inferred, demonstrated and authorised capability.

Today it acquires a limit on interpretation.

More observation does not automatically mean more understanding. As agents absorb the repeatable centre of work, the organisation must protect the ambiguous human edge from being compressed into a falsely precise score. Evidence may travel widely; judgement must return to the context, person and consequence.

The emerging North Star is an enterprise in which machines can notice at scale, people can challenge with dignity and no measure of humanity becomes more authoritative than the human relationship it was meant to improve.

## Questions for the executive table

- Which quality of human work are we now scoring without a shared definition of what good looks like?

- Has automation changed the difficulty and emotional weight of the cases that reach people?

- Can an employee see, explain and correct an AI inference before it affects coaching, opportunity or discipline?

- Which entry-level role is quietly shrinking, and what live-work bridge exists into the occupations that are growing?

- When an agent gains a role, identity and memory, who measures the additional exception and relational work inherited by the team?

Evidence note. The Co-op implementation is described in Guardian reporting and includes an anonymous worker's account and the company's response; no independent validation of the model or employee-experience prevalence is available. Revelio Labs uses observational workforce data and cannot identify firm-level substitution by AI. Google's capabilities, customer volumes and benefits are supplier and customer claims, not independent outcome evaluations. The ICO survey measures public attitudes and its guidance is not a legal opinion about the reported implementation.

The concepts relational compression, synthetic emotional labour and the proposed separation of service assurance from judgement about the person are original AyEye analysis, not claims made by the cited sources.

Canonical: https://www.lecxie.com/publications/ayeye-workforce/2026-10-09.html
