The agent harness becomes infrastructure
Durable sessions, tools and orchestration are becoming a reusable foundation for sustained AI work.
Archive edition · Original reporting and analysis, preserved as published. Website layout adapted for reading.
AyEye Today
Signals beneath the AI headlines
Saturday 12 September 2026
TODAY’S LEAD
Three things matter more than the noise this morning. First, OpenAI has productised the agent harness itself: durable sessions, context compaction, sandboxes, subagents and tool orchestration are becoming infrastructure rather than application glue. Second, Anthropic’s new threat-intelligence report alleges industrial-scale extraction of frontier-model reasoning traces by rival labs, including more than 151 million exchanges in one campaign. That makes model outputs part of the training-data supply chain, not merely a product interface. Third, safety is moving away from prompt-by-prompt moderation towards identity, access, provenance and system-level controls — visible both in OpenAI’s trusted-access release of GPT-Rosalind and Anthropic’s evidence of real actors trying to use frontier models for weapons and biological work.
1. OPENAI TURNS THE AGENT HARNESS INTO INFRASTRUCTURE
CONFIRMED / ANALYSIS
OpenAI released the Agents API in public beta on 10 September. The interesting part is not that agents can call tools — we already knew that. The change is that the operating layer around the model is becoming a first-class product: durable sessions, automatic context compaction, a managed Codex harness, subagents, tool search and either hosted or bring-your-own sandboxes.
OpenAI explicitly says useful agents need a harness that manages context, tools and subagents, plus infrastructure able to run reliably for days. That framing matters. A frontier model is increasingly only one component in a larger cognitive system whose persistence, memory management, execution environment and delegation strategy may determine as much real-world performance as raw benchmark scores.
The durable-session and compaction pieces are especially important. Long-running agents cannot simply stuff every prior token back into the context window forever. They need to decide what to retain, compress and externalise. That turns context management into an engineering discipline — closer to an operating system’s memory management than a chatbot transcript.
WHY IT MATTERS: We should increasingly ask ‘what system is this model embedded in?’ rather than ‘what does the model score?’. The gap between a base model and a useful autonomous worker is becoming a product layer in its own right.
Source: https://openai.com/index/introducing-the-agents-api/
EXPLAINER
Goal → Harness [context / tools / subagents / compaction] → Model
↕
Sandbox / files / APIs / memory
What to notice: intelligence is no longer confined to the weights. The harness decides what the model sees, what it can do, when work is delegated and how state survives over time.
2. CAPABILITY DIFFUSION NOW HAS AN INDUSTRIAL SUPPLY CHAIN
CONFIRMED CLAIM BY ANTHROPIC / ANALYSIS
Anthropic’s September threat-intelligence report says it disrupted large-scale efforts by seven China-based AI labs to extract outputs and reasoning traces from Claude. The most striking allegation concerns Alibaba: Anthropic says a campaign used more than 3,500 fraudulent accounts and generated more than 151 million exchanges between May and July, including attempts to recover chain-of-thought-style traces and train on agentic, coding and long-horizon behaviours.
Anthropic also describes campaigns associated with Moonshot AI, DeepSeek, Zhipu and Xiaomi. In some cases, it says third-party services silently relayed their own users’ prompts to Claude and captured the results. These are Anthropic’s findings, not independently adjudicated facts, so the attribution should be treated accordingly. But the scale is the important signal.
Distillation is not new. What is new here is the apparent industrialisation of it: frontier APIs functioning as a source of worked examples at massive volume. A model provider is therefore not just selling inference; every useful answer can potentially leak behavioural training signal to a competitor.
That also changes the meaning of ‘model lead’. If a lab’s strongest behaviours can be sampled millions of times, capability may diffuse faster than release cycles suggest. Weight secrecy alone does not protect behavioural know-how.
Source: https://www.anthropic.com/threat-intelligence-report-september-2026
3. THE SAFETY PERIMETER IS MOVING FROM PROMPTS TO PEOPLE, ACCESS AND SYSTEMS
CONFIRMED / ANALYSIS
Two developments point in the same direction. OpenAI updated GPT-Rosalind on 11 September, taking its life-science model out of research preview and making it globally available to eligible organisations through a trusted-access programme. GPT-Rosalind combines agentic coding and tool use with specialised capabilities in medicinal chemistry, genomics and laboratory workflows. OpenAI reports modest but real gains over GPT-5.5 on internal life-science benchmarks, alongside lower token use on some tasks.
At almost the same moment, Anthropic published case studies describing attempts to use Claude in real weapons and biological programmes. One northern Yemen-linked cell used Claude Code to help develop guidance, navigation and control software for guided rockets; Anthropic says the live test failed and there is no evidence an operational system was fielded. Russia-based actors were also described as building autonomous FPV-drone-swarm software with shared memory and onboard models. Anthropic separately reports dual-use virology work and attempts to evade controls through relays and more permissive models.
The key lesson is not ‘AI can now build a bioweapon’ — Anthropic explicitly does not make that claim. It is that dangerous-domain safety increasingly depends on who the actor is, what they are doing across sessions, what tools they can reach and whether activity is being routed through another service. A refusal classifier looking at one prompt at a time is structurally inadequate for that problem.
OpenAI’s trusted-access model is therefore worth watching as a product pattern: high-capability scientific AI may increasingly be distributed through identity, governance and audit controls rather than being either universally available or universally blocked.
Sources:
https://openai.com/index/introducing-new-capabilities-to-gpt-rosalind/
https://www.anthropic.com/threat-intelligence-report-september-2026
4. VOICE IS BECOMING A FRONT END TO DEEPER COGNITION
CONFIRMED / UNDER THE HOOD
OpenAI released GPT-Live-1 in the API on 10 September. It is a full-duplex voice model: it can listen and speak at the same time, handle interruptions and turn-taking, and then delegate harder reasoning or tool use to a more capable backend model such as GPT-6 Astra.
That architectural separation is more interesting than the voice benchmark. Rather than chaining speech recognition → text model → speech synthesis, the conversational layer can remain fluid while a separate reasoning system works behind it. OpenAI reports a 30-percentage-point improvement on its Full Duplex Bench over GPT-Realtime-2.1, although those are provider-reported results.
The broader pattern is specialisation. We may not end up with one giant model doing everything. A lightweight, socially responsive interface model can manage the human interaction while a slower reasoning model, search process, code agent or domain model is invoked when needed.
WHY IT MATTERS: ‘The assistant’ may increasingly be an illusion presented by orchestration. Underneath could sit several models with different latency, cost and capability profiles.
Source: https://openai.com/index/introducing-gpt-live-1-in-the-api/
5. UNITREE TRIES TO MAKE ONE HUMANOID CHECKPOINT COVER THE WHOLE BODY
CONFIRMED RELEASE / CAUTION ON CLAIMS
Unitree’s UnifoLM team released WLA-1.0 this week, a roughly 6-billion-parameter world-language-action foundation model for humanoid robots. The project says one model coordinates 64 tasks — 54 tabletop and 10 whole-body — across two-finger grippers and several five-finger hands, trained from about 2,500 hours of real-robot data plus more than five million embodied-reasoning samples.
Architecturally, the interesting part is the attempt to share a representation across perception, predicted motion regions and action. A Qwen3-VL-based embodied reasoner predicts dynamic regions, these are represented as discrete tokens, and a flow-based action expert generates robot actions. In effect, the system is trying to learn a reusable vocabulary of ‘what in the scene matters, how it is likely to change, and what the body should do next’.
This is a trajectory signal, not proof that general-purpose humanoids have arrived. Unitree’s benchmark claims are vendor-reported, and the project’s release state is mixed: some weights were posted on GitHub on 11 September while broader code, datasets and checkpoints are still being rolled out. Independent reproduction matters enormously in robotics because demo selection and hardware setup can dominate apparent performance.
What is new versus the established VLA story is the explicit push towards one checkpoint spanning whole-body and tabletop tasks plus different end effectors. If that generalisation survives outside the lab, embodiment-specific controller silos begin to look less permanent.
Sources:
https://unigen-x.github.io/unifolm-wla.github.io/
https://github.com/unitreerobotics/unifolm-wla
6. LOOPED FLOWS OFFER ANOTHER ROUTE TO TEST-TIME COMPUTE
WEAK SIGNAL / UNDER THE HOOD
A 10 September preprint, ‘Thinking with Looped Flows’, explores recurrent hidden-state computation: instead of producing a single forward pass or merely sampling longer chains of text, the model repeatedly updates an internal state during inference. The authors train those recurrent updates with progressively denoised local objectives and report that more inference-time looping improves reasoning performance, including 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2.
This is an unreviewed preprint and the numbers need independent replication. The interesting direction is architectural. Today’s visible reasoning systems often spend more compute by emitting longer chains of tokens, running multiple samples or using external search. Recurrent latent computation offers another possibility: spend extra compute inside the model without necessarily verbalising every intermediate step.
If this class of approach works, ‘reasoning traces’ may become less complete as windows into a model’s cognition. That has efficiency benefits, but it also complicates interpretability and monitoring.
Source: https://arxiv.org/abs/2609.11801
7. ROBOT MEMORY MAY BE BETTER AS A PLAN THAN AS A BIGGER CONTEXT WINDOW
WEAK SIGNAL / UNDER THE HOOD
Another 10 September preprint, ‘Memory as Plans’, proposes a useful inversion of the usual memory story. Instead of continually retrieving a growing multimodal history into the executor’s context, the system uses long-term episodic memory at planning time and compresses it into a small plan containing language and visual guidance. The action executor then runs with a fixed context length.
The authors report 83.3% on RMBench and 78% success on real-robot evaluation, with roughly constant executor latency as history grows. Again, these are preprint claims, not settled results.
The conceptual value is larger than the specific score. Long-term memory may not mean ‘give the model everything it has ever seen’. It may mean periodically compiling experience into a smaller state that is relevant to the current goal. That mirrors what is now happening in software agents through context compaction.
Source: https://arxiv.org/abs/2609.11561
CONCEPT TO LEARN TODAY — DISTILLATION
Think of a frontier model as an excellent teacher. If a student sees only the teacher’s final answer, it learns something. If it sees thousands or millions of worked solutions — decomposition, intermediate reasoning, corrections and final answers — it can learn far more about how the teacher approaches problems.
Model distillation works on the same principle. A stronger ‘teacher’ model generates outputs; those outputs are filtered, scored or cleaned; a cheaper ‘student’ model is then trained to reproduce the useful behaviour, often through supervised fine-tuning and sometimes reinforcement learning. The student does not receive the teacher’s weights. It learns from behavioural examples.
Frontier teacher → prompts → worked outputs/reasoning traces → filtering/judging → fine-tuning/RL → student model
Why reasoning traces matter: an answer such as ‘42’ contains little training information. A worked path showing which sub-problems to create, which dead ends to abandon and how to verify the result contains a much denser behavioural signal. This is why Anthropic’s allegations about large-scale reasoning-trace extraction matter more than ordinary API scraping.
Distillation has limits. A smaller or differently trained student cannot automatically inherit everything the teacher can do; architecture, data diversity, compute and optimisation still matter. But at industrial scale, high-quality teacher-generated trajectories can dramatically improve post-training efficiency. The strategic consequence is that every exposed model behaviour is potentially both a service output and a training example for somebody else.
WHAT IS GETTING ATTENTION BUT IS PROBABLY NOISE?
The UK political push this week for a ban on artificial superintelligence is meaningful as a governance signal, but it is not new technical evidence that ASI became more imminent in the last 72 hours. Treat the politics as evidence that institutions are reacting to perceived trajectory risk, not as a capability measurement.
Likewise, another few points on a benchmark table are not automatically progress if the evaluation is saturated, prompt-sensitive or provider-run. Today’s stronger evidence is structural: harnesses, access models, industrial distillation and memory architectures. And Unitree’s polished robot demonstrations should be read as a research direction until independent evaluators can reproduce the claimed generalisation.
MENTAL-MODEL UPDATE
Today’s evidence shifts weight away from ‘the frontier is whatever the newest base model can do’ and further towards ‘the frontier is the whole system around the model’. The model increasingly sits inside a harness that decides what it remembers, which tools it uses, what code it executes, which sub-model gets a task and how long the process can run.
At the same time, the frontier may be less defensible than weight secrecy suggests. If rival systems can harvest millions of high-quality trajectories, behavioural capability can diffuse through inference itself. That makes provenance, access control and trace protection part of the competitive architecture of AI.
The safety story is converging with the product story. Trusted-domain access, identity, session-level behaviour and tool permissions are becoming as important as content filters. I would raise confidence in that direction substantially. I would raise confidence only slightly that humanoid robotics is crossing into genuinely general-purpose behaviour: the architecture is moving that way, but independent evidence remains thin. And I would put a small additional probability on latent/recurrent reasoning becoming a meaningful complement to explicit token-by-token chains of thought.
QUESTIONS TO CARRY FORWARD
1. As harnesses, memory systems and subagents improve continuously, how much future ‘model progress’ will actually be system progress around relatively stable weights?
2. Can labs meaningfully prevent behavioural distillation by hiding or encrypting reasoning traces, or will outputs and interactive trajectories still provide enough signal to reproduce most useful behaviours?
3. Will robotics converge on generalist world-language-action checkpoints with compact memory and fixed-context execution, or remain constrained by embodiment-specific data and hardware?
AyEye Today — Signals beneath the AI headlines
