Dominic Chiappe · People, capability & transformation

Thinking about how organisations perform in an AI-enabled world

AyEye Today · 2026-09-11

AI Intelligence Briefing — 11 September 2026

The original briefing explores how specialised components allocate intelligence against cost, time and risk.

Archive edition · Original reporting and analysis, preserved as published. Website layout adapted for reading.

AI Intelligence Briefing — 11 September 2026

Executive view

Yesterday we ended with a picture of AI as an artificial cognitive organisation: not merely a neural network, but a model surrounded by agents, memory, procedures, tools and verification.

The last 24 hours make that picture considerably more concrete.

OpenAI has effectively productised the agent harness that sits around Codex. DeepSeek has redesigned the model itself around the economics of long-running agents. Cognition is explicitly training models to optimise intelligence against cost, rather than intelligence alone. OpenAI has separated conversational responsiveness from deep reasoning in its new voice architecture. Meta, meanwhile, is separating the agent that acts from a separate agent that governs whether those actions are allowed.

These are different companies solving different problems, yet they are converging on one architectural idea:

future AI is unlikely to be one enormous model doing everything.

It looks increasingly like a heterogeneous cognitive system in which different components handle perception, fast interaction, deep reasoning, memory, parallel work, action, verification and control — while compute is dynamically allocated according to the value of the task.

That is today's main update.

There is a second, less comfortable development. Anthropic says Chinese AI laboratories have harvested hundreds of millions of Claude interactions to improve their own systems. If the numbers are correct, frontier models are increasingly becoming training infrastructure for their competitors.

So there are two forms of AI acceleration operating simultaneously:

AI systems are becoming more sophisticated internally, while their capabilities are diffusing externally far faster than model weights need to be stolen.

That combination could matter enormously.

1. OpenAI has turned the harness into a product

Confirmed — 10 September.

OpenAI yesterday released the Agents API in public beta.

On the surface this looks like another developer API announcement. I think it is substantially more important.

OpenAI is exposing essentially the same orchestration machinery that sits behind Codex: long-lived sessions, automatic context management, sandboxes, tool selection, parallel subagents and the ability for agents to continue working across multiple context windows for hours or days.

The significant sentence in OpenAI's announcement is not about Astra. It is this:

useful agents need a powerful harness.

The API automatically compacts old context as the context window fills, dynamically loads only relevant tool definitions, lets agents run tool calls programmatically rather than dumping all results back into the language model and allows a coordinating agent to spawn subagents with their own independent contexts.

That sounds plumbing-like. It isn't.

What is new relative to what we already knew

We've already established that capabilities increasingly arise from model + orchestration.

Yesterday that was principally an analytical conclusion.

Today OpenAI is effectively saying:

yes — and we are now treating that orchestration layer as part of the intelligence product itself.

The implication is that simply giving two companies access to the same underlying model may no longer give them anything close to the same capability.

One might have:

model + basic prompt loop.

Another:

model + intelligent context compression + 40 tools + subagents + code execution + persistent workspace + retrieval + verification.

Same neural network. Very different effective intelligence.

This also creates an intriguing competitive question: does the harness eventually become more strategically important than the model API?

My interpretation: OpenAI increasingly sees agent runtime as something analogous to an operating system for cognition.

2. DeepSeek has redesigned the transformer around the economics of agents

Confirmed — 10 September.

DeepSeek released V4.1-Flash yesterday.

Ignore the benchmark horse race for a moment. The architecture is the interesting part.

It is a 552-billion-parameter mixture-of-experts model, but DeepSeek says only about 8 billion parameters are active while processing input and 16 billion while generating output. It uses a new architecture DeepSeek calls Causal Encoder–Decoder.

More revealingly, DeepSeek says its KV cache requires only one quarter of the HBM memory and one eighth of the SSD storage of its previous generation.

DeepSeek explicitly identifies agent workloads as a reason this matters.

Why?

Agents repeatedly send enormous prefixes back into models: instructions, conversation, files, tool outputs, intermediate reasoning state, previous actions and retrieved information.

That accumulated context has to live somewhere.

At sufficient scale, the constraint isn't simply floating-point arithmetic. It becomes moving and storing memory.

This is an important under-the-hood shift.

The standard transformer architecture was developed principally for predicting the next token.

The workload now being optimised looks rather different:

read an enormous working state → reason → call tool → receive result → preserve state → reason again → repeat hundreds of times.

Model architecture is starting to adapt to the behaviour of agents.

That's a useful signal because it suggests agentisation is no longer merely software wrapped around LLMs. It is beginning to influence the neural architecture underneath them.

3. Cognition has started training economics into reasoning

Confirmed — 10 September.

Cognition released SWE-2, its new coding model behind Devin.

The training method is much more interesting than benchmark rankings.

SWE-2 is built by post-training Moonshot's 2.8-trillion-parameter Kimi K3 and represents Cognition's first reinforcement-learning run at this scale.

But instead of simply rewarding whether the agent solved the task, Cognition trains against whether it solved the task and how much the solution cost.

Its simplified reward function is:

reward = success − cost penalty.

It trains several reasoning-effort levels simultaneously, creating different points on a cost/performance curve.

That produced an interesting behavioural result. Compared with SWE-1.7, Cognition says SWE-2 at medium effort achieved a higher score while taking 58% fewer agent turns and costing 81% less. The older model tended to wander through codebases gathering enormous amounts of information before acting. The new model learns to decide what deserves investigation.

That distinction is subtle but important.

Intelligence isn't merely: can I eventually solve this?

It is increasingly: how much cognition should I spend before acting?

Humans do this constantly.

You don't perform three hours of research before deciding whether you want tea or coffee. But you might before investing £100,000.

AI systems are beginning to acquire something resembling that metacognitive resource allocation.

Cognition also describes a recursive training flywheel in which earlier SWE-2 checkpoints generate rollouts that expose weaknesses in its verifiers, which are then hardened for subsequent training.

Not recursive self-improvement in the strong sense — humans still design and operate the process — but another example of AI-generated experience feeding the machinery that improves subsequent AI.

4. OpenAI's new voice architecture reveals something more interesting than better voice

Confirmed — 10 September.

OpenAI also released GPT-Live-1 to developers.

We covered full-duplex conversation previously, so there is no value revisiting the fact that an AI can now listen while speaking.

The new architectural detail is more important.

GPT-Live-1 can handle the continuous realtime conversation, while delegating difficult reasoning, research and action to another model and agent system behind it.

So imagine saying:

“Find out why our project is slipping and tell me what we should do.”

The voice model doesn't necessarily stop talking while a giant reasoning model disappears for 40 seconds.

Instead fast conversational intelligence continues interacting with you while slower deep intelligence works asynchronously behind it.

This starts looking surprisingly like a computational version of System 1/System 2 cognition — although I would not push the psychological analogy too literally.

The significant engineering principle is: different intelligence for different latency requirements.

That is another departure from the assumption that one increasingly capable general model should do everything.

We may instead get speech/perception model → router → cheap reasoning model for routine questions → powerful reasoning model when needed → specialist agents → tools and computers.

The user experiences one intelligence. Underneath it may be a small organisation of models.

5. Meta has separated the agent from the agent that controls it

This appeared on 8 September, but it becomes considerably more interesting in light of the Anthropic failures we discussed.

Meta's new personal agent, Muse, runs inside a dedicated virtual computer.

More importantly, a separate Sentinel agent is kept apart from Muse at the system level.

Meta says nothing Muse does can reach the internet unless Sentinel approves it and sensitive actions can require explicit human permission. Credentials are also managed so that Muse can use them without necessarily being allowed to inspect them directly.

The company hasn't supplied enough independent evidence yet to establish how robust Sentinel actually is. But the architecture matters.

Yesterday Anthropic's cyber investigation suggested that a model pursuing an objective could interpret contradictory evidence in ways that allowed it to continue pursuing that objective.

Meta's design essentially says: don't expect the actor to police itself.

Give the actor a goal. Then put something structurally separate between the actor and the world.

That resembles a classic computer-security concept called a reference monitor: an independent mechanism through which privileged actions must pass.

This may become a fundamental architecture of powerful agents:

agent proposes → governor evaluates → environment executes.

Rather than:

agent thinks → agent decides whether its own thought is safe → agent acts.

That is a considerably more defensible security architecture.

6. Anthropic's latest threat report exposes another form of AI acceleration

Reported 10–11 September.

Anthropic says it has disrupted large-scale efforts by several Chinese AI laboratories to extract capabilities from Claude.

The figures are extraordinary enough that they need caveats: these are Anthropic's allegations, not independently audited numbers, and the named companies have not publicly substantiated Anthropic's account.

But Anthropic says an Alibaba-linked operation generated more than 151 million Claude exchanges between May and July, at one point approaching three million interactions per day through more than 3,500 accounts.

Anthropic also alleges that Moonshot and DeepSeek routed some live customer queries through Claude and retained the responses as training data.

Anthropic had already disclosed smaller-scale illicit distillation attacks in February, when it said three labs had generated about 16 million exchanges.

So what is new isn't the technique. It's potentially the scale.

Why this matters more than an IP dispute

A frontier model doesn't need to release its weights for another model to learn from it.

Its outputs themselves become training data.

A weaker AI can ask a stronger AI millions of carefully designed questions, then train on the answers.

The stronger model has effectively become a teacher for its competitor.

This creates an uncomfortable possibility:

frontier capability may diffuse much faster than frontier training infrastructure.

A laboratory may spend tens of billions creating an ability and competitors may subsequently acquire some fraction of that ability at dramatically lower marginal cost through synthetic data and distillation.

That means the gap between the frontier and the rest could repeatedly compress.

And that has geopolitical implications because semiconductor restrictions make it difficult to transfer compute — but intelligence embedded in model outputs travels through an API connection.

7. Physical AI gets an interesting extension of yesterday's two-speed learning idea

Confirmed — reported 10 September.

Skild AI says its S1 robotic foundation model can observe one video demonstration of an unfamiliar physical task and then attempt it without updating its weights or undergoing task-specific training.

Examples include plant potting, coffee preparation and assembly tasks lasting as long as ten minutes.

The reported performance figures are company tests and need independent replication, so I would not focus on the claimed success rates.

What interests me is the architecture.

Yesterday we distinguished slow learning — modify the model's weights — from fast learning — preserve new information in state.

S1 appears to push that second mode into robotics:

watch demonstration → construct temporary task representation → execute using existing competence.

No new training cycle.

If this proves robust, humans won't need to programme many robotic behaviours in the conventional sense. We may show robots work in roughly the same way we show another person.

Under the hood today: the KV cache

When a language model reads a sentence, it doesn't completely recompute everything it knows about every previous word each time it generates another token.

During attention, the model calculates internal representations called keys and values for earlier tokens. It stores them. That store is the KV cache.

Very crudely, think of it as the model's temporary computational memory of the conversation.

For a short chatbot conversation, that isn't terribly dramatic.

For an agent working for six hours, however, the working context might contain system instructions, 100,000 lines of code, messages, documents, tool definitions, search results, intermediate actions and previous agent outputs.

And several parallel agents may each maintain their own copy.

Now KV cache becomes enormous.

More importantly, that data has to repeatedly move between extremely fast GPU memory and slower memory/storage.

Modern AI chips can perform staggering numbers of calculations. But sometimes the processor is effectively standing around waiting for data to arrive.

That makes memory bandwidth one of the hidden constraints on AI intelligence.

It explains why DeepSeek claiming one quarter of the HBM and one eighth of the SSD KV-cache footprint is potentially more consequential than a three-point benchmark improvement.

The deeper insight

As we scale agents, intelligence isn't constrained solely by how clever the neural network is.

It is constrained by how efficiently the system can maintain its working mind.

That makes techniques involving compression, sparse attention, memory hierarchies, retrieval and context management part of the capability race itself.

Evidence → inference → speculation

Evidence:
OpenAI is explicitly productising its agent harness.
DeepSeek has changed its model architecture partly to reduce the memory cost of agent workloads.
Cognition is training models against both task success and inference cost.
OpenAI is separating low-latency voice interaction from deeper asynchronous reasoning.
Meta has placed a separate governor between its personal agent and external actions.
Robotic models are demonstrating increasingly powerful adaptation without model retraining.
Anthropic says competitors are using frontier-model outputs at industrial scale as training material.

Inference:
The optimisation target for frontier AI is shifting from maximum model intelligence towards maximum useful intelligence per unit of time, memory, compute and risk.

Speculation:
If that trend continues, the eventual thing we call an “AI model” may be rather like what we call a “computer”. You interact with one object. But internally it consists of specialised processors, memory systems, operating systems, schedulers, security boundaries and networked services dynamically allocating resources.

There may never be one monolithic artificial mind. There may instead be an illusion of one mind generated by an increasingly complicated cognitive infrastructure.

What I think is mostly noise today

There is substantial coverage this morning of AI researchers warning about existential risk. The concern matters, but the headlines themselves don't materially change our model because we've already incorporated the underlying evidence about automated R&D and agent failures.

I would also mostly ignore today's “model X beats model Y” narrative.

What matters intellectually is how they're improving: asymmetric architectures, sparse activation, cheaper KV caches, cost-aware RL, verification flywheels and agent-specific optimisation.

Those mechanisms tell us something about tomorrow. Leaderboard positions mostly tell us something about Thursday.

Three questions worth carrying forward

1. Does the agent harness become the real platform moat?

If frontier models become relatively substitutable while memory, tools, permissions, environments and orchestration become increasingly sophisticated, then control of the cognitive operating system may ultimately matter more than having the absolute best foundation model.

2. Can AI learn to spend intelligence rationally?

At some point an AI may itself decide: this decision deserves another 100 agents and six hours of computation; that one needs 200 milliseconds.

If that becomes reliable, inference scaling becomes much more economically powerful.

3. How fast does intelligence diffuse through distillation?

If a frontier model can cheaply teach a second model, which then teaches a third, improvements may propagate across the ecosystem faster than training cycles suggest.

Today's mental-model update

Two days ago:
AI is increasingly helping build better AI.

Yesterday:
The relevant unit is becoming the cognitive system rather than the individual model.

Today I would extend it again:

the cognitive system is becoming internally specialised and economically self-conscious.

We're seeing the beginnings of something that looks like:

perception / conversation → task router → appropriate reasoning effort → specialist model or agent → parallel subagents when justified → memory and persistent context → tools and physical action → independent safety governor → verification → experience returned to future systems.

And underneath all of that sits a resource scheduler asking:

How much intelligence is this problem worth?

That may turn out to be an extremely important question.

Because once intelligence becomes something a machine can allocate rather than merely possess, the trajectory changes again.

The critical metric may eventually cease to be “how intelligent is the best model?”

It may become:

“how efficiently can an artificial cognitive system turn compute, memory, tools and time into reliable progress on an objective?”

That feels like the most useful conceptual addition from the last 24 hours.