Dominic Chiappe · People, capability & transformation

Thinking about how organisations perform in an AI-enabled world

AyEye Today · 2026-09-09

AI Intelligence Briefing — Inaugural edition

The emerging unit of capability is an organised cognitive system, with feedback between research, tools and verification.

Archive edition · Original reporting and analysis, preserved as published. Website layout adapted for reading.

AI Intelligence Briefing — 9 September 2026 ## Inaugural edition

Executive view

The most important change in the AI picture this week is not GPT‑6 Astra itself. It is evidence that the frontier labs are moving from one model answering one prompt towards large populations of agents conducting research, using tools, sharing intermediate discoveries, checking one another and feeding results back into subsequent work.

Yesterday OpenAI disclosed that its Navier–Stokes result was produced using an internal model it says is significantly more capable than GPT‑6 Astra, despite Astra having been released only days earlier. The internal model is still being trained. Around 10,000 concurrent agents were coordinated on the problem, with useful insights consolidated and redistributed between groups. This is potentially a more important capability transition than another jump in raw benchmark scores.

The emerging hypothesis I would keep in mind is this:

AI progress may increasingly come from the system surrounding the model — inference compute, agent populations, memory, tool use, verification, orchestration and AI-assisted research — rather than simply from making one neural network larger.

That does not mean model scaling is over. Quite the opposite: the model is improving at the same time as the machinery around it becomes much more powerful.


1. The strongest signal: OpenAI has a model beyond Astra — and used ~10,000 agents on Navier–Stokes

FACT — 8 September. OpenAI published a proposed solution to the Navier–Stokes Millennium Prize problem, including a conventional mathematical write-up and a formal proof in Lean. OpenAI says the system demonstrates finite-time singularity for a form of the three-dimensional Navier–Stokes equations.

The scientifically important claim will now have to survive independent mathematical scrutiny. The AI-development details, however, are independently interesting regardless of whether every part of the proof survives that process.

OpenAI says:

  • the work used an internal model significantly more capable than GPT‑6 Astra;
  • training on that model began on 28 August and is still continuing;
  • approximately 10,000 agents were involved in the Navier–Stokes effort;
  • the agents could read cached internet material, run code and communicate in groups;
  • useful intermediate results were consolidated and fed back to other groups;
  • the decisive result appeared about 88 hours after the first agents were launched;
  • GPT‑6 Astra then helped formalise and verify the result in Lean;
  • across the wider mathematical experiment, agents generated 4.9 million messages and roughly 300 billion output tokens.

Source: https://openai.com/index/navier-stokes-solution/

Why this matters

We are used to thinking of capability as a property of the model: GPT‑4 → GPT‑5 → GPT‑6.

This experiment suggests that the relevant unit may increasingly be the cognitive system:

model + thousands of parallel attempts + tools + communication + selection + consolidation + verification + huge inference budgets.

That is closer to an artificial research organisation than a chatbot.

There is also an important economics point hiding underneath this. Training compute gets most of the attention, but this approach throws enormous amounts of inference compute at a problem after training. Intelligence may therefore become increasingly elastic: spend ten seconds and get one level of answer; spend three days and billions of tokens coordinating thousands of agents and get something qualitatively different.

My interpretation

The frontier may be moving from “How intelligent is this model?” towards “How much reliable cognitive work can this system marshal against one objective?”

That is a substantial conceptual shift.


2. Recursive self-improvement: yesterday’s discussion now has a harder empirical foundation

You already know the basic RSI argument, so I will not re-explain it.

WHAT IS NEW: OpenAI has now published internal operating data showing that AI agents are materially entering the AI-development process rather than merely being used as coding assistants.

OpenAI says it has reached what it calls an automated research intern: a system capable of carrying out well-defined research tasks under human direction that might otherwise take a skilled researcher several days. Its stated target is an automated AI researcher by March 2028.

By mid-August, OpenAI reports that its research organisation was consuming roughly 3.1 agent-workdays for every human workday. Researchers are running more experiments and increasingly using several agents concurrently.

Humans still decide research priorities, judge results and decide whether systems should be scaled or deployed. So this is not autonomous RSI.

But the loop is now visible:

AI helps conduct AI research → research improves models and research tooling → improved AI can conduct more ambitious AI research → repeat.

Source: https://openai.com/index/research-acceleration-view-inside-openai/

Evidence / inference / speculation

Evidence: AI agents are already increasing research throughput inside OpenAI.

Inference: as agent capability rises, a progressively larger fraction of model-development work can be delegated.

Speculation: at some threshold, improvement in research capability could shorten the time required to produce the next improvement in research capability. That is where ordinary automation begins to resemble genuine recursive acceleration.

We do not yet have convincing public evidence that this acceleration curve has become self-sustaining.


3. A revealing change in OpenAI’s language

OpenAI Chief Scientist Jakub Pachocki published an essay on 6 September called *An Alien Mind*.

The notable part is not philosophical language about AGI. It is that he explicitly says that, based on internal results, he has a strong expectation that current progress could be sustained into recursive self-improvement, and that AI systems will increasingly drive their own development.

He also says OpenAI is orienting research towards RSI because it believes this may be necessary to remain at the research frontier, while simultaneously arguing that scaling may need to be deliberately slowed when alignment and monitoring cannot keep pace.

Source: https://openai.com/index/an-alien-mind/

Why I think this is worth noticing

Labs have talked abstractly about self-improving AI for years. This is different in tone because it is being said alongside:

  • a working automated-research-intern claim;
  • measured AI participation in research;
  • an unreleased model beyond Astra;
  • very large multi-agent scientific experiments.

None proves runaway RSI.

Together they do suggest that automated AI R&D has moved from a future scenario into an active engineering programme.

That is the conceptual update I would retain.


4. AI science is separating into two forms

Google DeepMind yesterday launched the AlphaGenome Atlas, containing model predictions for the effects of roughly nine billion possible single-letter variants in the human genome.

Source: https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/

This is different from the Navier–Stokes story in an instructive way.

There are increasingly two kinds of AI-for-science:

A. Specialist scientific models — AlphaFold, AlphaGenome, WeatherNext, AlphaEvolve and similar systems that learn highly structured scientific domains.

B. General reasoning/agent systems — frontier language models plus code, search, tools, formal verification and large inference budgets.

The interesting future may be their convergence.

Imagine a general research agent able to choose and operate dozens of specialised scientific models, formulate experiments, operate laboratory equipment, analyse results and update the research plan.

That starts to resemble something much closer to an automated laboratory than an AI assistant.


5. The physical-world bridge is quietly being standardised

This is slightly older — 27 August — but I am including it in the inaugural briefing because it is an important baseline signal rather than today’s news.

Anthropic has introduced a research preview of its Model Hardware Standard (MHS): a common interface intended to let AI agents operate laboratory and manufacturing equipment including microscopes, liquid handlers and robotic arms.

Anthropic describes agents orchestrating multiple instruments, adjusting experimental parameters and in some circumstances recovering from hardware errors without intervention.

Source: https://www.anthropic.com/news/model-hardware-standard-research-preview

Why it belongs in our model

Software agents become much more economically consequential when the action layer stops at the edge of a browser.

Standards such as MHS potentially give the model an actuator layer into science and manufacturing.

Put that beside automated AI research and specialist scientific models and a plausible architecture appears:

hypothesis → simulation/model → experiment design → physical execution → observation → analysis → revised hypothesis

with humans supervising the overall programme rather than manually performing every step.

That loop is worth watching very carefully.


6. Under the hood: voice is becoming interaction, rather than speech generation

Meta published Alignment-Free Text-Audiobox on 6 September.

The details are more interesting than another voice demo. It uses a latent diffusion architecture in which 48 kHz audio is represented at only 25 latent steps per second, providing more than ten-fold compression relative to Meta’s previous representation. It learns text/speech alignment through cross-attention rather than requiring explicit duration alignment.

The model is trained on 480,000 hours of speech and supports full-duplex dialogue: overlapping speech, back-channel signals, turn-taking and emotional dynamics.

Source: https://ai.meta.com/research/publications/alignment-free-text-audiobox-for-voice-dubbing-and-full-duplex-dialogue-synthesis/

Why this matters

Human conversation is not:

person finishes sentence → machine computes → machine speaks → person waits.

We interrupt, hesitate, change direction, say “mm-hm”, infer when somebody has finished and react to tone while the other person is still talking.

Full-duplex models begin moving AI interaction towards that continuous loop.

That sounds like UX polish, but it may have deeper consequences. Once models continuously perceive and act rather than process discrete turns, the boundary between assistant and ongoing cognitive companion/agent becomes less distinct.


Concept to learn today: inference scaling

We have spent several years talking about training scaling: more compute, more data, larger or better-trained models produce greater capability.

Increasingly we also need to think about inference scaling.

Inference is what happens after the model has been trained. Instead of asking the model for one answer, a system can spend additional compute to:

  • reason for longer;
  • generate many candidate solutions;
  • critique those solutions;
  • run experiments or code;
  • search external information;
  • create sub-agents;
  • compare approaches;
  • verify promising results formally;
  • recursively investigate failures.

The Navier–Stokes experiment is an extreme illustration.

One copy of even a very intelligent model might fail. Ten thousand copies exploring a solution space, exchanging discoveries and backed by verification machinery may solve problems that no individual instance reliably can.

A rough analogy is the difference between asking one scientist for an answer during lunch and funding an entire institute for three years.

The underlying human brain has not become smarter. The amount and organisation of cognition brought to bear on the problem has changed.

For AI this difference can be compressed dramatically in time because agents can run simultaneously.

Why this concept matters for forecasting

If inference scaling works well, future capability will not be described by a single IQ-like number for a model.

A system may have several operating points:

cheap/instant → competent

minutes of computation → expert

hours + tools + agents → research team

days + massive parallel inference → frontier discovery system

That means benchmarks measured with a fixed inference budget may increasingly understate what a model/system can accomplish when given much larger resources.


Weak signal worth watching

The most interesting weak signal this week is orchestration.

For several years, almost all attention was directed towards the foundation model itself. Increasingly the frontier systems contain another layer:

task decomposition → sub-agents → context management → communication → result selection → verification → memory.

If this continues, some of the most important AI advances may occur in the architecture around the neural network rather than inside it.

That would also explain why apparent capability can jump dramatically without a proportionate jump in base-model benchmark scores.

I would not yet call this a settled conclusion. But it is now a hypothesis worth tracking.


What is probably noise today

I would largely ignore routine announcements that a company has added an agent, launched another image generator or achieved another small benchmark lead unless the underlying mechanism is novel.

Similarly, enormous AI funding rounds tell us a lot about capital concentration but relatively little about the shape of intelligence itself.

The higher-value questions right now are:

What can systems do for longer? What can they verify? How many agents can cooperate effectively? What tools can they operate? How much of AI research itself can they perform?

Those are much better indicators of the capability trajectory than another chatbot leaderboard.


Three open questions to carry forward

1. Does multi-agent scaling continue to work at 10,000+ agents?

Parallelism creates enormous search capacity, but it also creates coordination noise and duplicated effort. We need to understand whether capability scales smoothly with agent population or quickly hits organisational limits.

2. Is AI research becoming the dominant accelerator of AI progress?

If AI-assisted research moves from a productivity multiplier to the primary source of algorithmic improvement, our existing expectations about model-generation timelines may become unreliable.

3. What becomes the real bottleneck?

If cognition becomes abundant, the limiting factor might shift from intelligence to verification, experimental throughput, compute, energy, physical equipment, trusted data or human permission.

The bottleneck moving is often more important than the capability increasing.


Today’s mental-model update

Yesterday our model was roughly:

AI is beginning to help build better AI.

After this week’s evidence, I would sharpen it to:

Frontier AI development is becoming a hybrid research system in which humans set objectives and exercise judgement, while increasingly large amounts of experimentation and cognitive labour are performed by AI. At the same time, inference-time orchestration is turning a single model into something resembling a scalable research organisation.

The question is no longer whether that loop exists.

The important questions are how quickly more of the loop becomes automated, whether the loop itself accelerates and whether verification/alignment can keep pace.

That is the thread I think we should keep updating rather than repeatedly rediscovering the fact that agents are getting better.

— End of inaugural briefing