Dominic Chiappe · People, capability & transformation

Thinking about how organisations perform in an AI-enabled world

AyEye Today ·

The clever bit is moving out of the conversation

AI systems are beginning to shift expensive reasoning into training, persistent memory and specialised components, making reusable intelligence as important as live model deliberation.

Signals beneath the AI headlines

LEAD SUMMARY. Three developments matter this morning. Google Research has shown a way to spend expensive reasoning during training and then compile the behaviour into a tiny, fast retriever. New robotics work is showing something similar at the system level: useful intelligence increasingly comes from persistent memory and reusable experience, not from starting every task from scratch. Meanwhile, a reported decision inside Google to give engineers broad access to Anthropic's Claude is a quiet reminder that even frontier labs are becoming multi-model organisations.


The clever bit is moving out of the conversation

CONFIRMED · UNDER THE HOOD. Google Research published Retrieve-for-Train on 15 September. The system uses reinforcement learning offline to discover good query fan-outs, turns those results into training data, then trains a 53.9-million-parameter diffusion retriever to generate the whole set in parallel. Google reports roughly 12–20× faster fan-out than autoregressive approaches in its experiments.

The important part is not the search use case. It is the architectural pattern: pay for difficult reasoning once, then compile the result into something cheaper that can be reused at scale. We have spent much of the past two years watching models think longer at inference time. Retrieve-for-Train points in the opposite direction: move some of that intelligence into training, specialised components and precomputed structure.

This is not a universal substitute for reasoning. Novel, open-ended problems will still require deliberation. But many enterprise tasks are repetitive enough that the expensive exploratory phase can potentially be amortised across thousands or millions of later actions.

Source: Google Research, 15 September 2026


Robots are beginning to remember what happened last time

CONFIRMED · MEMORY. MessyMem, submitted on 14 September, gives a mobile manipulator persistent spatial memory that is updated by interaction. If the robot discovers that a cabinet is locked or finds an object in a drawer, that information can be reused in later tasks rather than rediscovered from scratch. In a continuous 25-task simulation, the authors report 80% task progress, 28.9 percentage points above their strongest external baseline, with a smaller physical-robot evaluation as well.

This connects directly to the shift we have been tracking in software agents. Memory is not merely a longer transcript. Useful memory is selective, structured and action-linked: what was encountered, what changed, what failed, what worked and where that evidence belongs in the environment.

The broader implication is that persistent experience may become a scaling axis in its own right. Bigger models know more in general; experienced systems know more about this environment.

Source: MessyMem, arXiv, submitted 14 September 2026


Touch is becoming an evaluation surface, not just a sensor

CONFIRMED · ROBOTICS. Bench2Dex, also submitted on 14 September, provides a shared simulation benchmark for visual-and-tactile two-handed manipulation across 12 different dexterous hands and 26 tasks. It includes around 1,300 human-teleoperated demonstrations and tests robustness under several perturbation types.

The quiet significance is methodological. Embodied AI has been difficult to compare because every hand, sensor, task and success criterion is different. Standardising the evaluation surface does not solve dexterity, and simulated touch is not real touch, but it makes progress more falsifiable.

That matters because physical AI is moving towards the same discipline that software agents now need: not simply impressive demos, but repeatable evidence under changed conditions.

Source: Bench2Dex, arXiv, submitted 14 September 2026


Google reportedly gives its engineers access to Claude

REPORTED · ORGANISATIONAL SIGNAL. Business Insider reports that Google has widened internal access to Anthropic's Claude through its engineering environment, after previously limiting external coding tools more tightly. Google remains a major Anthropic investor, and Gemini remains central to Google's own stack, so this should not be read as a model verdict.

The interesting signal is more general: frontier organisations themselves may be moving towards model portfolios rather than model monocultures. If different models have different strengths, the rational enterprise architecture may be to route work across them rather than force one model to be best at everything.

That changes the competitive question from “which model wins?” to “who can orchestrate a changing portfolio of models, tools, memory and policy with the least friction?”

Source: Business Insider, 15 September 2026


AI safety agreement is becoming an implementation problem

CONFIRMED REPORTING · ANALYSIS. AP reports today that leading AI companies have shown unusual convergence around the need for stronger safety mechanisms, while the practical questions remain unresolved: who evaluates whom, how independent oversight should work, what can be standardised across companies and how international competition affects compliance.

This is a meaningful update to yesterday's signal about a possible shared safety body. The debate is moving from “do we need assurance?” towards the much harder institutional question: who has authority to declare a frontier system safe enough for a particular kind of deployment?

The likely answer will not be one universal safety certificate. It is more likely to resemble layered assurance: model evaluations, deployment-specific controls, external tests, incident reporting and continuously updated evidence.

Source: Associated Press, 16 September 2026


Concept to learn today: amortised intelligence

In computing, amortisation means paying a cost once and spreading its benefit over many later operations. AI is increasingly doing the same thing with reasoning.

Explore expensively once → capture what worked → compile or store the result → reuse cheaply many times.

Retrieve-for-Train does this through training: expensive RL discovers useful behaviour, then a small diffusion model reproduces it quickly. Persistent agent memory does it through experience: an agent learns something about an environment once, then retrieves it later instead of rediscovering it. Tool libraries do it through procedural knowledge: a successful sequence becomes a reusable capability.

The architectural implication is easy to miss. Future AI systems may not simply be “a model plus more thinking tokens”. They may contain several forms of stored intelligence, each optimised for a different timescale: model weights for broad prior knowledge, memory for local experience, specialised components for repeatable reasoning patterns and live deliberation for genuinely new problems.


Weak signal: the product may include humans for longer than the demos imply

WEAK SIGNAL. OpenAI is reportedly partnering with London legal-technology firm Telon, whose “legal engineers” work inside law firms to configure models, workflows and adoption. The specific partnership is commercially narrow, but the pattern matters.

As models become more capable, implementation expertise does not necessarily disappear. It may move closer to the work: people translating professional practice into tools, policies, evals and workflows. In other words, the near-term AI product may often be a socio-technical system rather than a self-installing model.

Source: Business Insider, 16 September 2026


Noise: “more reasoning” is not a universal architecture

Longer chain-of-thought, larger context and more test-time compute remain important, but they are becoming one option among several. If a task repeats, storing or compiling successful behaviour may beat reasoning from first principles every time. If a task acts on the world, verification and memory may matter more than eloquence. If several models are available, routing may matter more than squeezing another few benchmark points from one of them.


Mental-model update

The model we have been building over the past week now needs one more layer. We moved from model → working system, then from capability → competence. Today's evidence adds experience and compilation.

The emerging AI system is becoming a small cognitive economy: some intelligence is expensive and general, some is cheap and specialised, some is accumulated through memory, some is verified before action, and some is selected dynamically from a portfolio of models.

Short term: more routing, caching, memory and specialised inference components around frontier models. Medium term: organisations deliberately compile recurring expert work into cheaper reusable capabilities while reserving frontier reasoning for novelty and exceptions. Longer term: the unit of AI advantage may be less “who owns the smartest model?” and more “who turns expensive intelligence into reliable reusable capability fastest?”

Questions to carry forward

  1. Which categories of reasoning can be safely compiled or memorised, and which must remain live because the environment changes too quickly?
  2. Does persistent experience become as important a scaling axis as parameter count and inference compute?
  3. If enterprises become multi-model by default, what evidence should govern routing between models, humans, agents and specialised components?