Dominic Chiappe · People, capability & transformation

Thinking about how organisations perform in an AI-enabled world

AyEye Today ·

Capability is not competence

Raw model capability is becoming only one component of useful, governable intelligence. The harder question is whether a particular deployed system can perform a particular role, under defined conditions, with current evidence that it remains competent to do so.

Signals beneath the AI headlines

Frontier labs move from safety statements towards shared assurance infrastructure

CONFIRMED · ANALYSIS. On 14 September, the Washington Post reported that leaders at Anthropic, OpenAI and Google had discussed a new AI safety body. The interesting shift is institutional: frontier assurance is beginning to look like something that may need shared mechanisms and independent access rather than each laboratory grading its own homework.

This extends yesterday’s signal about embedded evaluators. The new element is the possibility of cross-lab infrastructure. If it develops, evaluation could become a continuing function around frontier systems rather than a pre-release ceremony.

Source: Washington Post, 14 September 2026


Long rulebooks remain brutally hard

CONFIRMED · EVALUATION. Tasks over Application Manuals (TAM), submitted 11 September, tests GPT-5 on ICD-10-CM clinical coding and US federal sentencing rules. The best prompting baseline achieved only 1% exact match on clinical coding and 15.5% on sentencing.

This is a useful antidote to broad claims that professional reasoning is solved. Retrieval is not the same as procedural competence. In long rule systems, an early mistake propagates, exceptions interact and exactness matters.

Source: TAM paper, arXiv


DeepSeek replaces Pro traffic with Flash

UNDER THE HOOD. DeepSeek’s V4.1-Flash was announced on 10 September, but the operational change arrived on 14 September: V4-Pro API requests began routing to V4.1-Flash. DeepSeek describes V4.1-Flash as a 552B-parameter MoE using only 8B active parameters for input and 16B for output, with a KV cache requiring one quarter of the HBM and one eighth of the SSD storage of the previous generation.

The important update is architectural: spend computation asymmetrically where it creates value rather than activating everything everywhere.

Source: DeepSeek


Agents are beginning to optimise the substrate they run on

AI-ASSISTED AI RESEARCH. ForgeMegakernel uses coding agents to generate high-performance autoregressive decode megakernels, with an independent mid-state test oracle checking correctness while the agent searches for implementations. Across the reported tests, generated kernels decoded faster than the comparison SGLang engine at comparable answer accuracy.

This is a credible form of bounded AI self-improvement: AI improving the software infrastructure that makes subsequent AI inference faster.

Source: ForgeMegakernel paper, arXiv


Capability is not competence: the missing assurance layer for agents

ORIGINAL SYNTHESIS · AGENT ASSURANCE. A model benchmark establishes evidence about model capability. It does not establish that a deployed agent is competent for a role. Agent competence is conditional on the model + harness + tools + permissions + memory + environment + supervision + recent performance.

Capability — what the underlying model can potentially do.
Competence — evidence that the deployed system can perform defined tasks under defined operating conditions.
Authorisation — which of those tasks it is currently permitted to perform.
Currency — whether the evidence remains recent enough to trust after model, tools, data or environment have changed.

This resembles competency assurance in regulated human work because both require role-specific evidence rather than a generic label of ability. The analogy breaks in important places: agents can be copied exactly, updated simultaneously, instrumented continuously and can change behaviour when a shared model or harness changes. Human competence degrades and develops differently.

The stronger idea is a common evidence architecture for a mixed workforce. Human evidence might include observed performance, assessment, qualifications and recent practice. Agent evidence might include evals, execution traces, incidents, supervision rates, environmental versions and recent successful task histories. Both could feed continuing decisions about what work each actor is trusted and authorised to perform.


Concept to learn today: Currency

Competence is not permanent. In safety-critical human systems, currency asks whether someone’s evidence is recent enough for the task they are about to perform. The same idea may become crucial for agents.

An agent that passed an evaluation three months ago may now use a different model, prompt, memory store, tool version, permission set or data source. Its old score has not necessarily become false; it has become evidence about a different configuration.

Configuration identity → recent task evidence → incident history → supervision outcomes → current confidence → permitted authority.


Harness self-improvement is becoming a separate research surface

WEAK SIGNAL. MetaRSI-v1 argues that recursive improvement can occur at three surfaces: data, model weights and the harness around the model. Its open-source RSI Harness makes prompts, tools, skills, memory and runtime policies versionable and editable. The claims are early and need independent reproduction, but the framing is important: practical self-improvement may often mean improving the cognitive machinery around a model rather than the model rewriting itself.

Source: MetaRSI-v1 project


Noise: treat benchmark superlatives as measurements, not identities

The week is full of models described as best, frontier, expert or autonomous. Those labels increasingly collapse model, harness, task distribution and scoring method into one adjective. The more agentic systems become, the less useful that shorthand becomes.


Mental-model update

The emerging unit of AI progress is becoming a governed working system rather than a model. Today sharpens that further: the useful unit of trust may be a role-specific, version-specific competence claim backed by continuously refreshed evidence.

Short term: better agent evals and configuration-aware testing. Medium term: authorisation layers connecting eval evidence to tool permissions and autonomy. Longer term: mixed human-agent organisations may maintain a live capability and competency graph in which authority follows evidence rather than whether the worker happens to be biological or digital.

Questions to carry forward

  1. What minimum evidence should be required before an agent moves from assistance to delegated execution to autonomous authority?
  2. Can agent competence be continuously inferred from normal work traces without turning every deployment into an expensive evaluation programme?
  3. If a shared foundation model changes underneath thousands of agents, how quickly should their existing authorisations expire?