Intelligence is getting cheaper. Reliability is not.
Frontier-adjacent capability is falling sharply in price while open training stacks and agent-ready robotics show the other half of the story: useful autonomy increasingly depends on the harness, runtime controls and execution environment around the model.
Signals beneath the AI headlines
LEAD SUMMARY - ANALYSIS. Since the last AyEye Today edition, the most important shift is not a single benchmark win. It is the changing economics of intelligence. OpenAI released GPT-6 Sol and Luna at prices that put strong reasoning capability into much cheaper tiers, while Anthropic says Claude Opus 5.5 reaches roughly Fable 5.1-level performance on most work at 40% lower running cost than Opus 5. At almost the same moment, Xiaomi opened more of the reinforcement-learning and agent-harness stack around MiMo-V2.6, while NVIDIA made its robotics tooling explicitly agent-ready.
The useful synthesis is this: capable intelligence is becoming abundant faster than dependable autonomy is. That moves the scarce value outward from the base model towards data, memory, tools, harnesses, verification, identity, permissions and operating discipline.
1. The near-frontier is moving down the cost curve
CONFIRMED. OpenAI's 22 September API changelog lists GPT-6 Sol at $2 per million input tokens and $10 per million output tokens, with GPT-6 Luna at $0.10 input and $0.50 output for standard prompts within the stated context threshold. OpenAI API changelog
INTERPRETATION. Lower unit cost changes architecture, not just procurement. It becomes rational to use intelligence more often - for checking, routing, comparing, monitoring and supervising as well as for producing the final answer. Total token consumption may therefore rise even while unit prices fall.
2. Anthropic is compressing price without retreating from capability
CONFIRMED. Anthropic introduced Claude Opus 5.5 on 22 September and says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. Anthropic also highlights external evaluation and its automated behavioural audit. Those are vendor claims, but the economic direction matches the OpenAI signal. Anthropic: Claude Opus 5.5
INFERENCE. The economic frontier may matter as much as the capability frontier. A model that is fractionally weaker but dramatically cheaper can be more transformative inside organisations because it can be placed in more loops and used for supervision as well as production.
3. Xiaomi opens more of the improvement loop
CONFIRMED - UNDER THE HOOD. Xiaomi's MiMo-V2.6 material describes open task environments and lightweight, composable mini-harnesses that decouple system prompts, tools and context management. It also describes multi-harness training intended to improve generalisation across different agent frameworks, including unseen ones. Xiaomi MiMo-V2.6
ANALYSIS. This matters because the harness is moving from deployment plumbing into the learning system itself. The model is no longer simply trained and then placed inside an agent. Agent environments, tool interfaces and orchestration choices are becoming part of how capability is produced and generalised.
4. Robotics makes the harness visible because mistakes move atoms
CONFIRMED. NVIDIA's Isaac ROS 5.0 release introduces agentic workflows and AI-agent skills for robotics development, including setup and migration tasks. The release also expands open physical-AI libraries and deployment support across Jetson. NVIDIA: Isaac ROS 5.0
ANALYSIS. Robotics is a useful preview of enterprise agent engineering. Physical systems force explicit treatment of state, timing, interfaces and recovery. The closer AI gets to consequential action, the less sufficient "the model is smart" becomes as an assurance argument.
Concept to learn today: Capability compression
CAPABILITY COMPRESSION is the process by which a level of AI performance that was recently rare, expensive or frontier-only becomes available in cheaper models, smaller systems or more accessible deployment stacks.
It has two consequences. First, intelligence gets embedded in more places. Second, differentiation moves outward from the model towards data, memory, tools, workflow design, verification, identity, permissions and operating discipline. The cheapest token is not automatically the cheapest outcome if a weakly governed agent repeats work, calls the wrong tools or needs expensive human checking.
Noise: benchmark decimals without an operating context
NOISE. Small differences between headline scores matter less when cost, latency, tool behaviour, recovery and task-specific reliability can dominate the real result. A leaderboard can tell us something about capacity. It tells us much less about whether a deployed system is competent to perform a role repeatedly under real constraints.
Mental-model update
Previous: model capability -> working system -> demonstrated competence -> accumulated experience -> verified trajectory.
Now add: capability compression -> abundant intelligence -> more frequent delegation -> greater dependence on harness reliability and authority controls.
The practical implication is counter-intuitive: cheaper models can make governance more important, not less. When intelligence is expensive, organisations ration its use naturally. When it becomes cheap, the volume of autonomous decisions can rise much faster than the quality of the controls around them.
Questions to carry forward
- Which workloads become economical only because AI can now be used as a checker or supervisor as well as a producer?
- Should agent competence evidence attach to the whole deployed configuration - model, harness, tools, permissions and version - rather than the model alone?
- What assurance practices from robotics will migrate back into ordinary enterprise software agents?
