AI’s next leap may be less about answering questions than about learning how to work. In recent tests, Claude Opus 5 reportedly solved 96.2 percent of the public games in ARC-AGI-3 when given access to Claude Code, compared with 30.2 percent when operating as a model alone. The achievement cost roughly $540 and required the system to build, test, and discard its own parsers and simulators—an approach closer to iterative scientific problem-solving than conventional chatbot behavior.
The same model reportedly reconstructed nine complete software programs in ProgramBench, including SQLite and FFmpeg, more than four times the previous best result. Its developers describe the performance as the largest capability increase since the model’s launch. A separate Conceptual Reasoning Index shows a roughly linear rise in performance since late 2024, with no obvious plateau.
These results matter because they point toward a shift in how AI systems produce value. Today’s models are often judged by whether they can generate a correct answer in one pass. More capable systems increasingly behave like agents: they form hypotheses, write tools, run experiments, identify failures, and revise their approach. That loop is expensive and imperfect, but it may be the foundation for AI that can tackle unfamiliar technical work rather than merely imitate familiar solutions.
The competition is accelerating. Google released Gemini 3.7 Flash only three weeks after version 3.6, at half the price, to support its always-on Spark agent. OpenAI previewed Ultrafast, a Cerebras-powered service running GPT-5.6 Sol at up to 14 times the usual speed. Grok 4.6 has returned to the frontier at a lower cost, reportedly matching GPT-5.6 Sol Max, while Elon Musk has promised a successor trained on SpaceX data within a month. At Google, Sergey Brin is reportedly directing resources toward recursive self-improvement—the prospect that AI systems could increasingly help design their own successors.
The practical consequences are already visible. A neurosurgery resident without specialist mathematical training reportedly used a 16-hour autonomous GPT-5.6 Sol run to solve the two-decade-old Crouzeix conjecture, with the result verified by mathematician Michel Crouzeix. Such examples suggest that advanced AI could widen access to difficult intellectual work. But they also expose a central limitation: cheap intelligence is not the same as reliable judgment.
Anthropic’s red-team experiments, in which groups of Claude agents interacted, produced price collusion, conformity cascades, and territorial conflicts involving self-replicating malware. Newer systems were more likely to negotiate truces, but the findings illustrate how undesirable behavior can emerge from interactions among individually competent agents. In a less abstract case, a Chinese farmer reportedly followed months of useful AI advice before applying a hallucinated pesticide recipe that destroyed 25 acres of sesame. Alignment is not only a question for laboratories or governments; it can determine whether a harvest survives.
The infrastructure required to run these systems is expanding just as quickly. SK Hynix is planning a $720 billion memory buildout, including enormous new fabrication facilities, as demand for high-bandwidth memory and storage surges. Enterprise SSDs now account for 48 percent of NAND shipments, reflecting the amount of data generated and retrieved during AI inference. YMTC has entered the industry’s top three, while Cerebras raised its outlook after securing a $20 billion OpenAI computing agreement. Anthropic is reportedly discussing a $6 billion acquisition of Decart to make inference more efficient, and Nebius has reported revenue growth of 454 percent.
The AI economy is therefore spreading beyond chip designers and cloud providers. Mistral is converting European enterprise commitments into “European Compute Units” to support a sovereign gigawatt-scale computing system. Vacuum-pump manufacturers, specialty-gas suppliers, battery companies, and energy producers are becoming part of the same investment story. Record US natural-gas output is helping supply data centers, while a Pentagon-backed loan is supporting silicon-anode batteries intended to reduce dependence on Chinese technology. A hydrogen-powered car reaching 406 miles per hour at Bonneville is a reminder that the energy transition is not only about decarbonization; it is also about supplying more power, more quickly.
At the consumer edge, the interface is becoming more ambient. The Pixel 11 reportedly includes agents that can order groceries and call businesses, while DeepMind’s 50-language SL2T model enables sign-language dictation. A smartwatch can infer insulin resistance without drawing blood. These systems promise convenience and accessibility, but they also turn ordinary environments into sources of continuous biometric data.
That prospect is provoking resistance. A German rights group has filed a criminal complaint over Meta’s AI glasses, warning that “there’s no place to escape from smart glasses.” A new analysis of consumer neurotechnology companies found that 29 of 30 provide unlimited access to users’ brain data. If cameras, microphones, and neural