The most important AI upgrade may not be a larger benchmark score. It may be judgment.
Anthropic’s Claude Opus 5.5 reportedly matches Fable 5.1 on many tasks while costing 40 percent less than its predecessor. It leads several evaluations in coding, knowledge work, and computer use, and debuted at the top of the Artificial Analysis index. Yet users’ most enthusiastic reactions have focused elsewhere: the model can edit a startup video in about a minute, generate sophisticated artwork through code, animate interactive scenes in JavaScript, and build an Antikythera-inspired game in a single HTML file.
Those demonstrations point to a shift in what people expect from general-purpose AI. Models are no longer judged only by whether they can answer questions. They are increasingly evaluated on whether they can make coherent choices—about composition, interface design, narrative, and presentation. In other words, “taste” is becoming a practical capability.
That does not mean the systems possess human aesthetic understanding. Their results still depend on training data, prompts, and selection by users. But the ability to produce a usable first draft, recognize patterns of quality, and revise work across several media could reshape creative and professional workflows. Ten Opus agents reportedly spent 15 hours developing a shortest-path algorithm that asymptotically improves on Dijkstra’s algorithm. The result has been formally verified in Lean, although it is not yet practical. Such projects illustrate both the promise and the limitation: AI can explore large spaces of possibilities, but human researchers still have to decide which solutions matter.
The economics are changing just as quickly. OpenAI’s GPT-6 Sol and Luna are reported to cut API prices substantially, with Sol allegedly outperforming Opus 5 on some business workflows at a fraction of the cost. GPT-6 Astra, meanwhile, completed DrivingBench’s cone course in a real Toyota Corolla in 5 minutes 22 seconds, at a reported cost of $7.74. Fable 5.1 completed 45 percent of the same evaluation. These results support a broader trend: generalist models are becoming cheaper, more capable, and more useful in physical as well as digital environments.
That trend is challenging the assumption that specialized systems will always dominate. A general model that can reason, use tools, interpret images, and control a vehicle may eventually replace several narrower systems. But reliability remains the central question. A model can perform impressively in a controlled test and still fail unpredictably on unfamiliar roads, ambiguous instructions, or unusual operating conditions. The cost of a mistake may matter more than the cost of an API call.
Forecasting has also become a humbling exercise. Experts who expected machines to achieve gold-medal performance at the International Mathematical Olympiad around 2030 or later have seen those timelines compressed. AI labs’ revenues have also exceeded earlier estimates by wide margins. Newer evaluations attempt to measure progress more carefully. HLE-Diamond, a 1,000-question version of Humanity’s Last Exam, places GPT-6 Astra at 60.6 percent. An offline Kaggle system reportedly came within two percentage points of the ARC-AGI-2 grand prize.
Yet benchmark gains can obscure as much as they reveal. Models may optimize for the structure of an exam rather than the underlying ability researchers want to measure. One proposed solution to a Navier–Stokes challenge, generated by a 10,000-agent OpenAI system, was estimated to be months ahead of a single Astra run, but mathematicians questioned whether it addressed the intended problem or exploited a contrived force. The episode captures a recurring difficulty: as systems become more capable, evaluating them becomes a contest between scientific measurement and strategic adaptation.
The cost of intelligence is falling particularly fast. One estimate puts the price of a given level of capability on a roughly 13-fold annual decline. If that trend persists, tasks that currently require teams of analysts, programmers, or researchers could become accessible to small organizations and individuals. The social consequences will depend on who controls the infrastructure, who owns the resulting intellectual property, and whether workers receive a share of the productivity gains.
Some laboratories are already treating AI agents as research staff. Anthropic has opened a life-sciences effort in which roughly 950 Claude agents searched 200,000 reverse transcriptases and identified a new enzyme system with CRISPR-like arrays. AI-designed drugs are entering first-in-human trials faster than traditional programs—by as much as 80 percent in some accounts—although the hardest problem may remain choosing the correct disease mechanism. Anthropic and OpenEvidence are also expanding access to clinical AI in about 100 lower-income countries.
The scientific applications are broadening beyond biology. AI agents are being used to search mathematical spaces, reconstruct damaged ancient Greek papyri, and automate engineering simulations. The value of these systems is not simply that they work faster. They can pursue thousands of hypotheses in parallel, making previously impractical searches feasible. But speed does not remove the need for experimental validation. A plausible enzyme, proof, or medical hypothesis is only the beginning of a scientific process.
Governments are struggling to regulate technologies advancing faster than legislation. At the United Nations, the US administration has promoted “super intelligence” while rejecting global control. Twenty-two countries, excluding the United States, China, and the United Kingdom, have called for human control and floated the idea of an international agency modeled on the IAEA. In Washington, proposals range from strict oversight to a ban on superintelligence and a “corporate death penalty.”
The political conflict is also becoming entangled with national security and industrial policy. Officials are weighing export controls, data-routing investigations, and rules governing algorithmic decisions. A potential Trump–Xi summit is expected to focus as much on crisis communication as on slowing development. Meanwhile, reports of stress leave at the UK AI Security Institute and allegations that major donors placed allies inside regulatory institutions illustrate how difficult it is to separate technical governance from power.
Infrastructure is another constraint. Alibaba is expanding cloud regions from Turkey to Finland as it works toward 20 gigawatts of capacity. Texas has frozen some data-center permits, putting nearly 50 gigawatts of proposed capacity at risk. Chip licensing and geopolitical bargaining are creating new computing hubs, including a planned 70,000-GPU cluster in Armenia. Even quantum computing is beginning to look more robust: IonQ has demonstrated real-time error decoding using a conventional CPU.
The bottleneck is therefore not only algorithms. It is electricity, chips, cooling, land, skilled labor, and public consent. A world in which intelligence becomes cheap may still be one in which access to the machines that provide it remains concentrated among a handful of companies and states.
Public attitudes are more complicated than the usual narrative of either enthusiasm or panic. Gallup found optimism about AI exceeding fear in 34 of 37 countries, although the United States was among the more cautious wealthy nations, with only 36 percent expressing a positive view. Students are shifting away from computer science toward broader engineering, perhaps reflecting uncertainty about what entry-level programming work will look like. Schools are responding unevenly: New York City and Los Angeles have restricted classroom AI, while Alabama is introducing it in kindergarten.
The educational challenge is not merely cheating. One professor who identified 60 AI-generated essays in a single batch plans to eliminate written assignments. If automated text makes conventional homework unreliable, educators will need to redesign assessment around oral defense, project work, observation, and demonstrated reasoning. At the same time, speech itself is changing: one analysis found that daily speech declined by 338 words per year between 2005 and 2019, while increasingly capable text-to-speech systems can now generate almost any described voice in more than 100 languages. The danger is not that machines will speak too much, but that people may lose opportunities to practice thinking aloud.
AI is also moving into the physical world. Anduril won the autonomous track of the XPRIZE Wildfire competition by detecting and suppressing Alaskan fires. NASA-funded researchers have found an amoeba reproducing at 63 degrees Celsius, expanding the known limits of life. Scientists have engineered microbes intended to help build habitats on Mars, and transplant patients have remained rejection-free for five years on a single drug regimen. Even planetary science is revising old assumptions: Mercury appears to be shrinking faster than once estimated.
Taken together, these developments suggest that the AI story is no longer about a single race toward a hypothetical superintelligence. It is about a broad decline in the cost of cognition, arriving alongside advances in robotics, biotechnology, energy infrastructure, and scientific automation. The benefits could be enormous: faster discovery, cheaper services, and new tools for people who have never had access to expert labor.
But capability alone is not a strategy. The systems will need dependable evaluation, accountable deployment, and institutions capable of distinguishing genuine progress from impressive demonstrations. The central question is no longer whether machines can do more. It is whether societies can decide, quickly and fairly, what they should do—and who gets to decide.