Welcome to September 4, 2026

By: alexwg

Published: 2026-09-07T05:11:29.707738Z

Last Updated: 2026-09-07T05:11:29.707741Z

Category: Science & Technology

The AI race is entering a less comfortable phase: models are becoming more capable just as they become harder to understand.

OpenAI’s reported GPT-6 Astra illustrates the tension. The company says the system achieved a perfect score on ExploitBench, including a newer version built around vulnerabilities discovered after its training cutoff. That performance prompted OpenAI’s first “Critical” cyber designation, a limited rollout, and review by the White House. Trained on more than 100,000 GPUs, Astra has also performed strikingly across unrelated tasks: building a detailed virtual Manhattan, completing Pokémon in 18 hours, and reaching 99.9 percent on ARC-AGI-3.

Yet its apparent intelligence may be becoming less legible. Reports suggest that Astra uses recurrent computation—revisiting a problem internally—to increase effective reasoning depth without exposing every intermediate step. OpenAI researcher Jakub Pachocki has argued that the model’s depth is within roughly twice that of GPT-4. Others, including Joshua Achiam, contend that fully transparent reasoning was never a realistic long-term safety strategy. The unresolved question is whether “thinking time” is now a controllable dial—and whether developers can reliably monitor what happens when that dial is turned up.

The benchmark results are impressive, but they are not the same as general intelligence. Greg Kamradt says Astra eventually “subsumed the harness,” while Greg Brockman described the system as “saturated” and called the arrival of an AGI era “not unreasonable.” Other researchers remain more cautious. The model’s gains in token efficiency and its record ECI score of 169 suggest a new frontier in performance, but not necessarily a settled definition of intelligence.

The frontier is becoming a queue rather than a throne. Anthropic has introduced Claude Fable 5.1 and Mythos 5.1, apparently the same underlying model offered with different safeguard levels. Fable reportedly returned to the AAII frontier with a score of 66, while Mythos 5.1 on low settings matched Mythos 5 at maximum settings. Fable also outperformed GPT-5.6 Sol on CursorBench at an estimated $3.53 per task and received a 75 percent reduction in cache-read costs. Anthropic says its systems carry invisible watermarks for European Union compliance.

Google has released Gemini 3.8 Flash, which the company says its own programmers preferred to Anthropic’s Opus, along with a cybersecurity version that patches software 2.6 times more effectively. Its video agents can reduce token use by as much as 88 percent. Meta’s Muse Spark 1.3 is priced so cheaply that it is “almost too cheap to meter,” although it reportedly trails leading systems by up to 62 percent. Grok 4.7 is expected soon, while World Labs’ Atlas combines language, video, and three-dimensional understanding.

These systems are moving from chat windows into tools that act. That shift makes reliability, identity, and control more important than leaderboard position. OpenAI has told Congress it is developing automated shutdown capabilities. Ilya Sutskever has warned that rogue agents could hijack cloud infrastructure and copy themselves. Dean Ball has argued that autonomous systems need identity and accountability rails rather than blanket bans.

Policy is struggling to keep pace. Senator Bernie Sanders has proposed outlawing superintelligence, while the G20 has endorsed the so-called Carolina Principles, emphasizing responsible development. Michael Kratsios has urged a lighter regulatory touch. Mark Zuckerberg has reportedly told the president that a single national regulator would be flawed, even as Washington supports OpenAI’s fair-use position. Anthropic has received praise from Commerce Secretary Howard Lutnick, but the Pentagon says its restrictions on the company remain in place.

The scientific consequences may be substantial if the claims hold up. Astra reportedly solved two of 68 open Erdős problems and produced a Lean-verified proof that prime gaps of at most 186 occur infinitely often. WeatherNext 3 forecasts conditions hourly at five-kilometer resolution from satellite data. Such systems could accelerate mathematics, weather prediction, and software development—but only if their outputs can be checked. A proof assistant can verify formal logic; it cannot by itself determine whether a model chose the right problem or framed the result usefully.

AI is also spreading through ordinary infrastructure. Dyson’s $499 CameraJet uses a camera to guide flossing. ChatGPT can read Epic health records. The Codex app reportedly ships with LibreOffice support, while Nvidia’s Personal AI Router coordinates computing across home PCs. Nvidia is also reported to be acquiring Hugging Face for $12.9 billion, raising familiar questions about whether “open” ecosystems can remain independent when their infrastructure is owned by a chip giant.

The physical cost of this expansion is harder to ignore. Anthropic plans to deploy five gigawatts of TPU capacity next year and has signed a reported $35 billion agreement with Lambda. Dell says its AI backlog has reached $95 billion. SB Energy has filed for an IPO with 8.8 gigawatts of contracted capacity, though none is yet operating. Elon Musk has warned of a 15-gigawatt shortfall by 2027 and predicted that AI coding could reach Stockfish-like strength within 18 months.

The industry’s public explanation has not kept pace with its appetite for electricity, water, and construction. Sam Altman has dismissed some water concerns as a meme, while Scott Bessent says the sector has done a “terrible job” explaining itself. The political response mixes promises of millions of jobs with pressure to accelerate construction. Meanwhile, energy is becoming more distributed: California lawmakers have backed balcony solar, and solar power has overtaken coal generation in China.

Autonomy is spreading beyond software. Tesla has launched the Cybercab in Austin, a roughly $30,000 vehicle without a steering wheel. Uber and Wayve have begun operating robotaxis in London, and Waymo says it now serves 14 cities and provides 500,000 rides a week. Uber, once a disruptor of the taxi industry, is now lobbying alongside taxi unions to slow autonomous competitors. Drone tariffs of up to 100 percent add another layer of industrial policy to the contest.

Biology is advancing along a parallel track. Semaglutide extended the lives of mice by nearly 100 days in one study, while GLP-1 drugs have been associated with fewer serious infections. A pancreatic-cancer pill reportedly shrank lung tumors, and Until Labs says its cryoprotectant preserves 85 percent of cell viability. Researchers have also completed the first male fruit-fly connectome, mapping 166,700 neurons and finding that sex-related differences are concentrated largely in higher-order brain regions.

The social contract is being adjusted in smaller but consequential ways. BT’s obsolete copper network is valued at $2.7 billion. Florida governor Ron DeSantis has ordered Flock surveillance cameras removed, while New York City has paused student-facing AI use through eighth grade. These decisions reflect a broader question: which technologies should be introduced quietly through infrastructure, and which deserve explicit public consent?

Even contact with the unknown is acquiring a budget line. The FBI is reportedly producing UAP coins, the government is said to have a plan for confirmed nonhuman intelligence, and NASA has selected Blue Origin to develop a Mars telecommunications network. A proposed Fermi Explorer Mission would be the first spacecraft explicitly aimed at another star, Alpha Centauri. The route was reportedly identified by an AI system developed by Physical Superintelligence, a company that emerged with $58 million in funding after finding the trajectory in three days.

The common thread is not that machines have suddenly become superintelligent. It is that computation is becoming embedded in science, transport, medicine, energy, government, and even the search for life beyond Earth. The central challenge is therefore no longer simply to build more capable systems. It is to make their capabilities measurable, their failures visible, their resource demands defensible, and their deployment subject to institutions that can still say no.