跳到正文
Epoch AI:Gradient Updates· Elliot Stewart·· 2 小時前AI 評分57

Epoch AI 發佈 2026 年 10 月 8 日《The Epoch Brief》,彙整 AI 研究與基準更新

The Epoch Brief - October 8, 2026

AI 導讀

Epoch AI 於 2026 年 10 月 8 日發佈《The Epoch Brief》,彙整 AI 智能體規模、成本、產業影響、使用情況及基準測試等研究更新。

正文

Welcome back to the Epoch Brief. In this edition, we’re covering:

  • Proliferating agents and plummeting costs. Our new report estimates how many AI agents could soon exist in the world: nearly 2 billion at the upper bound, assuming agents run on cheap models. Another report shows that prices for a given level of performance are dropping faster for AI than for any other transformative technology in history.

  • New research on China’s AI developers and economy. How developers in China’s heavily open-weight ecosystem make money, and research showing China is more exposed to semiconductor supply-chain disruptions than the US. Plus, trade data suggesting $3B of chips were smuggled to China through Malaysia.

  • How AI is impacting people and changing behaviors. New data on whether advancing cyber capabilities are affecting people's lives, AI's growing role in math publications, near-daily AI use more than doubling, and OpenAI's soaring internal spending on AI coding. Plus our new ChatGPT Usage explorer, built on metadata from 8.3 million messages.

  • Updates from our growing benchmarking teams. The upward trend on the Epoch Capabilities Index continues unabated; models have gotten a lot better at spotting mistakes in how you assemble your IKEA furniture; GPT-6 Astra discovered a creative way to break our board game benchmark; and our newest benchmark, InnovationEval, tests whether frontier AI can automate AI R&D (short answer: not yet).

  • The launch of Benchmark Reviews. Our AI benchmark auditing initiative to help provide a consistent source of information on benchmark quality.

  • New open roles. We continue to grow, with multiple open roles across research, design, product, and communication.

Subscribe now

Proliferating agents and plummeting costs

In How many AI agents could run on the AI chips shipped through 2027?, researcher Jason Li estimates that the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, which, running nonstop, would supply as many weekly working hours as roughly 140 to 720 million full-time human employees. Running cheaper models instead, the same chips could support almost 2 billion concurrent agents. Whether demand for such a large agent population will exist is the open question.

The demand question will hinge a lot on how capable AI agents are and how much that capability costs. In The plunging price of thought, researchers David Roodman and Luke Emberson find that the cost of achieving a fixed level of AI performance has fallen about 47% per quarter over the past three years, around 13× per year. That means AI is getting cheaper faster than any other transformative technology they compare against, four times faster than DNA sequencing and 18 times faster than lithium batteries.

Share the Epoch Brief to help our work reach more people.

Share

New research on China’s AI developers and economy

In How do Chinese AI companies make money?, senior Epoch researcher Anson Ho and GovAI research scholar Cheryl Wu find that six leading Chinese AI companies (Alibaba, ByteDance, Z.ai, Moonshot, DeepSeek, and MiniMax) collectively generate roughly one-tenth the AI-related revenue of OpenAI and Anthropic combined.

In China is more exposed to semiconductor supply-chain disruptions than the US, researcher Daniel Carey estimates China is more exposed than the US to semiconductor supply disruptions, because semiconductors account for a larger share of the cost of the goods it buys: for every $1,000 of Chinese final demand, $15.20 flows to chipmakers, versus $5.70 for the US. By our measure, China’s exposure to potential semiconductor supply shocks is 2.7× that of the US. Carey also simulates how different shock scenarios affect prices in China and other countries.

Our economics team also investigated global trade data and found evidence consistent with more than $3 billion of chips smuggled into China via Malaysia. Between April 2024 and June 2025, China recorded $3.8 billion of server imports from Malaysia, at AI server prices of $80,000–$150,000 per unit, yet Malaysia recorded only $0.6 billion of exports to China. While this does not prove diversion, the pattern is consistent with established cases of chip smuggling into China, in which intermediaries have routed AI servers through Malaysia and declared them on export paperwork as ordinary servers.

How is AI impacting people and changing behaviors?

The spike in serious cyber vulnerability disclosures hasn’t translated to more personal cyber incidents. In our latest polling with Ipsos, we found that the share of US adults reporting a cyber incident was essentially unchanged between June and September 2026 (46% vs 45%). This contrasts with the spike in reported serious cyber vulnerabilities over the same period, when advanced AI cyber capabilities became more widely available. We detailed that spike at the time, and maintain open data on disclosures in our Cyber Vulnerabilities explorer.

More AI assistance in math publications. AI capability advances have had dramatic impacts on mathematics in recent months, including on our FrontierMath: Open Problems benchmark. But beyond more AI-generated solutions, AI is also changing how human mathematicians are producing mathematical knowledge. Our analysis of every math preprint posted to arXiv since January 2025 shows acknowledgments of AI use jumping from 4% in April to 25% in August 2026, with around 6% crediting AI at the level of a coauthor.

Share

We’re also tracking intensifying AI use in the broader public. Our latest polling with Ipsos shows the share of US adults who report using AI six to seven days during the previous week rose from 8% in March to 19% in August 2026, while the share reporting use just one day a week fell from 17% to 10%.

These polling results are complemented by findings from our new ChatGPT Usage explorer, an interactive resource built on aggregate metadata from 8.3 million messages across 660,000 conversations, shared by 5,000 US YouGov panelists. It shows median monthly messages per active user climbing from 14 in January 2023 to 36 by December 2025, with the most active 10% of panelists sending 63% of all messages. Participation was opt-in, and the data is unweighted, meaning it has not been adjusted to match the demographics of US adults or of ChatGPT users as a whole. You can explore the data on the ChatGPT Usage hub.

OpenAI’s internal AI spending is soaring. Perhaps the most dramatic trend in AI use is within AI labs themselves. Based on OpenAI’s September research post, the median OpenAI researcher’s daily coding-agent usage, valued at API prices, grew from under $1 in January 2026 to about $600 by mid-August, with 90th-percentile users spending over $7,000 per day.

So how good is AI getting?

The upward trend on the Epoch Capabilities Index (ECI), our measure of general capabilities, continues unabated.

And our benchmarking team is designing tests to understand AI capability progress across a number of consequential domains, from spatial reasoning to continuous learning to autonomous AI R&D that many believe could trigger recursive self-improvement (RSI).

Using IKEA furniture to test AI’s spatial reasoning.

We built our new Furniture Assembly Benchmark from 60 photos of three IKEA builds. We seeded the photos with deliberate assembly errors and asked models to find every mistaken step. The top score on our spatial reasoning benchmark jumped from 28% to 80% in 10 months.

Using board games to measure continuous learning.

We provided an update on EBR-bench, a benchmark that we designed to test models’ ability to learn from experience by playing the board game Earthborne Rangers. GPT-6 Astra initially achieved a perfect score by exploiting an overpowered card that enables unlimited turns; with that card banned, it averages 16/21, still about 50% above the previous best of 10.5/21 set by Claude Opus 5. Astra remains slightly behind top human players in tactical decision-making.

Can AI automate AI R&D yet?

AI developers are aiming to automate AI researchers. Our newest evaluation, InnovationEval, tests how close they are by tasking models with discovering a post-training innovation comparable in magnitude to a recently published advancement. Given a budget of 3,000 GPU-hours, no AI model came close. GPT-5.6 Sol managed only incremental tweaks resembling prior work, Claude Fable 5 gamed the setup by selecting its best results across runs, and both models made misleading claims about their achievements.

How good are benchmarks, though?

Our work to better understand AI capabilities also includes an effort to provide the community with a consistent source of information on benchmark quality. Benchmark Reviews is our initiative to audit AI benchmarks. We launched with reviews of 15 benchmarks, grading four as verified, nine as flawed, and two as having insufficient information for a review. This is an ongoing project with new reviews available on our website.

Each review is of a specific benchmark version, and if a benchmark is updated to fix errors, we may review the new version. We share findings with developers and will publish a link to their response, if they’d like. You can also find the full auditing methodology on our site.

Share

Careers

Epoch is growing quickly, and we’re looking for exceptional people to help us meet the moment. We’re hiring for:

Not seeing a fit? Submit an Expression of Interest. Applications are rolling, so apply soon!

Thanks for reading! Subscribe for free to receive new posts.

來源:Epoch AI:Gradient Updates · epochai.substack.com