Artificial Analysis 將可信存取模型納入 Cyber Index,GPT-6 Sol (Daybreak Blue, max) 登頂
Introducing trusted-access models to the Artificial Analysis Cyber Index
Artificial Analysis 擴大 Cyber Index,納入防護限制較少的可信存取模型;首個納入的 GPT-6 Sol (Daybreak Blue, max) 目前領先榜單。
GPT-6 Sol (Daybreak Blue) now leads the Cyber Index
The Artificial Analysis Cyber Index Alliance brings together industry partners to evaluate how AI models perform on enterprise cyber defense tasks. The Index launched with only publicly accessible models, giving users the best indication of publicly available capability on agentic cyber defense. Now, we’re expanding the leaderboard to include models with fewer cyber guardrails, providing a broader view of frontier model cyber capability. The first of these is GPT-6 Sol (Daybreak Blue, max) which is only accessible through OpenAI’s Daybreak program

GPT-6 Sol (Daybreak Blue, max) now leads the Index and sits on the Cost vs. Capability Pareto frontier. It hits no safety blocks across the entire Index, with its overall score improving by 32 points over the publicly available GPT-6 Sol (max)
At a Cost per Task of $1.77, it is more cost-effective than other leading models, including Grok 4.7 (xhigh) which costs $11.67 per task

Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

We will continue to expand the Artificial Analysis Cyber Index to cover both publicly available and trusted-access models
For full evaluation results, see here: https://artificialanalysis.ai/evaluations/artificial-analysis-cyber-index
來源:Artificial Analysis 完整文章 · artificialanalysis.ai