評測來源官方評測
Arena Creative Writing
LMArena · 真人盲選故事、詩歌和其他創意表達,觀察作品是否吸引讀者。
在 ATech Hot 中參與排名
證據預算5%
上游數據時間10/02 08:00
最近成功同步10/06 14:05
它測什麼,怎麼測
讀取官方數據集 text_style_control 的 creative_writing 分類,採用風格控制後的 Bradley–Terry 分與置信區間;不借用整體文本或職業寫作分類成績。
如何使用這份證據
正式計分;從原有真人偏好 10% 中拆分 5%,整體文本保留 5%,總預算不增加。與其他 Arena 任務共用證據家族。
評測成績
按固定規則,每個公開模型採用一套代表配置;沒有可採用配置的模型標為未計入並寫明原因。匿名測試型號不展示,保留原榜名次,每項最多 30 個。
| 原榜名次 | 原榜型號 | 原始成績 | 代表配置 |
|---|---|---|---|
| 1 | gemini-4-argon-highgoogle | 1,518.8 | High 推理 |
| 2 | claude-opus-5.5-highAnthropic | 1,516.3 | High 推理 |
| 3 | claude-fable-5-highAnthropic | 1,502.6 | High 推理 |
| 4 | claude-opus-4-6-highAnthropic | 1,500.8 | High 推理 |
| 5 | gemini-3.7-flash-highGoogle | 1,493.4 | High 推理 |
| 6 | claude-opus-4-7-highAnthropic | 1,489 | High 推理 |
| 7 | gemini-3.8-flash-highGoogle | 1,484.5 | High 推理 |
| 8 | gemini-3-proGoogle | 1,484.4 | 來源預設配置 |
| 10 | claude-fable-5.1-maxAnthropic | 1,481.8 | Max 推理 |
| 11 | gemini-3.1-pro-previewGoogle | 1,480.3 | 來源預設配置 |
| 14 | gemini-3.6-flash-highGoogle | 1,470.9 | High 推理 |
| 15 | claude-opus-5-maxAnthropic | 1,470.6 | Max 推理 |
| 16 | claude-opus-4-5-20251101-high-32kAnthropic | 1,469.5 | High 推理 · 32k |
| 17 | claude-opus-4-8-highAnthropic | 1,469.5 | High 推理 |
| 19 | gpt-5.6-sol-xhighOpenAI | 1,468.3 | xHigh 推理 |
| 20 | qwen3.8-maxAlibaba | 1,467.9 | 來源預設配置 |
| 21 | muse-sparkMeta | 1,465.5 | 來源預設配置 |
| 22 | gemini-3.5-flash-highGoogle | 1,463.9 | High 推理 |
| 24 | grok-4.20-beta1xAI | 1,463 | 來源預設配置 |
| 26 | kimi-k3-maxMoonshot AI | 1,460.2 | 來源預設配置 |
| 27 | gpt-6.1-sol-maxopenai | 1,459.7 | Max 推理 |
| 28 | muse-spark-1.3-maxMeta | 1,458.7 | Max 推理 |
| 29 | gemini-3-flashGoogle | 1,456.8 | 來源預設配置 |
| 30 | gpt-5.5-instantOpenAI | 1,456.4 | 來源預設配置 |
| 31 | muse-spark-1.2 (xHigh)Meta | 1,455.5 | xHigh 推理 |
| 32 | glm-5.2-maxZ.ai | 1,454.2 | 來源預設配置 |
| 34 | glm-5.3-maxZ.ai | 1,453.9 | 來源預設配置 |
| 35 | gpt-5.5-highOpenAI | 1,451.3 | High 推理 |
| 36 | grok-4.5xAI | 1,450.2 | 來源預設配置 |
| 37 | claude-sonnet-4-5-20250929-high-32kAnthropic | 1,450 | High 推理 · 32k |
評測侷限與數據署名
開放用户提示由分類器歸類,並非全部為長篇小説或中文創作;投票包含個人偏好。與 Arena 整體文本及職業寫作分類存在重疊,不能作為額外獨立機構。
數據許可:CC BY 4.0 · Arena 官方 Hugging Face leaderboard-dataset;保留署名、來源鏈接並説明聚合改動。
成績由 LMArena 發佈,原始分數與 ATech Hot 共識分使用不同尺度,不能直接相加。