美團 LongCat 發佈 LongCat-2.0 MoE 大語言模型
meituan-longcat/LongCat-2.0
美團 LongCat 發佈 LongCat-2.0,這款 MoE 大語言模型總參數量為 1.6 trillion,每個 token 約啟用 48 billion 參數。
文中按編碼、Agent、搜尋及基礎能力列出 LongCat-2.0 與多款專有模型的基準分數,並標明部分分數引用自官方報告。
模型介紹
我們推出 LongCat-2.0,這是一款大規模 MoE 語言模型,總參數達 1.6 trillion,每個 token 啟用約 48 billion 個參數,較以往的 LongCat 模型大幅提升,並採用多項架構改良。
完整訓練過程及大規模部署均完全建基於 AI ASIC superpods。預訓練橫跨數百萬個 accelerator-days,涵蓋超過 35 trillion 個 token,期間沒有回退,也沒有無法恢復的損失尖峯,展現我們在替代硬件平台上進行前沿規模訓練的能力。
為提升模型處理長時程任務的能力,我們推出 LongCat Sparse Attention,並以數千億個 1M-context 資料 token 訓練 LongCat-2.0。配合專門的後訓練,LongCat-2.0 在編碼及 Agent 任務方面表現出色。
LongCat-2.0 深度整合 Claude Code、OpenClaw 和 Hermes 等主流 harness,在程式碼理解、儲存庫層級編輯、自動化任務執行及 Agent 工作流程等方面表現出色,為開發者帶來更穩定、高效的協作體驗。
主要功能
🌟 LongCat Sparse Attention
為解決 DSA 中 Lightning Indexer 的輸出不連續問題及二次方評分瓶頸,我們推出 LongCat Sparse Attention (LSA)。LSA 有三項互不重疊的改良:
- Streaming-aware Indexing (SI) 會重新配置 token 選取預算,結合與硬件對齊的連續存取及動態隨機選取。這將分散的記憶體存取轉化為可預測的循序讀取,實現合併 HBM 存取及高有效頻寬。
- Cross-Layer Indexing (CLI) 利用注意力顯著性在相鄰層之間的實證穩定性,分攤索引成本:在推理時,一次索引即可供數個連續層使用;訓練期間的跨層蒸餾則讓這項做法得以實現。
- Hierarchical Indexing (HI) 採用由粗至細的兩階段評分方案:先以區塊層級的近似評分進行粗略召回,再於召回的候選項目中精細選取 token,從而縮小索引器每次查詢須處理的候選範圍。
所有策略均可無縫延伸至用於推測解碼的 3-step Multi-Token Prediction 模組。對 CLI 而言,目標模型每 2 層共用一個索引,而全部 3 個 MTP 草稿步驟則共用一次索引處理。
🌟 N-gram Embedding
LongCat-2.0 承襲 LongCat-Flash-Lite 的 N-gram Embedding,透過在與 MoE 正交的稀疏維度擴展參數,提高參數使用效率。模型包含 135B 個 N-gram Embedding 參數,並遵循以下縮放原則:
- MoE 的稀疏度已超出最佳點。
- N-gram Embedding 的佔比控制在最佳範圍內。
這兩項原則確保 N-gram Embedding 相較於規模相當的純 MoE 模型,具有穩健優勢。
詳情請參閲我們的網誌。
評測結果
我們在 Agent 能力、程式編寫、搜尋、生產力及基礎能力方面,將 LongCat-2.0 與領先的專有模型進行評估。除非標有 *,否則所有分數均在統一評測框架下由內部測得。
基準測試 |
LongCat-2.0 |
Gemini 3.1 Pro |
GPT-5.5 |
Claude Opus 4.6 |
Claude Opus 4.7 |
Claude Opus 4.8 |
|---|---|---|---|---|---|---|
程式碼 Agent | ||||||
Terminal-Bench 2.1 |
70.8 |
70.7* |
73.8* |
- |
71.7* |
78.9* |
SWE-bench Pro |
59.5 |
54.2* |
58.6* |
57.3* |
64.3* |
69.2* |
SWE-bench Multilingual |
77.3 |
76.9* |
- |
77.8* |
80.5* |
84.8* |
通用 Agent | ||||||
FORTE ↗ |
73.2 |
70.3 |
77.8 |
73.2 |
77.6 |
77.2 |
BrowseComp |
79.9 |
85.9* |
84.4* |
84.0* |
79.3* |
84.3* |
RWSearch ↗ |
78.8 |
76.3 |
85.3 |
81.3 |
79.3 |
77.3 |
基礎能力 | ||||||
IFEval |
90.0 |
96.1 |
95.0 |
92.2 |
88.7 |
86.0 |
Writing Bench |
83.8 |
83.7 |
84.7 |
- |
85.3 |
85.2 |
IMO-AnswerBench |
81.8 |
90.0 |
79.5 |
75.3* |
81.8 |
75.3 |
GPQA-diamond |
88.9 |
94.3* |
93.6* |
91.3* |
94.2* |
92.4 |
備註:* — 引自模型的官方報告;- — 沒有可比較的公開分數。
\n\t\n\t\n\t\t聊天網站\n\t\n
你可在我們的官方網站 https://longcat.ai/ 與 LongCat-2.0 聊天。
\n\t\n\t\n\t\t部署\n\t\n
LongCat-2.0 可部署於 GPU 及 NPU 平台。
\n\t\n\t\n\t\tGPU\n\t\n
如要部署至 GPU,請參閲 SGLang 指南。
\n\t\n\t\n\t\tNPU\n\t\n
如要部署至 NPU,請參閲 SGLang-FluentLLM。
\n\t\n\t\n\t\t聊天範本\n\t\n
我們在 tokenizer_config.json 檔案中提供 LongCat-2.0 的聊天範本,可用來將訊息列表編碼為單一字串,作為模型輸入。
以下簡單示範如何使用此範本:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("meituan-longcat/LongCat-2.0", trust_remote_code=True)
tools = [
{
"type": "function",
"function": {
"name": "func_add",
"description": "Calculate the sum of two numbers",
"parameters": {
"type": "object",
"properties": {
"x1": {"type": "number", "description": "The first number to add"},
"x2": {"type": "number", "description": "The second number to add"},
},
"required": ["x1", "x2"],
},
},
},
{
"type": "function",
"function": {
"name": "func_multiply",
"description": "Calculate the product of two numbers",
"parameters": {
"type": "object",
"properties": {
"x1": {"type": "number", "description": "The first number to multiply"},
"x2": {"type": "number", "description": "The second number to multiply"},
},
"required": ["x1", "x2"],
},
},
},
]
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Calculate 1+1"},
{
"role": "assistant",
"reasoning_content": "Calling func_add to calculate 1+1",
# Note: unlike the standard OpenAI format, we expect `arguments` to be a dict rather than a string.
"tool_calls": [
{"type": "function", "function": {"name": "func_add", "arguments": {"x1": 1, "x2": 1}}},
],
},
{"role": "tool", "name": "func_add", "content": '{"ans": 2}'},
{"role": "assistant", "reasoning_content": "The result is 2", "content": "2"},
{"role": "user", "content": "Check your answer, is it correct?"},
]
# thinking mode on
prompt_think = tokenizer.apply_chat_template(
messages,
tools=tools,
tokenize=False,
enable_thinking=True,
add_generation_prompt=True
)
# thinking mode on, keeping all reasoning content for better performance
prompt_full = tokenizer.apply_chat_template(
messages,
tools=tools,
tokenize=False,
enable_thinking=True,
add_generation_prompt=True,
save_reasoning_content=True
)
# thinking mode off, for better token efficiency
prompt_no_think = tokenizer.apply_chat_template(
messages,
tools=tools,
tokenize=False,
enable_thinking=False,
add_generation_prompt=True
)
\n\t\n\t\n\t\t授權協議\n\t\n
模型權重以 MIT License 授權發布。
除非另有説明,對此儲存庫作出的任何貢獻均以 MIT License 授權。此授權不授予使用 Meituan 商標或專利的任何權利。
請參閲 LICENSE 檔案,瞭解完整授權條文。
\n\t\n\t\n\t\t使用注意事項\n\t\n
此模型並非專為所有可能的下游應用而設計,亦未針對這些應用進行全面評估。
開發者應考慮大型語言模型已知的限制,包括模型在不同語言中的效能差異,並在將模型部署於敏感或高風險情境前,仔細評估其準確度、安全性及公平性。\n開發者及下游用戶有責任瞭解並遵守與其使用情境相關的所有適用法律及規例,包括但不限於資料保護、私隱及內容安全要求。
本模型卡的任何內容均不應解讀為修改或限制模型發布時所依據的 MIT License 條款。
\n\t\n\t\n\t\t聯絡\n\t\n
如有任何疑問,請透過 longcat-team@meituan.com 聯絡我們,或提交 issue。
來源:美團 LongCat:HuggingFace 新模型 · huggingface.co