跳到正文
原文
Google DeepMind·· 2025-12-09AI 評分59

FACTS Benchmark Suite 系統評估大語言模型事實準確性

FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

AI 導讀

Google DeepMind 與 Kaggle 推出 FACTS Benchmark Suite,透過四項基準評估大語言模型的事實準確性。套件包含更新版 Grounding Benchmark - v2,以及 Parametric、Search、Multimodal 三項基準,並公開提供 3,513 個例子;Kaggle 管理私有保留集並維護公開排行榜。

來源:Google DeepMind · deepmind.google