Google DeepMind·· 2025-12-09AI 評分59
FACTS Benchmark Suite 系統評估大語言模型事實準確性
FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
AI 導讀
Google DeepMind 與 Kaggle 推出 FACTS Benchmark Suite,透過四項基準評估大語言模型的事實準確性。套件包含更新版 Grounding Benchmark - v2,以及 Parametric、Search、Multimodal 三項基準,並公開提供 3,513 個例子;Kaggle 管理私有保留集並維護公開排行榜。
來源:Google DeepMind · deepmind.google