跳到正文
Cohere Labs:官方研究博客·· 2 小時前AI 評分35

為模型排名動態分配評估工作量

Dynamically Allocating Evaluation Effort for Model Ranking

AI 導讀

研究將多模型人工評估形式化為具相關臂的多臂 bandit 最佳臂識別問題,並提出按中途排名自適應分配評估工作的演算法。這種取樣方式把標註預算集中於最具競爭力的模型。研究證明所提演算法具最優性,並顯示可加快評估、降低成本及提升對頂尖模型的區分能力。

正文

Aug 04, 2026

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach.

Authors

Vilém Zouhar, Julia Kreutzer, Alon Lavie, Tom Kocmi, Matt Post, Ondřej Bojar, Mrinmaya Sachan

Abstract

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.

Related works

來源:Cohere Labs:官方研究博客 · cohere.com