跳到正文
原文
英國 AI Security Institute:Blog·· 3 小時前AI 評分69

AI 智能體評測為何需要納入測試時計算量:更多計算帶來更強能力

More compute, more capability: Why AI agent evaluations need to account for test-time compute

AI 導讀

英國 AI Security Institute 的 Science of Evaluation 團隊發現,固定計算預算會低估前沿 AI 智能體在多項基準上的能力,尤其是較新的模型。

來源:英國 AI Security Institute:Blog · aisi.gov.uk