跳到正文
原文
Anthropic:Alignment Science Blog·· 4 天前AI 評分71

SLEIGHT-Bench 用於找出 AI 監控器的盲點

SLEIGHT-Bench: Finding Blind Spots in AI Monitors

AI 導讀

Anthropic Fellows Program 研究團隊發佈 SLEIGHT-Bench,以 40 組合成攻擊測試前沿 AI 監控器的 11 類盲點。在 1% 假陽性率下,Opus 4.6 監控器有 50% 的攻擊在 10 次測試中一次也未被捕捉,只有 8/40 組攻擊能穩定檢出。

來源:Anthropic:Alignment Science Blog · alignment.anthropic.com