跳到正文
原文
Goodfire Research·· 5 小時前AI 評分8

對抗樣本不是錯誤,而是疊加現象

Adversarial Examples Are Not Bugs, They Are Superposition

AI 導讀

文章主張,對抗樣本並非錯誤,而是疊加現象。頁面另列出以蛋白質嵌入向量改善 AI 智能體生物安全監測、規模化捕捉模型獎勵作弊,以及估算文字生成不確定性動態的研究。頁面亦介紹 Silico,稱其為可解釋性智能體,可用於解釋、除錯及精確控制模型行為,並提供可解釋性方法與基礎設施。

來源:Goodfire Research · goodfire.com