Goodfire Research·· 5 小時前AI 評分8
對抗樣本不是錯誤,而是疊加現象
Adversarial Examples Are Not Bugs, They Are Superposition
AI 導讀
文章主張,對抗樣本並非錯誤,而是疊加現象。頁面另列出以蛋白質嵌入向量改善 AI 智能體生物安全監測、規模化捕捉模型獎勵作弊,以及估算文字生成不確定性動態的研究。頁面亦介紹 Silico,稱其為可解釋性智能體,可用於解釋、除錯及精確控制模型行為,並提供可解釋性方法與基礎設施。
來源:Goodfire Research · goodfire.com