跳到正文
原文
Anthropic:Alignment Science Blog·· 4 天前AI 評分59

壓力測試模型規範揭示不同語言模型的行為特徵差異

Stress-testing model specs reveals character differences among language models

AI 導讀

Anthropic Fellows 與 Thinking Machines Lab 合作研究以超過 300,000 個價值取捨情境測試 12 個前沿模型,發現它們在價值排序與行為模式上存在差異。

來源:Anthropic:Alignment Science Blog · alignment.anthropic.com