Anthropic:Alignment Science Blog·· 4 天前AI 評分59
壓力測試模型規範揭示不同語言模型的行為特徵差異
Stress-testing model specs reveals character differences among language models
AI 導讀
Anthropic Fellows 與 Thinking Machines Lab 合作研究以超過 300,000 個價值取捨情境測試 12 個前沿模型,發現它們在價值排序與行為模式上存在差異。
來源:Anthropic:Alignment Science Blog · alignment.anthropic.com