跳到正文
原文
Zyphra Research·· 5 小時前AI 評分29

Zyphra 公佈 Zyda:供語言模型使用的 1.3T tokens 開放數據集

ResearchZydaZyphra is pleased to announce Zyda, a 1.3T trillion-token open dataset for language modeling. Zyda combines the existing suite of high-quality open datasets together and merges them through a uniform and thorough filtering and deduplication process. The goal of Zyda is to provide a simple, accessible, and highly performant dataset for language modeling experiments and training up to the 1 trillion scale. In our ablation studies, Zyda outperforms all existing open datasets including the Dolma, Fineweb, Pile, RefinedWeb, and SlimPajama.June 7, 2024

AI 導讀

Zyphra 公佈 Zyda,一個包含 1.3T tokens、供語言模型訓練使用的開放數據集。它整合七個開放數據集,經統一的品質過濾及跨數據集去重,並以開放且寬鬆的授權發布。消融研究顯示,Zyda 表現勝過 Dolma、Fineweb、Pile、RefinedWeb 和 SlimPajama 等現有開放數據集。

來源:Zyphra Research · zyphra.com