Zyphra Research 介紹 PUFFER 增量模糊去重系統及其效能測試
ResearchPUFFER: Incremental Fuzzy Deduplication for Continuously Evolving CorporaZyphra introduces PUFFER, a provenance-aware incremental fuzzy-deduplication system for continuously growing training corpora. PUFFER maintains a faithful, disk-resident MinHash locality-sensitive hashing history that can ingest new datasets, recover deterministically from interrupted jobs, and remove individual datasets without repeatedly rebuilding the full corpus index. PUFFER achieves an 11x-35x speedup over existing CPU-based deduplication methods when tasked with processing up to one billion documents over a ten-hour throughput test. At the deduplication (index) layer, PUFFER ingests one billion documents in 1.75 hours on a single process, and is deployed in production on more than 30 billion documents. Available as open-source software under a permissive Apache 2.0 license.September 1, 2026
Zyphra Research 介紹 PUFFER,一套為持續增長的訓練語料設計的增量模糊去重系統,可逐批加入數據集而無須重建完整索引。
來源:Zyphra Research · zyphra.com