跳到正文
HuggingFace Daily Papers(社區熱門論文)·· 7 天前AI 評分51

DMAD:以對抗蒸餾進行分佈匹配,實現快速視覺生成

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

AI 導讀

論文提出 DMAD,以判別器分類方式進行分佈匹配蒸餾,直接學習對數密度比,省去輔助分數擬合。在 ImageNet-64x64 上,DMAD 一步生成的 FID 為 1.04;四步 SDXL 在 COCO-10K 的 FID 為 14.47,四步 Wan2.1-T2V-14B 的 VBench 總分為 85.15,均為文中比較的少步方法及多步教師模型中的最佳成績。

正文

Published on Oct 1

Authors:

,

,

,

,

,

,

,

,

,

Abstract

Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.

View arXiv page View PDF Project page GitHub Add to collection

Models citing this paper 2

ZhengmingYu/DMAD

Text-to-Video •

Updated about 21 hours ago

•

79

nixsn/DMAD

Text-to-Video •

Updated 2 days ago

Datasets citing this paper 1

ZhengmingYu/DMAD-H3-data

Preview •

Updated about 24 hours ago

•

152

•

2

Spaces citing this paper 3

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co