Scalable Diffusion Models with Transformers
William Peebles, Saining Xie
Berkeley College University of California, Berkeley New York University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
We explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops—through increased transformer depth/width or increased number of input tokens—consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512×512 and 256×256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
生物医学Advanced Neuroimaging Techniques and Applications
Generative Adversarial Networks and Image Synthesis · Music and Audio Processing
参考文献 59
此处列出前 3 条
引用本文 1,941
按被引量排序,此处列出前 3 条