Diffusion Transformers with Representation Autoencoders VAE (e.g., DINO, SigLIP, MAE) with (RAEs)
Автор: Byte Goose AI.
Загружено: 2025-10-14
Просмотров: 281
Описание:
Diffusion Transformers with Representation Autoencoders.
This podcast provides an overview and detailed study guide for a research paper on Representation Autoencoders (RAEs), which are proposed as a superior alternative to traditional Variational Autoencoders (VAEs) for use with Diffusion Transformers (DiTs). The core innovation is replacing VAEs with pretrained representation encoders (like DINO) coupled with lightweight decoders, creating semantically rich and highly efficient latent spaces that solve the VAE limitations of outdated backbones and low capacity. Because RAEs produce high-dimensional latents, the paper introduces critical modifications for DiTs, including matching model width to the token dimension, implementing a dimension-dependent noise schedule shift, and using noise-augmented decoding to ensure stability and efficiency. A new architecture, DiTDH (DiT with a shallow, wide head), is shown to scale more efficiently and achieve state-of-the-art FID scores on ImageNet, such as 1.13 with guidance.
https://arxiv.org/pdf/2510.11690
Повторяем попытку...
Доступные форматы для скачивания:
Скачать видео
-
Информация по загрузке: