arxivcs.LG2026-06-29
Quality-Aware Modulation for Diffusion Transformers
Luke Budny, Yuhong Guo, Kevin Cheung
Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep. While this modulation communicates the current noise level, it does not provide any quality-aware…