The computational demands of latent diffusion models hinder their practical deployment despite remarkable generation capabilities. We present a comprehensive optimization framework that synergistically enhances both the UNet denoiser and VAE decoder in Stable Diffusion. Our key innovation lies in a spatial-aware architecture co-design addressing two fundamental bottlenecks: For the UNet, we propose Spatial-Reduced Attention (SRA) that strategically manipulates feature map dimensions. By replacing standard 1 × 1 convolutions with strided 3 × 3 filters for query/key/value projections, followed by low-rank attention computation in reduced space and transposed convolution-based resolution recovery. The decoder optimization introduces Multi-Space Distillation, a data-free compression paradigm that leverages synthetic data generation from teacher’s latent space. Through hybrid supervision combining pixel-level reconstruction, perceptual feature alignment, and latent cycle consistency, our distilled decoder attains parameter reduction without quality degradation. Theoretical analysis reveals our spatial-adaptive processing minimizes Wasserstein distance between original and optimized models, establishing a new efficiency-quality Pareto frontier for diffusion-based generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DiffEngine: Holistic Optimization of Attention and Decoder in Stable Diffusion

  • Zebing Wei,
  • Tao Zhang,
  • Biyi Chen,
  • Chengguang Wang

摘要

The computational demands of latent diffusion models hinder their practical deployment despite remarkable generation capabilities. We present a comprehensive optimization framework that synergistically enhances both the UNet denoiser and VAE decoder in Stable Diffusion. Our key innovation lies in a spatial-aware architecture co-design addressing two fundamental bottlenecks: For the UNet, we propose Spatial-Reduced Attention (SRA) that strategically manipulates feature map dimensions. By replacing standard 1 × 1 convolutions with strided 3 × 3 filters for query/key/value projections, followed by low-rank attention computation in reduced space and transposed convolution-based resolution recovery. The decoder optimization introduces Multi-Space Distillation, a data-free compression paradigm that leverages synthetic data generation from teacher’s latent space. Through hybrid supervision combining pixel-level reconstruction, perceptual feature alignment, and latent cycle consistency, our distilled decoder attains parameter reduction without quality degradation. Theoretical analysis reveals our spatial-adaptive processing minimizes Wasserstein distance between original and optimized models, establishing a new efficiency-quality Pareto frontier for diffusion-based generation.