Diffusion models have achieved significant progress in image generation, with backbone architectures evolving from U-Net to Transformers. However, the quadratic complexity of Transformer-based diffusion models limits their scalability and efficiency, and this limitation becomes more prominent with increasing resolution. Linear complexity models such as Mamba partially address this issue but struggle with spatial continuity when applied to two-dimensional image data. To tackle these challenges, we propose T4Di, a hybrid backbone architecture combining the efficiency of Test-Time Training (TTT) with the global modeling capability of Transformers. By introducing multidirectional scanning and lightweight local feature enhancement modules, T4Di adapts TTT to 2D image signals, improving spatial continuity and local coherence. Moreover, we explore adaptive block composition, adjusting the ratio between Transformer and TTT components to achieve a favorable balance between generation quality and computational cost. We evaluate T4Di on both unconditional and class-conditional image generation tasks across CIFAR-10, CelebA, and ImageNet benchmarks. Experimental results demonstrate that T4Di consistently outperforms existing diffusion models in terms of both generation quality and computational efficiency, establishing it as a scalable and effective solution for image synthesis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

T4Di: A Hybrid TTT-Transformer Backbone for Scalable and Efficient Diffusion Model

  • Xirui Wu,
  • Haixia Pan,
  • Ruijun Liu,
  • Biao Dong,
  • Ying Zheng,
  • Huolong Ye

摘要

Diffusion models have achieved significant progress in image generation, with backbone architectures evolving from U-Net to Transformers. However, the quadratic complexity of Transformer-based diffusion models limits their scalability and efficiency, and this limitation becomes more prominent with increasing resolution. Linear complexity models such as Mamba partially address this issue but struggle with spatial continuity when applied to two-dimensional image data. To tackle these challenges, we propose T4Di, a hybrid backbone architecture combining the efficiency of Test-Time Training (TTT) with the global modeling capability of Transformers. By introducing multidirectional scanning and lightweight local feature enhancement modules, T4Di adapts TTT to 2D image signals, improving spatial continuity and local coherence. Moreover, we explore adaptive block composition, adjusting the ratio between Transformer and TTT components to achieve a favorable balance between generation quality and computational cost. We evaluate T4Di on both unconditional and class-conditional image generation tasks across CIFAR-10, CelebA, and ImageNet benchmarks. Experimental results demonstrate that T4Di consistently outperforms existing diffusion models in terms of both generation quality and computational efficiency, establishing it as a scalable and effective solution for image synthesis.