TexStFusion : a controllable diffusion model using textural, structural, and textual feature fusion
摘要
Recent advances in Text-to-Image (T2I) diffusion models enable highly realistic image generation from text. However, long and intricate descriptions often struggle to provide precise controls. To address this, we propose TexStFusion (TEXtural, STructural, TEXtual feature FUSION), a method that adds conditional controls to pre-trained T2I models. Unlike existing approaches relying on visual cues, we introduce composite maps, which fuse texture and structure-text maps derived from TextureNet and StructureNet encoders. This integration occurs without fine-tuning the T2I model, preserving prior knowledge. Our method achieves 25% better FID, 33% better SSIM, and 5% better CLIP-T scores with a dataset of just 30k images, in the best case.