MCT-Net: Multiscale Convolution-Transformer Network for Defect Image Generation Using Segmentation Maps
摘要
Industrial defect detection is a critical component of quality assurance in manufacturing. While generative modeling has shown promise in synthesizing artificial defect images to enhance anomaly detection training, existing methods often lack structured control over defect spatial distribution. To address this limitation, we propose MCT-Net (Multi-scale Convolutional-Transformer Network), a pixel-level controlled defect generation framework that combines the complementary strengths of Convolutional Neural Networks (CNNs) for local feature extraction and Transformers for global contextual modeling. By leveraging multi-scale features from segmentation maps, MCT-Net guides the image generation process within a Diffusion Transformer architecture, enabling precise spatial control and enhanced realism. Our key innovation lies in exploiting spatial hierarchies and contextual information from segmentation maps to improve both the fidelity and diversity of generated defects. Additionally, we introduce a two-stage generation strategy: (1) encoding noisy defect images using a Variational Autoencoder (VAE), followed by (2) latent-space denoising via a Diffusion Transformer. Extensive experiments demonstrate that MCT-Net outperforms existing methods in image quality, diversity, and downstream defect detection performance, establishing a new benchmark for controlled defect synthesis.