Revisiting U-Net: a foundational backbone for modern generative AI
摘要
This survey explores the evolution and application of U-Net in generative AI, highlighting its success across various modalities, including image, text, audio, video, 3D, and pose/action generation. Initially designed for biomedical segmentation, U-Net has been adapted and enhanced with architectural innovations such as normalization techniques, self and cross-attention mechanisms, and residual connections. These advancements have made U-Net a powerful backbone for modern generative models in diffusion-based frameworks, GANs, and autoregressive architectures. The survey comprehensively reviews U-Net’s modality-specific applications, from high-resolution image synthesis and text-to-image generation to speech enhancement, video generation, 3D reconstruction, and pose/action generation. Despite its widespread success, U-Net faces challenges in computational efficiency, contextual understanding, and scalability for multimodal tasks. Future directions focus on optimizing U-Net for lightweight and real-time applications, enhancing its contextual awareness, and improving its integration with emerging architectures like transformers and diffusion models.