Artificial Intelligence Generated Content (AIGC) has achieved remarkable progress in recent years, particularly in unimodal applications like image generation. However, single modality approaches are inherently constrained in their ability to capture and represent comprehensive multimodal information. To overcome these limitations, researchers have increasingly integrated multimodal data, enhancing the understanding and generative capabilities of models. This paper reviews modern AIGC approaches, focusing on foundational models including Transformer, Generative Adversarial Networks and diffusion models, highlighting their roles in image and 3D shape generation. We further delve into examining methodologies to achieve high-fidelity and controllable 3D generation. Mainstream baselines and research directions are discussed, alongside critical challenges and potential solutions, aiming to advance multimodal AIGC technologies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Pixels to Shapes: Generative AI for 2D Images and 3D Models

  • Jianhui Huo,
  • Shijian Xu,
  • Xi Chen,
  • Weilong Peng,
  • Yangtao Wang,
  • Yan Wang,
  • Meie Fang

摘要

Artificial Intelligence Generated Content (AIGC) has achieved remarkable progress in recent years, particularly in unimodal applications like image generation. However, single modality approaches are inherently constrained in their ability to capture and represent comprehensive multimodal information. To overcome these limitations, researchers have increasingly integrated multimodal data, enhancing the understanding and generative capabilities of models. This paper reviews modern AIGC approaches, focusing on foundational models including Transformer, Generative Adversarial Networks and diffusion models, highlighting their roles in image and 3D shape generation. We further delve into examining methodologies to achieve high-fidelity and controllable 3D generation. Mainstream baselines and research directions are discussed, alongside critical challenges and potential solutions, aiming to advance multimodal AIGC technologies.