This paper introduces ViscoNet, a novel one-branch-adapter architecture for concurrent spatial and visual conditioning. Our lightweight model requires trainable parameters and dataset size multiple orders of magnitude smaller than the current state-of-the-art IP-Adapter. However, our method successfully preserves the generative power of the frozen text-to-image (T2I) backbone. Notably, it excels in addressing mode collapse, a pervasive issue previously overlooked. Our novel architecture demonstrates outstanding capabilities in achieving a harmonious visual-text balance, unlocking unparalleled versatility in various human image generation tasks, including pose re-targeting, virtual try-on, stylization, person re-identification, and textile transfer. Demo and code are available from project page https://soon-yau.github.io/visconet/ .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Flow chart depicting a process with a central node labeled "VisionNet." The chart likely represents a system or model related to vision processing, with "VisionNet" as the focal point. The layout suggests a network or sequence of steps, though specific connections or additional nodes are not visible. : Bridging and Harmonizing Flow chart with a central node labeled "Visual" in pink text. The chart likely represents a process or concept related to visual elements, but no additional nodes or connections are visible. The background is white, emphasizing the central text. and Textual Conditioning for  Flow chart with a central node labeled "ConvexNet." The chart likely represents a process or system related to this term, with potential connections or pathways not visible in the image.

  • Soon Yau Cheong,
  • Armin Mustafa,
  • Andrew Gilbert

摘要

This paper introduces ViscoNet, a novel one-branch-adapter architecture for concurrent spatial and visual conditioning. Our lightweight model requires trainable parameters and dataset size multiple orders of magnitude smaller than the current state-of-the-art IP-Adapter. However, our method successfully preserves the generative power of the frozen text-to-image (T2I) backbone. Notably, it excels in addressing mode collapse, a pervasive issue previously overlooked. Our novel architecture demonstrates outstanding capabilities in achieving a harmonious visual-text balance, unlocking unparalleled versatility in various human image generation tasks, including pose re-targeting, virtual try-on, stylization, person re-identification, and textile transfer. Demo and code are available from project page https://soon-yau.github.io/visconet/ .