FashionAtlas: Enhancing Semantics and Control in Multimodal Fashion Image Editing
摘要
In recent years, with the rapid advancement of diffusion-based image generation and virtual try-on technologies, novel research methodologies integrating multimodal inputs such as text and images have emerged. These approaches have become a significant direction in fashion image editing. However, most existing datasets only provide images with basic attribute annotations, lacking systematic support for explicit structural constraints such as sketches and structure maps. This makes it difficult to ensure structural consistency and controllability during the editing and generation processes. Simultaneously, existing textual descriptions often remain confined to superficial semantics like color and category, lacking multidimensional annotation information regarding material properties and design styles. This limits the model’s ability to understand and generalize complex instructions and authentic fashion contexts. To address this, we propose constructing a novel benchmark dataset: FashionAtlas. This dataset provides more refined sketches as structural constraints to enhance the controllability of the generated results. Simultaneously, we introduce more detailed and layered prompts generated by large language models (LLMs), integrating structured clothing attribute ontologies with specialized terminology systems. This enhances the model’s comprehension of fashion semantics and styling logic, thereby better supporting complex editing and generation requirements. Experimental results demonstrate that FashionAtlas outperforms traditional image-annotation-only baseline dataset in terms of generation accuracy, semantic consistency, detail fidelity, and user controllability. This benchmark dataset aims to fill the current gap in combining structured constraints with fine-grained semantic prompts, providing a new research standard for multimodal control editing and in-context generation tasks in the fashion domain.