DreamDiffuser: From Textual Prompts to Multimodal Story Continuations
摘要
This research paper delves into the comprehensive exploration of DreamDiffuser, an innovative storytelling platform that seamlessly integrates the pipelines of Stable Diffusion XL for image generation and an advanced Language Model (LLM) for dialogue creation. To enhance user engagement, DreamDiffuser introduces the concept of interactive and creative storytelling, allowing users to input brief storylines that are transformed into visually engaging comics. With a user-friendly interface and a commitment to clarity, the platform empowers users to express their narrative ideas without relying on industry-specific jargon. The inclusion of features like anime-style background image generation and DuBaGAN, a specialised module for creative story continuation based on user-provided images, further amplifies the storytelling experience. Through AI multimodal analysis, DreamDiffuser interprets user-provided images, extracting essential features and generating contextually relevant text for a comprehensive and engaging narrative. The paper introduces a novel method for generating multimodal story continuations by converting textual prompts into captivating narratives. Leveraging the capabilities of Stability AI and OpenAI pipelines, this method combines language prompts with visual components, going beyond conventional limits to present a vibrant fusion of creativity and technology. The resulting narratives reflect the synergy between the Stability AI and OpenAI platforms, offering users an enriched storytelling experience that transcends traditional boundaries.