GLANCE—Guided Language Through Autoregression Establishing Natural and Classifier-Free Editing
摘要
In this study, researchers aimed to simplify text conversion into images using the latest text-to-image generation methods. While these methods have improved the quality and relevance of generated images, certain crucial questions remained unanswered, limiting their practicality and overall quality. To address these issues, the researchers introduced a novel text-to-image method. This method allows for better control of the scene depicted in the image through text, enhances the tokenization process by incorporating specific knowledge about key image regions such as faces and important objects, and provides guidance to the transformer model without needing a classifier. The outcome of this work was a model that achieved state-of-the-art results in terms of image quality and human evaluation, enabling the generation of high-fidelity 512 × 512-pixel images. Moreover, this method introduced new capabilities, including scene editing, text editing with reference scenes, handling out-of-distribution text prompts, and generating story illustrations.