To Enhance the Efficiency and Features of Text to Image Generation Using Neural Network Models
摘要
Generation of photorealistic images have multiple utilization in the field of photo editing, fashion, product, game designing, painting and so on. Individuals involved in these fields are in need of visualizing their own ideas. There isn’t any specific solution that serves them. Existing works result in limited features and low-quality images with less accuracy. It's important for developers to visualize their ideas before moving into the development phase. The proposed system is designed in such a manner to mitigate all these issues. It generates image with the help of a pre trained model. The accuracy is improved by ranking, denoising and upscaling the generated 2D image. By combining different models, a new algorithm is proposed for generating the high-resolution 2D image. To improve the features of the text to image generation, character 3D modelling and video generation are included. The 3D modelling feature displays 3D model by taking a 2D character or human image as input. The Video feature generates the video by taking list of text prompts as input. Best models are chosen from the survey to implement video generation and 3D model construction. The proposed system will stand as a unique solution for users to envision their thoughts and ideas. This aids several developers and designers, by simplifying work and effectively utilizing time.