Audioverse: Personalized Story-Telling Using Voice Cloning
摘要
The surge in global podcast and story listeners, now totaling 504.9 million and comprising 23.5% of internet users, highlights the growing demand for advanced narrative tools. This study explores the transformative potential of voice in storytelling by developing a deep learning model designed to generate diverse narratives based on user prompts. Utilizing advanced voice cloning techniques, the model aims to embed nuanced emotions within these narratives, merging human emotion with technological capabilities to create a sustainable tool for enhancing individual well-being. The model was constructed using PyTorch, TensorFlow, and Convolutional Neural Networks (CNN), with voice cloning powered by BARK, an open-source transformer-based model. Audioverse promises to significantly enrich human connections and revolutionize the storytelling landscape.