Multimodal Creativity State Detection from Speech and Voice
摘要
Modern society has been shaped and advanced by creativity, which has an impact on many facets of human expression and interaction. This research explores the intersection of technology and human interaction by proposing a computational multimodal creativity state identification from speech and voice. This study examines the relationship between language elements and emotion detection, as well as the emotional dimensions (arousal and valence), with a focus on the dynamic and process of identifying creative states through the analysis of speech and voice modalities. To achieve the creativity state detection, it uses a CNN model that was trained using four datasets: Crema-D, Ravdess, Savee, and Tess. A linguistic prompt creativity test is included in the study to provide external validation of the suggested model. The findings show a strong relationship between emotion and emotional dimension, alongside the machine-based detection methods.