Machine Learning Based Multimodal Opinion Mining
摘要
In today's digital era, understanding human emotions expressed through various communication channels is vital. Multimodal opinion mining method integrates audio, video, and visual information, each contributing a distinct perspective. In this paper, we have implemented models for each modality to capture the unique subtleties and patterns of various sets of information, improving opinion accuracy and depth. First, we adopt a multimodal fusion strategy for different perspectives on human expression and emotion using Tone, pitch, and patterns of speech that indicate emotions in audio data. Video data contains emotional information from facial expressions, body language, and context. Our approach involves developing machine learning models for each modality to capture its distinct traits and patterns. For audio, speech sentiment analysis models are implemented, and for video, face expression recognition and gesture models to provide a comprehensive model using machine learning. The final model score should indicate the importance of each modality in emotion interpretation in deep learning model.