Enhancing interpretability in video-based personality trait recognition using SHAP analysis
摘要
Deep learning models show promise for predicting personality traits from videos, but existing methods for understanding their predictions often lack global interpretability. This study addresses this gap by using SHAP (Shapley additive explanations) values to analyze a lightweight, multimodal deep learning model designed for personality recognition. Using the ChaLearn First Impressions V2 dataset, we investigated how visual, audio, and textual features contribute to the model’s predictions. SHAP analysis provided both global and local explanations of feature importance, offering insights into the model’s decision-making process. Feature reduction based on these insights was also explored, using KL divergence to assess the impact on prediction accuracy. This research demonstrates the effectiveness of SHAP in enhancing the interpretability of multimodal deep learning models for personality trait prediction. The findings can inform the development of more transparent and trustworthy AI tools for applications such as recruitment, human–computer interaction, psychological assessment, and personalized content recommendation.