Decoding deception: interpretable multimodal models for audio–visual Lie detection
摘要
Lie detection is a challenging problem at the intersection of psychology, ethics, law, and computational sciences. It requires approaches that balance predictive performance, interpretability, and operational efficiency, taking into account privacy and security issues. This study presents an interpretable framework based on classical machine learning classifiers and structured multimodal features. It has been designed to favor explainability and reproducibility over architectural complexity. Unlike recent deep learning-based systems, the proposed method leverages numerical audio embeddings extracted with VGGish and aggregated facial descriptors derived from OpenFace, combined at the feature level to enable robust, interpretable decision-making. A systematic evaluation is conducted on three benchmark datasets, Bag-of-Lies, Real-Life Trial Dataset, and DOLOS, considering both unimodal (audio or video) and multimodal configurations with feature-level fusion. Six machine learning classifiers (MLP, Random Forest, Logistic Regression, XGBoost, SVM, and LDA) were tested for each channel to determine suitability and performance. Feature-level multimodal integration merges audio and video representations into a shared space, preserving the interpretability and computational efficiency of the single channels. The proposed framework achieves competitive performance across all datasets, with peak multimodal Accuracies of 0.74 for Bag-of-Lies, 0.92 for the Real-Life Trial Dataset, and 0.70 for DOLOS. These results demonstrate that structured, understandable machine learning pipelines can offer transparency, efficiency, and methodological reliability appropriate for real-world implementation, making them an attractive alternative to more complex state-of-the-art designs.