<p>Federated learning for speech emotion recognition faces fundamental challenges in simultaneously achieving high performance, privacy preservation, and model interpretability. This paper introduces FedSER-XAI, a novel framework that integrates Particle Swarm Optimization (PSO)-based feature selection, multi-stream cross-attention mechanisms, and graph-based feature extraction within an explainable federated learning architecture. Our approach combines Vision Transformer processing of mel-spectrograms with temporal-spatial graph convolutional networks to capture both contextual and structural speech relationships. The PSO algorithm achieves 78.1% dimensionality reduction (228<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\rightarrow\)</EquationSource> </InlineEquation>50 features) while improving discriminative power. The multi-stream architecture processes traditional acoustic features alongside novel graph-based representations derived from visibility and correlation graphs, fused through Transformer-based cross-attention mechanisms. Extensive evaluation on EMODB and SAVEE datasets demonstrates exceptional performance: 99.9% and 97.2% accuracy in centralized settings, with remarkable federated performance achieving global model accuracies of 99.7% (EMODB) and 97.2% (SAVEE) across 8 emotion-specialized clients, representing only 0.2% and 0.0% degradation compared to centralized training. The framework achieves rapid convergence within 10 communication rounds, representing minimal performance degradation (0.2% for EMODB) while preserving privacy. Cross-dataset evaluation on CREMA-D yields 68% accuracy, demonstrating reasonable generalization. The comprehensive explainability framework using SHAP and LIME provides global and local interpretations, validating that graph-based features contribute significantly to emotion discrimination. FedSER-XAI represents the first explainable federated speech emotion recognition system, advancing trustworthy AI for sensitive healthcare and human-computer interaction applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FedSER-XAI: PSO-optimized multi-stream cross-attention transformer with graph features for explainable federated speech emotion recognition

  • Eman Abdulrahman Alkhamali,
  • Arwa Abdulaziz Allinjawi,
  • Rehab Bahaaddin Ashari,
  • Mohammed Tawfik

摘要

Federated learning for speech emotion recognition faces fundamental challenges in simultaneously achieving high performance, privacy preservation, and model interpretability. This paper introduces FedSER-XAI, a novel framework that integrates Particle Swarm Optimization (PSO)-based feature selection, multi-stream cross-attention mechanisms, and graph-based feature extraction within an explainable federated learning architecture. Our approach combines Vision Transformer processing of mel-spectrograms with temporal-spatial graph convolutional networks to capture both contextual and structural speech relationships. The PSO algorithm achieves 78.1% dimensionality reduction (228 \(\rightarrow\) 50 features) while improving discriminative power. The multi-stream architecture processes traditional acoustic features alongside novel graph-based representations derived from visibility and correlation graphs, fused through Transformer-based cross-attention mechanisms. Extensive evaluation on EMODB and SAVEE datasets demonstrates exceptional performance: 99.9% and 97.2% accuracy in centralized settings, with remarkable federated performance achieving global model accuracies of 99.7% (EMODB) and 97.2% (SAVEE) across 8 emotion-specialized clients, representing only 0.2% and 0.0% degradation compared to centralized training. The framework achieves rapid convergence within 10 communication rounds, representing minimal performance degradation (0.2% for EMODB) while preserving privacy. Cross-dataset evaluation on CREMA-D yields 68% accuracy, demonstrating reasonable generalization. The comprehensive explainability framework using SHAP and LIME provides global and local interpretations, validating that graph-based features contribute significantly to emotion discrimination. FedSER-XAI represents the first explainable federated speech emotion recognition system, advancing trustworthy AI for sensitive healthcare and human-computer interaction applications.