<p>Speaker identification is a fundamental task in speech processing, with applications spanning security, authentication, human-computer interaction, and personalized services. Conventional approaches often rely on handcrafted acoustic features, such as MFCC, chroma, and spectral descriptors, as well as deep learning-based representations, which capture certain characteristics of speech signals but may not fully exploit the complementary information available across feature types. Recent advances in large language models (LLMs) and Transformer-based architectures, such as Whisper, offer rich, high-level embeddings that encapsulate both linguistic and paralinguistic cues from audio signals. In this study, we propose an enhanced speaker identification framework that integrates Transformer-based Whisper embeddings with acoustic features to form a multimodal feature representation. To effectively combine these heterogeneous feature sets, Canonical Correlation Analysis (CCA) is employed, extracting components that are maximally correlated across the modalities, thereby capturing complementary information that enhances speaker discrimination. The proposed approach is evaluated on two benchmark speech dataset. On the TIMIT dataset, the proposed Acoustic+Whisper+CCA representation achieves up to 97.02% accuracy with SVM, On the VoxCeleb1 dataset, which includes more diverse and challenging real-world recordings, the proposed method reaches 92.60% accuracy with SVM. Experimental results indicate that fusing Whisper embeddings with acoustic features via CCA significantly improves identification accuracy, robustness, and generalization across different speakers. The contribution is a system-level multimodal fusion built from established components; the novelty lies in the CCA-based joint-space fusion and its robustness across controlled (TIMIT) and real-world (VoxCeleb1) conditions at modest cost, not in any new architecture.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced speaker identification using transformer-based whisper and acoustic multimodal features based on canonical correlation analysis

  • Mehmet Bilal Er

摘要

Speaker identification is a fundamental task in speech processing, with applications spanning security, authentication, human-computer interaction, and personalized services. Conventional approaches often rely on handcrafted acoustic features, such as MFCC, chroma, and spectral descriptors, as well as deep learning-based representations, which capture certain characteristics of speech signals but may not fully exploit the complementary information available across feature types. Recent advances in large language models (LLMs) and Transformer-based architectures, such as Whisper, offer rich, high-level embeddings that encapsulate both linguistic and paralinguistic cues from audio signals. In this study, we propose an enhanced speaker identification framework that integrates Transformer-based Whisper embeddings with acoustic features to form a multimodal feature representation. To effectively combine these heterogeneous feature sets, Canonical Correlation Analysis (CCA) is employed, extracting components that are maximally correlated across the modalities, thereby capturing complementary information that enhances speaker discrimination. The proposed approach is evaluated on two benchmark speech dataset. On the TIMIT dataset, the proposed Acoustic+Whisper+CCA representation achieves up to 97.02% accuracy with SVM, On the VoxCeleb1 dataset, which includes more diverse and challenging real-world recordings, the proposed method reaches 92.60% accuracy with SVM. Experimental results indicate that fusing Whisper embeddings with acoustic features via CCA significantly improves identification accuracy, robustness, and generalization across different speakers. The contribution is a system-level multimodal fusion built from established components; the novelty lies in the CCA-based joint-space fusion and its robustness across controlled (TIMIT) and real-world (VoxCeleb1) conditions at modest cost, not in any new architecture.