Speaker diarization is vital in contexts like police interrogations, where it enhances the security and personalization of data access and improves confidentiality in multi-speaker environments. The transcription of low-quality forensic audio recordings is challenging, as they are often marred by unclear speech and impede the accuracy of conventional Automatic Speech Recognition (ASR) systems. This paper evaluates the efficacy of traditional machine learning algorithms—Support Vector Machine (SVM), Decision Tree Classifier, Random Forest Classifier, and XGBoost in gender classification from voice samples for speaker diarization systems. These systems are critical in contexts like police interrogations, where they enhance data security and improve confidentiality in multi-speaker environments. We test these algorithms against real-world data, simulating practical conditions to ensure robustness. Our findings reveal that ensemble methods, particularly Random Forest and XGBoost, demonstrate high accuracy and strong generalizability when dealing with unfiltered, real-world audio data. XGBoost shows significant resistance to overfitting, making it highly suitable for secure voice-driven applications. This study aids in algorithm selection for speaker diarization tasks. It addresses gaps in forensic audio transcription accuracy, thereby enhancing the reliability of transcriptions and reducing risks of erroneous interpretations in legal contexts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Speaker Diarization in Forensic Audio: A Comparative Analysis of Machine Learning Algorithms for Gender Classification

  • Rahmat Ullah,
  • Ikram Asghar,
  • Gareth Evans,
  • Rab Nawaz,
  • Saeed Akbar,
  • Dorothy Anne Roberts

摘要

Speaker diarization is vital in contexts like police interrogations, where it enhances the security and personalization of data access and improves confidentiality in multi-speaker environments. The transcription of low-quality forensic audio recordings is challenging, as they are often marred by unclear speech and impede the accuracy of conventional Automatic Speech Recognition (ASR) systems. This paper evaluates the efficacy of traditional machine learning algorithms—Support Vector Machine (SVM), Decision Tree Classifier, Random Forest Classifier, and XGBoost in gender classification from voice samples for speaker diarization systems. These systems are critical in contexts like police interrogations, where they enhance data security and improve confidentiality in multi-speaker environments. We test these algorithms against real-world data, simulating practical conditions to ensure robustness. Our findings reveal that ensemble methods, particularly Random Forest and XGBoost, demonstrate high accuracy and strong generalizability when dealing with unfiltered, real-world audio data. XGBoost shows significant resistance to overfitting, making it highly suitable for secure voice-driven applications. This study aids in algorithm selection for speaker diarization tasks. It addresses gaps in forensic audio transcription accuracy, thereby enhancing the reliability of transcriptions and reducing risks of erroneous interpretations in legal contexts.