The task of speaker diarization aims to eliminate non-speaker vocal disturbances and label different speakers accordingly, ensuring that the audio contains only the voices of speakers, and addressing the challenge of “who spoke when”. In complex teaching scenarios, the existing models face the problem of gradient vanishing, which weakens the ability to effectively distinguish the speaker’s speech fragments. In addition, redundant speaker features significantly increase the dimensionality of the feature space, leading to the risk of dimensionality curse. To solve these problems, this paper proposes a classroom speaker diarization method based on multi-density feature clustering. Specifically, the feedforward sequential memory network is introduced, which combines cross-level hopping connections to enhance the information flow of deep modules, so as to solve the problem of gradient vanishing in the detection process. Secondly, inspired by manifold learning, the nonlinear dimensionality reduction is applied to the speaker embedding features, which solves the dimensionality problem caused by redundant speaker embedding. Extensive experiments on three public datasets and self-built datasets show that the proposed method has effectiveness and significant advantages in identifying speakers in the classroom teaching environment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Classroom Speaker Diarization Based on Multi-density Feature Clustering

  • Lin Jiang,
  • Jiatian Mei,
  • Ziyuan Cui,
  • Jingen Li,
  • Jun Wang

摘要

The task of speaker diarization aims to eliminate non-speaker vocal disturbances and label different speakers accordingly, ensuring that the audio contains only the voices of speakers, and addressing the challenge of “who spoke when”. In complex teaching scenarios, the existing models face the problem of gradient vanishing, which weakens the ability to effectively distinguish the speaker’s speech fragments. In addition, redundant speaker features significantly increase the dimensionality of the feature space, leading to the risk of dimensionality curse. To solve these problems, this paper proposes a classroom speaker diarization method based on multi-density feature clustering. Specifically, the feedforward sequential memory network is introduced, which combines cross-level hopping connections to enhance the information flow of deep modules, so as to solve the problem of gradient vanishing in the detection process. Secondly, inspired by manifold learning, the nonlinear dimensionality reduction is applied to the speaker embedding features, which solves the dimensionality problem caused by redundant speaker embedding. Extensive experiments on three public datasets and self-built datasets show that the proposed method has effectiveness and significant advantages in identifying speakers in the classroom teaching environment.