We propose a voiceprint recognition method of newborn giant pandas (NGPs), named Long Short Term Memory (LSTM) and hybrid local-global Transformer Parallel Net (LTPNet), which extracts voiceprint features by parallelly utilizing LSTM and hybrid Local-Global Transformer (LGT). The LSTM branch of the LTPNet focuses on capturing the dynamic patterns of voiceprint. The local self-attention (LSA) is introduced in the Transformer branch to form Hybrid Local and Global Multi head Self-Attention Transformer Encoder (HLGAE), which simultaneously focus on short-term details of voiceprint signals (such as rapid changes in phonemes) and global context (such as the structure and prosody of the entire voiceprint sequence). The features extracted by LSTM and HLGAE are fed into a feature fusion module composed of two Residual Blocks and Self-Attention (RBSA). The loss of LTPNet is composed by cross-entropy, and a new defined feature diversity loss (FDL) in order to enhance the feature difference between two branches. The experiments shown that the performance of our proposed method can accurately recognize the voiceprints of NGPs in complex environments compared to the state-of-arts methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LSTM and Hybrid Local-Global Transformer Parallel Net: A Novel Method to Recognize the Calls of Newborn Giant Pandas

  • Yajie Xu,
  • Zhiwu Liao

摘要

We propose a voiceprint recognition method of newborn giant pandas (NGPs), named Long Short Term Memory (LSTM) and hybrid local-global Transformer Parallel Net (LTPNet), which extracts voiceprint features by parallelly utilizing LSTM and hybrid Local-Global Transformer (LGT). The LSTM branch of the LTPNet focuses on capturing the dynamic patterns of voiceprint. The local self-attention (LSA) is introduced in the Transformer branch to form Hybrid Local and Global Multi head Self-Attention Transformer Encoder (HLGAE), which simultaneously focus on short-term details of voiceprint signals (such as rapid changes in phonemes) and global context (such as the structure and prosody of the entire voiceprint sequence). The features extracted by LSTM and HLGAE are fed into a feature fusion module composed of two Residual Blocks and Self-Attention (RBSA). The loss of LTPNet is composed by cross-entropy, and a new defined feature diversity loss (FDL) in order to enhance the feature difference between two branches. The experiments shown that the performance of our proposed method can accurately recognize the voiceprints of NGPs in complex environments compared to the state-of-arts methods.