In the realm of audio signal processing, isolating human speech from background noise poses a significant challenge. In noisy environments, multiple human audio signals may overlap, and the presence of ambient noise complicates accurate individual speech extraction. Recent advancements leverage Transformer based architectures to efficiently extract human speech signals from mixtures of overlapping background sounds, addressing the shortcomings of traditional techniques. However, existing Transformer models face limitations, including restricted scalability to scenarios with a high number of overlapping speakers, increased computational complexity, and parameter inefficiency. In this paper, we propose an architectural framework that leverages the combined speech separation capabilities of two such novel transformer architectures to enhance multi-speaker separation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FusionNet: Leveraging Dual Speech Separation Networks for Enhanced Multi-speaker Isolation

  • Sumedh Ravindran,
  • Shreyas Mallesh,
  • Raghavendra S. Bhatagunaki,
  • Srikrishna R. Chitnis,
  • Shylaja S. Sharath

摘要

In the realm of audio signal processing, isolating human speech from background noise poses a significant challenge. In noisy environments, multiple human audio signals may overlap, and the presence of ambient noise complicates accurate individual speech extraction. Recent advancements leverage Transformer based architectures to efficiently extract human speech signals from mixtures of overlapping background sounds, addressing the shortcomings of traditional techniques. However, existing Transformer models face limitations, including restricted scalability to scenarios with a high number of overlapping speakers, increased computational complexity, and parameter inefficiency. In this paper, we propose an architectural framework that leverages the combined speech separation capabilities of two such novel transformer architectures to enhance multi-speaker separation.