Audio Source Separation via Language-Queried Neural Networks
摘要
The goal of language-queried audio source separation (LASS), which is centred on separating a target audio source from a mixture using a natural language query, is presented in this work. The tight link between audio sources and natural language descriptions is what makes this endeavour so complex. We suggest LASS-Net, an end-to-end neural network built to jointly handle language and audio input, as a solution to this problem. To accurately separate the target source compatible with the language query, LASS-Net uses a Transformer-based query network and a ResUNet-based separation network. Using a dataset generated from AudioCaps, we assessed our method and found considerable improvements over baseline techniques. The outcomes further demonstrate LASS-Net’s strong generalization abilities in managing a variety of human-annotated descriptions, highlighting its potential for practical uses. Public access to the source code and split audio samples encourages more study and advancement in this cutting-edge area.