<p>Thanks to the advancement of artificial intelligence algorithms, music generation has made significant progress in both quality and duration. Currently, music generated based on human text prompts can achieve high preference in terms of continuity and melody. However, most studies have overlooked the importance of melody (high and low frequencies) in music generation. In fact, the output of current diffusion-based music generation models tends to average probabilities, making it difficult for the models to effectively generate or process high and low frequency signals. In this paper, we construct a second-order wavelet autoencoder to decompose waveforms into low and high frequency components. Subsequently, we use a diffusion model (DM) to model the concatenation of different frequency signals. Additionally, we propose a frequency-domain awareness module to enhance the DM’s ability to handle frequencies at different time steps. Compared to the most advanced work, our method performs better in preference under similar conditions of generation quality and duration.or equations are permitted.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Second-order wavelet-based multi-frequency music generation

  • Kun Bai,
  • Jiayu Zhang

摘要

Thanks to the advancement of artificial intelligence algorithms, music generation has made significant progress in both quality and duration. Currently, music generated based on human text prompts can achieve high preference in terms of continuity and melody. However, most studies have overlooked the importance of melody (high and low frequencies) in music generation. In fact, the output of current diffusion-based music generation models tends to average probabilities, making it difficult for the models to effectively generate or process high and low frequency signals. In this paper, we construct a second-order wavelet autoencoder to decompose waveforms into low and high frequency components. Subsequently, we use a diffusion model (DM) to model the concatenation of different frequency signals. Additionally, we propose a frequency-domain awareness module to enhance the DM’s ability to handle frequencies at different time steps. Compared to the most advanced work, our method performs better in preference under similar conditions of generation quality and duration.or equations are permitted.