Second-order wavelet-based multi-frequency music generation
摘要
Thanks to the advancement of artificial intelligence algorithms, music generation has made significant progress in both quality and duration. Currently, music generated based on human text prompts can achieve high preference in terms of continuity and melody. However, most studies have overlooked the importance of melody (high and low frequencies) in music generation. In fact, the output of current diffusion-based music generation models tends to average probabilities, making it difficult for the models to effectively generate or process high and low frequency signals. In this paper, we construct a second-order wavelet autoencoder to decompose waveforms into low and high frequency components. Subsequently, we use a diffusion model (DM) to model the concatenation of different frequency signals. Additionally, we propose a frequency-domain awareness module to enhance the DM’s ability to handle frequencies at different time steps. Compared to the most advanced work, our method performs better in preference under similar conditions of generation quality and duration.or equations are permitted.