EEG-Driven Music Generation with Latent Discrete Diffusion Models
摘要
Estimating music perceived by individuals from electroencephalography (EEG) signals holds considerable promise for both medical and engineering applications. In this study, we propose an EEG-driven music generation framework based on latent discrete diffusion models, designed to generate music conditioned on EEG recordings. The proposed method comprises two primary stages. First, a variational autoencoder (VAE) is employed to extract informative latent representations from EEG signals, yielding a conditional vector that encapsulates neural activity patterns associated with auditory perception. Second, a discrete diffusion model operates in the latent space, leveraging these EEG-derived feature vectors to generate music data that reflect the underlying neural responses to music stimuli. We evaluate the proposed approach using real EEG data recorded from human subjects during music listening. Given the inherently noisy nature of EEG signals, ensuring robustness to such noise is a critical challenge. Experimental results demonstrate that our model effectively generates music while maintaining resilience against noise, highlighting its potential as a novel machine learning-based approach for EEG-driven music generation and broader cross-modal neural-audio synthesis.